Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

AI summary

An inference engine called InferredThoughts has been developed for MoE models that exceed VRAM and RAM capacity. This engine allows models like Qwen3.8-Flash-Next 177B NVFP4(119GiB) to run on a single 16 GB RTX 5060 Ti with 32 GB RAM, achieving 9-10 tok/s via SSD streaming. Most of the model resides on the SSD, with experts loaded on demand. This setup utilizes 20 GiB in memory and 99 GiB on SSD, including 48.5 GiB for routed experts and a 50.7 GiB n-gram table.

Why this one

This report details a novel inference engine that allows large MoE models to run on consumer-grade GPUs by streaming most of the model from SSD, unlike traditional methods requiring full in-memory loading.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 28, 2026, 00:00 UTC

Ingested
Sep 28, 2026, 00:00
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com