Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080
A user shared their experience running Qwen 3.8 Flash Next q4_k_m with 130k context and q8 cache on a system with 16GB VRAM and 64GB RAM, achieving 15-20 t/s on a 4080. Key factors for success include using the AtomicChat AD-4.27bpw Q4_K_M target, the Unsloth MTP head from a specific pr-mtp-fix branch, and the --spec-draft-cpu-moe flag. This configuration allows draft experts to reside in RAM, enabling hot experts to utilize the GPU, resulting in 16.5 tg / 350 pp at 131k with q8 KV.
This report details specific, previously unshared configuration flags and branches that enable 130k context on 16GB VRAM, unlike general guides.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 24, 2026, 12:02 UTC
- Ingested
- Sep 24, 2026, 12:02
- Source type
- Dev community
Full text isn't available here.
Read at source →