RCreddit.com·
Not on the current live radar
Running Qwen3.8-Flash-Next locally on a 12GB VRAM card
A user successfully ran Qwen3.8-Flash-Next on an RTX 4070 12GB VRAM card, achieving over 20 tokens/second. This was accomplished by applying PR #28243, which enables a 1.78 GB shared-Q4_K_M compact head, and using the -ncmoe 45 setting. This configuration resulted in 77–96% acceptance rates across various tasks like coding, summarization, and creative generation.
This report uniquely details the specific PR (#28243) and configuration (-ncmoe 45) that enabled Qwen3.8-Flash-Next to break the 20 t/s barrier on 12GB VRAM, unlike general performance claims.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 15, 2026, 01:01 UTC
- Ingested
- Sep 15, 2026, 01:01
- Source type
- Dev community
Full text isn't available here.
Read at source →