Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs
Part 4 of the Qwen3.8-Flash-Next series details a method to achieve 2.2-2.5x faster prefill speeds by moving the expert cache off the GPU during prompt processing. This optimization significantly reduces the "time to first token" for prompts, cutting an 8k prompt from 82 seconds to 37 seconds and a 119k prompt from 1461 seconds to 575 seconds. While prefill performance saw substantial gains, decode speeds remained largely consistent, showing only minor changes of -1% to +2%.
Why this oneThis report details a specific optimization, unlike previous parts in the series that focused on other aspects like UD-Q4_K_XL and MTP stacking, by moving the expert cache off the GPU.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 10, 2026, 12:00 UTC
- Ingested
- Sep 10, 2026, 12:00
- Source type
- Dev community
Full text isn't available here.
Read at source →