Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs

AI summary

Part 4 of the Qwen3.8-Flash-Next series details a method to achieve 2.2-2.5x faster prefill speeds by moving the expert cache off the GPU during prompt processing. This optimization significantly reduces the "time to first token" for prompts, cutting an 8k prompt from 82 seconds to 37 seconds and a 119k prompt from 1461 seconds to 575 seconds. While prefill performance saw substantial gains, decode speeds remained largely consistent, showing only minor changes of -1% to +2%.

Why this one

This report details a specific optimization, unlike previous parts in the series that focused on other aspects like UD-Q4_K_XL and MTP stacking, by moving the expert cache off the GPU.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 10, 2026, 12:00 UTC

Ingested
Sep 10, 2026, 12:00
Source type
Dev community
Article

Full text isn't available here.

Read at source →
Source·reddit.com