Skip to content
RCreddit.com·
Not on the current live radar

You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

AI summary

Qwen3.8-Flash-Next, and potentially other qwen4exp-based models, can offload most of their KV cache to system RAM, significantly reducing VRAM requirements. This allows for running models at maximum context length without KV cache quantization, even with quantizations that barely fit in VRAM. While a full KV cache row is 2,048 B per token, only about 66 B per token per layer needs to remain on the GPU, leading to minimal decode slowdown despite the PCIe 4.0 x16 slot's 32 GiB/s bandwidth limiting host-resident cache to about 5 tokens per second.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 16, 2026, 20:00 UTC

Ingested
Sep 16, 2026, 20:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com