跳到正文
RCreddit.com·
暂不在当前实时榜单

You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

AI 摘要

Qwen3.8-Flash-Next, and potentially other qwen4exp-based models, can offload most of their KV cache to system RAM, significantly reducing VRAM requirements. This allows for running models at maximum context length without KV cache quantization, even with quantizations that barely fit in VRAM. While a full KV cache row is 2,048 B per token, only about 66 B per token per layer needs to remain on the GPU, leading to minimal decode slowdown despite the PCIe 4.0 x16 slot's 32 GiB/s bandwidth limiting host-resident cache to about 5 tokens per second.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月16日 20:00 UTC

收录
2026年9月16日 20:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com