RCreddit.com·
暂不在当前实时榜单
Qwen3.8-Flash-Next on 2x3090 + DDR4, part 4: 2.2-2.5x faster prefill by kicking the expert cache off the GPU while the prompt runs
Part 4 of the Qwen3.8-Flash-Next series details a method to achieve 2.2-2.5x faster prefill speeds by moving the expert cache off the GPU during prompt processing. This optimization significantly reduces the "time to first token" for prompts, cutting an 8k prompt from 82 seconds to 37 seconds and a 119k prompt from 1461 seconds to 575 seconds. While prefill performance saw substantial gains, decode speeds remained largely consistent, showing only minor changes of -1% to +2%.
为什么是这条This report details a specific optimization, unlike previous parts in the series that focused on other aspects like UD-Q4_K_XL and MTP stacking, by moving the expert cache off the GPU.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月10日 12:00 UTC
- 收录
- 2026年9月10日 12:00
- 来源类型
- 开发者社区
正文
本站未收录正文。
前往源站阅读 →