RCreddit.com·
暂不在当前实时榜单
Qwen 3.8 27B UD-IQ4_XS even faster on 16GB CUDA
A new development, building on Raymond's KV cache streaming fork, allows Qwen 3.8 27B UD-IQ4_XS to run even faster on 16GB CUDA. The innovation involves speculative decoding (MTP or DFlash2) by ejecting the speculative model from VRAM when not fully utilized and reloading it when context is low. Optimal performance is achieved by ejecting the model after a certain number of pages (256 bytes each) once KV streaming begins.
This report details a method for speculative decoding on 16GB CUDA, unlike earlier approaches that did not dynamically manage VRAM by ejecting and reloading the speculative model.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月14日 05:01 UTC
- 收录
- 2026年9月14日 05:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →