RCreddit.com·
暂不在当前实时榜单
Hot Expert Reload on GPU is what this community needs
A Reddit user suggests that Llama maintainers implement "Hot Expert Reload on GPU" to improve decode speed for Mixture-of-Experts (MOE) models. This feature would significantly enhance performance on GPUs like the 3090, making models such as Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, and GLM 5.3 Flash more usable locally, with speeds approaching full VRAM offload when using multiple cards.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月12日 06:01 UTC
- 收录
- 2026年9月12日 06:01
- 来源类型
- 开发者社区
正文
本站未收录正文。
前往源站阅读 →