跳到正文
RCreddit.com·
暂不在当前实时榜单

Hot Expert Reload on GPU is what this community needs

AI 摘要

A Reddit user suggests that Llama maintainers implement "Hot Expert Reload on GPU" to improve decode speed for Mixture-of-Experts (MOE) models. This feature would significantly enhance performance on GPUs like the 3090, making models such as Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, and GLM 5.3 Flash more usable locally, with speeds approaching full VRAM offload when using multiple cards.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月12日 06:01 UTC

收录
2026年9月12日 06:01
来源类型
开发者社区
正文

本站未收录正文。

前往源站阅读 →
来源·reddit.com