RCreddit.com·
暂不在当前实时榜单
MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?
A developer is seeking ideas to optimize MoE SSD streaming on a 64 GB Mac mini, where the GPU experiences a 27% decode wait on experts. They are already implementing lookahead guess fetching for subsequent layers' experts with 72% accuracy. The developer is considering expanding this lookahead strategy and using a separate staging buffer for guessed experts to prevent eviction of frequently used experts. They provided a GitHub link for further context.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月5日 18:00 UTC
- 收录
- 2026年10月5日 18:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →