RCreddit.com·
暂不在当前实时榜单
llama: add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp
A new GPU cache for Mixture-of-Experts (MoE) models, specifically for experts kept in host memory, has been added to llama.cpp via Pull Request #29887. This update, merged from https://github.com/ggml-org/llama.cpp/pull/30112, is expected to provide a significant speedup for MoE models that cannot fully fit into VRAM, potentially benefiting users with limited GPU resources.
This update specifically targets MoE models that exceed VRAM, unlike previous optimizations that focused on models fully fitting within GPU memory.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月7日 20:00 UTC
- 收录
- 2026年10月7日 20:00
- 来源类型
- 开发者社区
讨论趋势
→ 平稳
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →