llama: add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp
A new GPU cache for Mixture-of-Experts (MoE) models, specifically for experts kept in host memory, has been added to llama.cpp via Pull Request #29887. This update, merged from https://github.com/ggml-org/llama.cpp/pull/30112, is expected to provide a significant speedup for MoE models that cannot fully fit into VRAM, potentially benefiting users with limited GPU resources.
This update specifically targets MoE models that exceed VRAM, unlike previous optimizations that focused on models fully fitting within GPU memory.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 7, 2026, 20:00 UTC
- Ingested
- Oct 7, 2026, 20:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →