Skip to content
RCreddit.com·
Not on the current live radar

llama: add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp

AI summary

A new GPU cache for Mixture-of-Experts (MoE) models, specifically for experts kept in host memory, has been added to llama.cpp via Pull Request #29887. This update, merged from https://github.com/ggml-org/llama.cpp/pull/30112, is expected to provide a significant speedup for MoE models that cannot fully fit into VRAM, potentially benefiting users with limited GPU resources.

Why this one

This update specifically targets MoE models that exceed VRAM, unlike previous optimizations that focused on models fully fitting within GPU memory.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 7, 2026, 20:00 UTC

Ingested
Oct 7, 2026, 20:00
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Full text isn't available here.

Read at source →
Source·reddit.com