Faster prompt lookup drafting in llama.cpp
Llama.cpp has achieved 42x faster prompt lookup drafting through an updated approach that utilizes context, dynamic, and static caches. The system employs specific thresholds, such as $(a_1, a_2, a_3, a_4) = (2, 2, 1, 1)$ and $(p_1, p_2, p_3, p_4) = (0.66, 0.5, 0.5, 0.5)$ for the context cache, to determine draft tokens. If no candidate passes, it falls back to the static cache. An std::vector is used for memory efficiency, with a sorted std::vector addressing search latency for heavy-tailed n-gram distributions.
This report details the specific thresholds and cache prioritization (context, dynamic, static) that enable llama.cpp's 42x speedup, unlike general announcements of performance gains.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 27, 2026, 03:00 UTC
- Ingested
- Sep 27, 2026, 03:00
- Source type
- Dev community
Full text isn't available here.
Read at source →