Skip to content
RCreddit.com·
Not on the current live radar

Faster prompt lookup drafting in llama.cpp

AI summary

Llama.cpp has achieved 42x faster prompt lookup drafting through an updated approach that utilizes context, dynamic, and static caches. The system employs specific thresholds, such as $(a_1, a_2, a_3, a_4) = (2, 2, 1, 1)$ and $(p_1, p_2, p_3, p_4) = (0.66, 0.5, 0.5, 0.5)$ for the context cache, to determine draft tokens. If no candidate passes, it falls back to the static cache. An std::vector is used for memory efficiency, with a sorted std::vector addressing search latency for heavy-tailed n-gram distributions.

Why this one

This report details the specific thresholds and cache prioritization (context, dynamic, static) that enable llama.cpp's 42x speedup, unlike general announcements of performance gains.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 27, 2026, 03:00 UTC

Ingested
Sep 27, 2026, 03:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com