Faster prompt lookup drafting in llama.cpp
Llama.cpp has achieved 42x faster prompt lookup drafting through an updated approach that utilizes context, dynamic, and static caches. The system employs specific thresholds, such as $(a_1, a_2, a_3, a_4) = (2, 2, 1, 1)$ and $(p_1, p_2, p_3, p_4) = (0.66, 0.5, 0.5, 0.5)$ for the context cache, to determine draft tokens. If no candidate passes, it falls back to the static cache. An std::vector is used for memory efficiency, with a sorted std::vector addressing search latency for heavy-tailed n-gram distributions.
This report details the specific thresholds and cache prioritization (context, dynamic, static) that enable llama.cpp's 42x speedup, unlike general announcements of performance gains.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月27日 03:00 UTC
- 收录
- 2026年9月27日 03:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →