RCreddit.com·
Not on the current live radar
ggml-cpu: tiled mul_mat for k-quants by jbooth · Pull Request #27851 · ggml-org/llama.cpp
A recent Pull Request #27851 to ggml-org/llama.cpp by jbooth introduces a significant improvement for CPU prompt processing. This update, titled "ggml-cpu: tiled mul_mat for k-quants," aims to achieve 3-7x faster CPU mul_mat operations. The enhancement is attributed to the use of VNNI, with the developer noting its minimal complexity. This development is expected to boost the performance of llama.cpp on CPUs.
This pull request is the first to claim a 3-7x speedup for CPU mul_mat operations in llama.cpp, unlike previous updates that offered more modest gains.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 26, 2026, 20:00 UTC
- Ingested
- Sep 26, 2026, 20:00
- Source type
- Dev community
Full text isn't available here.
Read at source →