ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp
A recent pull request, #30087, titled "ggml-cuda: assign four GDN state columns per warp by SongXiaoXi," has been submitted to the ggml-org/llama.cpp repository. This update aims to significantly speed up Qwen 3.x prompt processing. Performance tests show a +5.5% improvement for pp512 and a +5.3% improvement for pp4096, with a minor +0.1% change for tg128, indicating substantial gains in prompt processing efficiency.
This update specifically targets Qwen 3.x, unlike previous general optimizations, and shows a notable +5.5% speedup in prompt processing for pp512.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 8, 2026, 10:00 UTC
- Ingested
- Oct 8, 2026, 10:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →