RCreddit.com·
Not on the current live radar
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
Official llama.cpp is not optimized for Strix Halo (gfx1151), achieving less than 50% of its hardware's theoretical performance. For optimal throughput, users should consider peonist-ai/halogen-flash-server, which is specifically optimized for Strix Halo and the Qwen 3.8 Flash Next (Q38FN) model. This setup, also known as "Ninfer for Strix Halo," can reach approximately 50 tokens/second for decoding and 1200 tokens/second for prefill, utilizing about 90% of the hardware's theoretical capacity.
Time & source
- Ingested
- 09/08, 08:00 UTC+0
- Source type
- Dev community
Discussion trend
→ Steady
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Article
Full text isn't available here.
Read at source →