Skip to content
RCreddit.com·
Not on the current live radar

For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

AI summary

Official llama.cpp is not optimized for Strix Halo (gfx1151), achieving less than 50% of its hardware's theoretical performance. For optimal throughput, users should consider peonist-ai/halogen-flash-server, which is specifically optimized for Strix Halo and the Qwen 3.8 Flash Next (Q38FN) model. This setup, also known as "Ninfer for Strix Halo," can reach approximately 50 tokens/second for decoding and 1200 tokens/second for prefill, utilizing about 90% of the hardware's theoretical capacity.

Time & source

Ingested
09/08, 08:00 UTC+0
Source type
Dev community

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Article

Full text isn't available here.

Read at source →
Source·reddit.com