RCreddit.com·
暂不在当前实时榜单
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
Official llama.cpp is not optimized for Strix Halo (gfx1151), achieving less than 50% of its hardware's theoretical performance. For optimal throughput, users should consider peonist-ai/halogen-flash-server, which is specifically optimized for Strix Halo and the Qwen 3.8 Flash Next (Q38FN) model. This setup, also known as "Ninfer for Strix Halo," can reach approximately 50 tokens/second for decoding and 1200 tokens/second for prefill, utilizing about 90% of the hardware's theoretical capacity.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月8日 08:00 UTC
- 收录
- 2026年9月8日 08:00
- 来源类型
- 开发者社区
讨论趋势
→ 平稳
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
正文
本站未收录正文。
前往源站阅读 →