RCreddit.com·
暂不在当前实时榜单
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)
Qwen3.8-Flash-Next achieved 38 tokens/second decode and an 18-minute prefill time at 1M context on Strix Halo, using halogen 0.12.0. These results were obtained on a Ryzen AI Max+ 395 with 128 GB, with the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576) for cold requests. A follow-up turn over the prompt cache at 1M reached its first token in approximately 0.55 seconds, while the 32k row maintained its standard ten-prompt served mean.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月20日 04:01 UTC
- 收录
- 2026年9月20日 04:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →