跳到正文
RCreddit.com·
暂不在当前实时榜单

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

AI 摘要

Qwen3.8-Flash-Next achieved 38 tokens/second decode and an 18-minute prefill time at 1M context on Strix Halo, using halogen 0.12.0. These results were obtained on a Ryzen AI Max+ 395 with 128 GB, with the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576) for cold requests. A follow-up turn over the prompt cache at 1M reached its first token in approximately 0.55 seconds, while the 32k row maintained its standard ten-prompt served mean.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月20日 04:01 UTC

收录
2026年9月20日 04:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com