RCreddit.com·
Not on the current live radar
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)
Qwen3.8-Flash-Next achieved 38 tokens/second decode and an 18-minute prefill time at 1M context on Strix Halo, using halogen 0.12.0. These results were obtained on a Ryzen AI Max+ 395 with 128 GB, with the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576) for cold requests. A follow-up turn over the prompt cache at 1M reached its first token in approximately 0.55 seconds, while the 32k row maintained its standard ten-prompt served mean.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 20, 2026, 04:01 UTC
- Ingested
- Sep 20, 2026, 04:01
- Source type
- Dev community
Full text isn't available here.
Read at source →