Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

AI summary

Qwen3.8-Flash-Next achieved 38 tokens/second decode and an 18-minute prefill time at 1M context on Strix Halo, using halogen 0.12.0. These results were obtained on a Ryzen AI Max+ 395 with 128 GB, with the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576) for cold requests. A follow-up turn over the prompt cache at 1M reached its first token in approximately 0.55 seconds, while the 32k row maintained its standard ten-prompt served mean.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 20, 2026, 04:01 UTC

Ingested
Sep 20, 2026, 04:01
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com