Skip to content
RCreddit.com·
Not on the current live radar

Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

AI summary

Qwen3.8-Flash-Next (125B MoE, 6B active) has been successfully run on a single AMD Strix Halo mini PC (Ryzen AI Max+ 395, 128 GB), achieving 44-59 tok/s with speculative decoding and approximately 1,400 tok/s prefill. The developers released 95 GB EXL3 weights and a new version of their open-source inference engine, Kyojin, built on ExLlamaV3. While Halogen 0.16.2 shows faster token generation in some scenarios, Kyojin demonstrates higher fidelity to the original model with 94.1% top-1 agreement.

Why this one

This report details the first public release of the Kyojin inference engine, which achieves higher fidelity to the original model compared to Halogen 0.16.2.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 5, 2026, 18:00 UTC

Ingested
Oct 5, 2026, 18:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com