Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open
Qwen3.8-Flash-Next (125B MoE, 6B active) has been successfully run on a single AMD Strix Halo mini PC (Ryzen AI Max+ 395, 128 GB), achieving 44-59 tok/s with speculative decoding and approximately 1,400 tok/s prefill. The developers released 95 GB EXL3 weights and a new version of their open-source inference engine, Kyojin, built on ExLlamaV3. While Halogen 0.16.2 shows faster token generation in some scenarios, Kyojin demonstrates higher fidelity to the original model with 94.1% top-1 agreement.
This report details the first public release of the Kyojin inference engine, which achieves higher fidelity to the original model compared to Halogen 0.16.2.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 5, 2026, 18:00 UTC
- Ingested
- Oct 5, 2026, 18:00
- Source type
- Dev community
Full text isn't available here.
Read at source →