Skip to content
RCreddit.com·
Not on the current live radar

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

AI summary

Researchers developed Kyojin, an engine built on ExLlamaV3 for AMD Strix Halo (gfx1151, ROCm), enabling two 300B-class Mixture-of-Experts (MoE) models to run on a single 128 GB mini PC. The GLM-5.3-Flash model achieves approximately 580 tok/s prefill, while the MiMo-V2.6-Flash model reaches up to 44 tok/s decode. These models utilize EXL3 weights, with MiMo being a custom quantization and GLM a mix of public turboderp tensors and custom tuning.

Why this one

This is the first time two 300B-class MoE models have been shown running on a single 128 GB mini PC, unlike previous demonstrations on more powerful hardware.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 00:00 UTC

Ingested
Oct 4, 2026, 00:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com