Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
A developer claims a new record for memory-constrained inference of Qwen3.8-Flash-Next on Apple Silicon, achieving 8-22 tg/s on a 32GB M4 MacBook Air with only 21GB of allocations. This is made possible by Cherenkov, an inference engine for Apple Silicon. Cherenkov uses predictive expert streaming and optional mixed-precision execution, keeping a bounded working set of experts in unified memory and predicting future expert needs to initiate SSD reads, with a fallback to smaller Q3/Q2 quantizations if full loading isn't possible.
Why this oneThis report details a new record for memory-constrained inference on Apple Silicon, unlike previous benchmarks that often focus on larger, dedicated AI hardware.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 11, 2026, 06:00 UTC
- Ingested
- Sep 11, 2026, 06:00
- Source type
- Dev community
Full text isn't available here.
Read at source →