Skip to content
RCreddit.com·
Not on the current live radar

Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

AI summary

A developer claims a new record for memory-constrained inference of Qwen3.8-Flash-Next on Apple Silicon, achieving 8-22 tg/s on a 32GB M4 MacBook Air with only 21GB of allocations. This is made possible by Cherenkov, an inference engine for Apple Silicon. Cherenkov uses predictive expert streaming and optional mixed-precision execution, keeping a bounded working set of experts in unified memory and predicting future expert needs to initiate SSD reads, with a fallback to smaller Q3/Q2 quantizations if full loading isn't possible.

Why this one

This report details a new record for memory-constrained inference on Apple Silicon, unlike previous benchmarks that often focus on larger, dedicated AI hardware.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 11, 2026, 06:00 UTC

Ingested
Sep 11, 2026, 06:00
Source type
Dev community
Article

Full text isn't available here.

Read at source →
Source·reddit.com