跳到正文
RCreddit.com·
暂不在当前实时榜单

Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

AI 摘要

A developer claims a new record for memory-constrained inference of Qwen3.8-Flash-Next on Apple Silicon, achieving 8-22 tg/s on a 32GB M4 MacBook Air with only 21GB of allocations. This is made possible by Cherenkov, an inference engine for Apple Silicon. Cherenkov uses predictive expert streaming and optional mixed-precision execution, keeping a bounded working set of experts in unified memory and predicting future expert needs to initiate SSD reads, with a fallback to smaller Q3/Q2 quantizations if full loading isn't possible.

为什么是这条

This report details a new record for memory-constrained inference on Apple Silicon, unlike previous benchmarks that often focus on larger, dedicated AI hardware.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月11日 06:00 UTC

收录
2026年9月11日 06:00
来源类型
开发者社区
正文

本站未收录正文。

前往源站阅读 →
来源·reddit.com