RCreddit.com·
Not on the current live radar
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
A developer has achieved running the 95.5 GiB Qwen3.8-Flash-Next model on a 64GB Mac at 41–52 tok/s, which is 1.76x faster than llama.cpp. This was accomplished through a Slipstream release, featuring 130k context scaling and a Swift variant. The new implementation shows improved decode speeds compared to a previous custom expert-streaming fork of llama.cpp, which capped at 23–27 tok/s and slowed with context growth. The Swift-Flash-Next V3 variant also demonstrated a +2.8% overall accuracy improvement.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 2, 2026, 03:00 UTC
- Ingested
- Oct 2, 2026, 03:00
- Source type
- Dev community
Full text isn't available here.
Read at source →