Skip to content
RCreddit.com·
Not on the current live radar

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

AI summary

A developer has achieved running the 95.5 GiB Qwen3.8-Flash-Next model on a 64GB Mac at 41–52 tok/s, which is 1.76x faster than llama.cpp. This was accomplished through a Slipstream release, featuring 130k context scaling and a Swift variant. The new implementation shows improved decode speeds compared to a previous custom expert-streaming fork of llama.cpp, which capped at 23–27 tok/s and slowed with context growth. The Swift-Flash-Next V3 variant also demonstrated a +2.8% overall accuracy improvement.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 2, 2026, 03:00 UTC

Ingested
Oct 2, 2026, 03:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com