RCreddit.com·
暂不在当前实时榜单
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
A developer has achieved running the 95.5 GiB Qwen3.8-Flash-Next model on a 64GB Mac at 41–52 tok/s, which is 1.76x faster than llama.cpp. This was accomplished through a Slipstream release, featuring 130k context scaling and a Swift variant. The new implementation shows improved decode speeds compared to a previous custom expert-streaming fork of llama.cpp, which capped at 23–27 tok/s and slowed with context growth. The Swift-Flash-Next V3 variant also demonstrated a +2.8% overall accuracy improvement.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月2日 03:00 UTC
- 收录
- 2026年10月2日 03:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →