Skip to content
RCreddit.com·

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

AI summary

Inco Splash, an open-source inference engine, achieves 144 tok/s with Qwen3.8-27B on an M5 Max MacBook Pro. It offers up to 3x the decode speed of Ollama and 2x oMLX, and nearly 4x when agents fan out. Requirements include an M3 or newer Mac with macOS 26.4+ and 36 GB RAM. Users can install it via brew install incoai/tap/splash and serve models like incoai/Qwen3.8-27B-Splash, or use it within LM Studio Bionic for local agent work.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 19, 2026, 02:11 UTC

IngestedOffset at this time: UTC+0Sep 19, 2026, 06:00 UTC

Published
Sep 19, 2026, 02:11
Ingested
Sep 19, 2026, 06:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon.

Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes

Settings > Runtime, download Splash, then download the model. The same engine, inside the app, for local agent work on your Mac.

Source·reddit.com