跳到正文
RCreddit.com·

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

AI 摘要

Inco Splash, an open-source inference engine, achieves 144 tok/s with Qwen3.8-27B on an M5 Max MacBook Pro. It offers up to 3x the decode speed of Ollama and 2x oMLX, and nearly 4x when agents fan out. Requirements include an M3 or newer Mac with macOS 26.4+ and 36 GB RAM. Users can install it via brew install incoai/tap/splash and serve models like incoai/Qwen3.8-27B-Splash, or use it within LM Studio Bionic for local agent work.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月19日 02:11 UTC

收录当时偏移:UTC+02026年9月19日 06:00 UTC

发布
2026年9月19日 02:11
收录
2026年9月19日 06:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

暂无对比
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon.

Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes

Settings > Runtime, download Splash, then download the model. The same engine, inside the app, for local agent work on your Mac.

来源·reddit.com