Skip to content
Rreddit·
Archived topic · source no longer tracked

Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable

AI summary

A developer successfully ran Qwen3.5-122B on a Mac Studio by fixing three bugs in their qMLX serving stack. These bugs, including prompt instability, interrupt path issues, and checkpoint poisoning, caused significant delays in long-context inference. After the fixes, prefill times dropped from minutes to sub-seconds. The developer open-sourced their qMLX fork, optimized specifically for Qwen, and a benchmark script to help others with similar hybrid attention caching issues.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Jul 13, 2026, 03:00 UTC

Ingested
Jul 13, 2026, 03:00
Source type
Unclassified