跳到正文
Rreddit·
Archived topic · 归档话题,来源已停止追踪

Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made long-context inference usable

AI 摘要

A developer successfully ran Qwen3.5-122B on a Mac Studio by fixing three bugs in their qMLX serving stack. These bugs, including prompt instability, interrupt path issues, and checkpoint poisoning, caused significant delays in long-context inference. After the fixes, prefill times dropped from minutes to sub-seconds. The developer open-sourced their qMLX fork, optimized specifically for Qwen, and a benchmark script to help others with similar hybrid attention caching issues.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年7月13日 03:00 UTC

收录
2026年7月13日 03:00
来源类型
未分类