Are you running Qwen 3.8 27b or Qwen Flash Next?
一位拥有 M3 Max 96GB 硬件的用户正在比较 Qwen 3.8 27b 和 Qwen Flash Next,发现 Qwen 27b 的预填充速度更快,尽管两者整体感觉相似。他们好奇是否有人正在致力于使用 MLX 改进预填充性能。该用户还询问是否有适用于无推理模型的工具,灵感来源于 Jetbrains 关闭推理功能使用 Qwen 3.6 的做法,并考虑如何协调使用有推理和无推理的模型。
- 发布
- 09/07 15:25 UTC+0
- 收录
- 09/07 21:00 UTC+0
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?
Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/
Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.