Are you running Qwen 3.8 27b or Qwen Flash Next?
A user with M3 Max 96GB hardware is comparing Qwen 3.8 27b and Qwen Flash Next, noting that Qwen 27b has faster prefill despite both feeling largely identical. They are curious if anyone is working on improving prefill performance with MLX. The user also asks if there's a harness for models without reasoning, inspired by Jetbrains' use of Qwen 3.6 with reasoning off, and considers orchestrating models with and without reasoning.
- Published
- 09/07, 15:25 UTC+0
- Ingested
- 09/07, 21:00 UTC+0
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?
Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/
Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.