Qwen 3.8 Next Flash is really really REALLY verbose..
A user migrating from Qwen 3.6 27b to Qwen 3.8 Next Flash reports that the new model is excessively verbose, taking up to 13 minutes for single-turn coding requests. Despite a decent output quality for straightforward tasks, the model struggles with decision-making, producing jargon-filled responses. The user is hesitant to lower the "thinking level" due to concerns about potential quality degradation, noting that Qwen 3.8 27b's performance is significantly affected by this setting.
- Published
- Sep 7, 2026, 09:31
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Times shown in UTC
More details
Long time user of 3.6 27b, switched over to Next Flash since it's a logical step up even from 3.8 27b. It's soooo verbose, i'm talking 13 minutes of thinking time on single turn coding requests at approximately 150 tokens per second tg and 7000 tokens per second pp. It's honestly kind of painful to use since I look back and it's still thinking, then when I go to check the output it's decent most of the time but if the task requires ANY decision making, it turns into alphabet soup where it's just buzzwords and jargon that nobody actually uses in the SWE space.
The runtime is actually shorter for me if I BYOK it to VSCode, but for pi.dev it's often takes 1 hour!
Before anyone tells me to lower the thinking level, I don't want to do that given the chance it makes the model worse. There's no solid benchmarks for how the model performs at different thinking levels yet, but looking towards 3.8 27b, it seems to affect the quality of the output quite a bit.