How far have ~30B open models actually come? Qwen3.8 vs Qwen3.6 vs Gemma 4
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
With Qwen3.8-27B out, I compared it with Qwen3.6-27B and Gemma 4 31B.
They’re unusually good models to compare because they’re all around the same size:
Qwen3.8: 27B, 262K context Qwen3.6: 27B, 262K context Gemma 4: 31B, 256K context
What’s interesting is where the gains are going.
Qwen3.8 pulls ahead particularly on coding and agentic benchmarks, while Gemma 4 is still very competitive on general reasoning. Comparing 3.8 directly with 3.6 also shows how much performance has moved in a single generation without increasing the parameter count.
And these aren’t datacenter-sized models. Quantized, this is roughly the class of AI you can run on a high-end consumer GPU.
The gap between “local model” and genuinely useful AI is getting pretty small.
Full benchmark + hardware comparisons:
https://canitrun.dev/models/qwen3.8-27b/
https://canitrun.dev/models/compare/qwen3.8-27b-vs-qwen3.6-27b/
https://canitrun.dev/models/compare/qwen3.8-27b-vs-gemma-4-31b/