UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh
UkisAI has post-trained Qwen 3.8 27B, creating Swift-Qwen3.8-27B, which achieves a 1.95x speed-up and 58% fewer thinking tokens while maintaining xhigh accuracy. This was accomplished by penalizing tokens linked to overthinking and using On-Policy Distillation. Benchmarks show significant reductions in thinking tokens across various tasks, with minimal accuracy changes. UkisAI is also working on Swift 3.8 Flash Next, aiming for further reductions in thinking token usage.
This report details a method to reduce thinking tokens by 58% and increase speed by 1.95x in Qwen 3.8 27B, unlike other optimizations that often sacrifice accuracy.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 17:01 UTC
- Ingested
- Sep 14, 2026, 17:01
- Source type
- Dev community
Full text isn't available here.
Read at source →