LessThink-Qwen3-4B: the same model, with far less thinking [P]
A developer has post-trained the Qwen3-4B model, creating "LessThink-Qwen3-4B" which significantly reduces token usage for reasoning by 44% while maintaining its original knowledge and answer style. This entire process was achieved using a single GPU. The developer invites interested individuals to explore this new model further on their website.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 30, 2026, 07:19 UTC
IngestedOffset at this time: UTC+0Sep 30, 2026, 14:00 UTC
- Published
- Sep 30, 2026, 07:19
- Ingested
- Sep 30, 2026, 14:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
I post-trained Qwen3-4B to spend 44% fewer tokens on reasoning, keeping its knowledge and answer style. The whole pipeline ran on one GPU.
folks, you can check it out on: https://5ivatej.com/lessthink/