Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090
Swift 1.5 and HyperQwen have been shown to reduce task completion time by 37% when running on an RTX 3090, achieving over 100 transactions per second (tps) with a 150k context. Tests across various benchmarks, including GSM8K, IFBench, LiveCodeBench, and custom tool-call/JSON evaluations, demonstrate strong performance. For instance, Swift 1.5 INT4 heads achieved 98.0% on GSM8K and 91% on LiveCodeBench, indicating efficient operation and high accuracy.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 28, 2026, 20:52 UTC
IngestedOffset at this time: UTC+0Sep 29, 2026, 20:00 UTC
- Published
- Sep 28, 2026, 20:52
- Ingested
- Sep 29, 2026, 20:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Hi everyone :)
The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.
To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.
Performance
Model Average time/task ↓ Average output tokens/task ↓ Decode tok/s ↑ Qwen - HyperQwen fast quant 108.1 s 8,985 112.1 Swift 1.0 + HyperQwen 66.2 s 5,245 105.9 Swift 1.5 + HyperQwen, INT8 heads 72.2 s 5,751 104.0 Swift 1.5 + HyperQwen INT4 heads 68.2 s 5,669 107.2 All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.
Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.
Quality
There are some minor quality and performance tradeoffs between the models:
Test Qwen HyperQwen fast Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads GSM8K, 200 questions 97.5% 98.0% 98.0% 97.5% IFBench, 300 prompts, strict 74.0% 73.3% 73.7% 72.3% LiveCodeBench, (100-problem subset) 90% 89% 89% 91% Custom tool-call/JSON eval 29/30 28/30 30/30 30/30 English/Python perplexity ↓ 6.551 6.605 6.643 6.679 Applied changes to Swift models to adapt for HyperQwen:
Changes to Swift models:
- Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.
- Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.
- Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.
Setup
If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.
All three models can be found here: https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks
Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!