Skip to content
RCreddit.com·

Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090

AI summary

Swift 1.5 and HyperQwen have been shown to reduce task completion time by 37% when running on an RTX 3090, achieving over 100 transactions per second (tps) with a 150k context. Tests across various benchmarks, including GSM8K, IFBench, LiveCodeBench, and custom tool-call/JSON evaluations, demonstrate strong performance. For instance, Swift 1.5 INT4 heads achieved 98.0% on GSM8K and 91% on LiveCodeBench, indicating efficient operation and high accuracy.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 28, 2026, 20:52 UTC

IngestedOffset at this time: UTC+0Sep 29, 2026, 20:00 UTC

Published
Sep 28, 2026, 20:52
Ingested
Sep 29, 2026, 20:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Hi everyone :)

The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.

To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.

Performance

Model Average time/task ↓ Average output tokens/task ↓ Decode tok/s ↑ Qwen - HyperQwen fast quant 108.1 s 8,985 112.1 Swift 1.0 + HyperQwen 66.2 s 5,245 105.9 Swift 1.5 + HyperQwen, INT8 heads 72.2 s 5,751 104.0 Swift 1.5 + HyperQwen INT4 heads 68.2 s 5,669 107.2 All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.

Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.

Quality

There are some minor quality and performance tradeoffs between the models:

Test Qwen HyperQwen fast Swift 1.0 Swift 1.5 INT8 heads Swift 1.5 INT4 heads GSM8K, 200 questions 97.5% 98.0% 98.0% 97.5% IFBench, 300 prompts, strict 74.0% 73.3% 73.7% 72.3% LiveCodeBench, (100-problem subset) 90% 89% 89% 91% Custom tool-call/JSON eval 29/30 28/30 30/30 30/30 English/Python perplexity ↓ 6.551 6.605 6.643 6.679 Applied changes to Swift models to adapt for HyperQwen:

Changes to Swift models:

- Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.

- Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.

- Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.

Setup

If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.

All three models can be found here: https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks

Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!

Source·reddit.com