Skip to content
HNHacker News·
Not on the current live radar

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

AI summary

The Qwen 3.8 Flash Next (125B) model can run on consumer hardware like the RTX 4090 at speeds up to 100 T/s. Performance metrics for different quantization levels (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, Coder) on NVIDIA and AMD GPUs detail tokens per second for answer generation and prompt reading. For instance, an RTX 3090 (24 GB) is expected to achieve 100-140 tokens per second. This is enabled by the open-source Strata engine. A detailed performance table is available in DETAILS.md, and GPUs with more VRAM generally offer faster performance.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 14:00 UTC

Ingested
Oct 4, 2026, 14:00
Source type
Dev community
Breakout verdict
Basis
Running about 6.6× the median of this source's recent listed items
Metric comparison
227 vs median 34.5 (20 baseline samples)
Detected
10/04, 14:00

Full text isn't available here.

Read at source →
Source·Hacker News·github.com