Skip to content
RCreddit.com·

Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

AI summary

Basalt, a Blackwell inference engine, achieves 665 tokens/s structured and 354 tokens/s prose on a 5090 + 5060 Ti setup, with a 400 W power consumption. It supports real concurrency for up to 8 users, offering 623 tokens/s total across 8 streams. Basalt also features a custom vision encoder, up to 3x faster on GPU and 4x on CPU than llama.cpp's, and provides an OpenAI + Anthropic compatible server UI for throughput and hardware statistics.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Oct 10, 2026, 00:58 UTC

IngestedOffset at this time: UTC+0Oct 10, 2026, 07:00 UTC

Published
Oct 10, 2026, 00:58
Ingested
Oct 10, 2026, 07:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:

- **SPEED:** 665 struct, 354 prose, 7,317 prefill at 64k, IQ3_XXS, 400 W, speed is the highlight
 - Single 5090 works too (no second card): 585 struct, 316 prose on IQ3_XXS, ~12% slower than with the 5060 Ti, prefill unchanged
 - Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare)
 - Fine-tuned MTP for lower quants, for increased throughput
 - Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU
 - A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness)
 - OpenAI + Anthropic compatible
 - Linux support (No Windows or Mac)

Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3_XXS and IQ3_S and UD-Q4_K_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.

+ Is it open source?
 - Yes, fully open source, MIT license:  https://github.com/jesdga95/basalt , fork it, improve it, share it with friends and foes.
+ Where are the weights?
 - Here:  https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt  pick your poison, fast and dumb or smart and slow. IQ3_S is a good middle ground (~89% top 1 agreement, 300 tok/s prose on my setup).
+ Why didn't you just contribute upstream to Strata?
 - This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.
+ Will you support my AMD 98123X?
 - Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.
+ This is vibe coded slop
 - Yes, but it's **fast** slop. Nobody is hand-writing cuda kernels anymore.
Source·reddit.com