Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)
Basalt, a Blackwell inference engine, achieves 665 tokens/s structured and 354 tokens/s prose on a 5090 + 5060 Ti setup, with a 400 W power consumption. It supports real concurrency for up to 8 users, offering 623 tokens/s total across 8 streams. Basalt also features a custom vision encoder, up to 3x faster on GPU and 4x on CPU than llama.cpp's, and provides an OpenAI + Anthropic compatible server UI for throughput and hardware statistics.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月10日 00:58 UTC
收录当时偏移:UTC+02026年10月10日 07:00 UTC
- 发布
- 2026年10月10日 00:58
- 收录
- 2026年10月10日 07:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:
- **SPEED:** 665 struct, 354 prose, 7,317 prefill at 64k, IQ3_XXS, 400 W, speed is the highlight - Single 5090 works too (no second card): 585 struct, 316 prose on IQ3_XXS, ~12% slower than with the 5060 Ti, prefill unchanged - Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare) - Fine-tuned MTP for lower quants, for increased throughput - Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU - A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness) - OpenAI + Anthropic compatible - Linux support (No Windows or Mac)
Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3_XXS and IQ3_S and UD-Q4_K_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.
+ Is it open source? - Yes, fully open source, MIT license: https://github.com/jesdga95/basalt , fork it, improve it, share it with friends and foes.
+ Where are the weights? - Here: https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt pick your poison, fast and dumb or smart and slow. IQ3_S is a good middle ground (~89% top 1 agreement, 300 tok/s prose on my setup).
+ Why didn't you just contribute upstream to Strata? - This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.
+ Will you support my AMD 98123X? - Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.
+ This is vibe coded slop - Yes, but it's **fast** slop. Nobody is hand-writing cuda kernels anymore.