Skip to content
RCreddit.com·

Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

AI summary

The 'Strata' inference engine significantly outperforms llama.cpp, achieving 51 tokens/second (t/s) with the ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/IQ3_XXS model at 43k context depth on a laptop with 12GB VRAM and 64GB RAM. This speed is more than double the 23 t/s reached by stock llama.cpp using the same quantization. The developer fixed initial bugs, making it possible to run frontier models locally on modest hardware, demonstrating impressive advancements in local inference capabilities.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 30, 2026, 04:01 UTC

IngestedOffset at this time: UTC+0Sep 30, 2026, 11:00 UTC

Published
Sep 30, 2026, 04:01
Ingested
Sep 30, 2026, 11:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them.

Inference engine (only runs on Nvidia for now; AMD support is experimental): https://github.com/Niko1221/Strata

Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD.

The model I used was https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/tree/main/IQ3_XXS which has a good quality for its size. Above, the model generated the aquarium test. At 43k context depth, it was running at 51 t/s. Stock llama.cpp reached only 23t/s with the same quant.

This quant could only reach 100t/s PP with stock llama.cpp using the same quant. Strata was reading 32k context text at 1500t/s. This is way above my expectation. This quant can load with up to 200k context at 8bit. However, I was only using 131k context.

As you can see it is utilizing 11GB VRAM and 56GB RAM (includes system/OS programs).

This engine is specifically built for one model only and only select ggufs (ISTA-DASLab) work with it. You can use IQ3_S from ISTA-DASLab which they claim recovers full model's performance on coding benchmarks. I tested IQ3_XXS for some time and I would say it is an excellent model.

I never thought 12GB VRAM would be enough to run frontier models from 6 months ago locally on a laptop. What a time to be alive!

Source·reddit.com