跳到正文
RCreddit.com·

Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact)

AI 摘要

A new engine achieves 20.14–21.13 tokens/s with Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS on an RTX 3060 12GB and 16GB DDR4 RAM. This performance is significantly faster than the stock llama.cpp (mmap) which runs at 1.41–2.12 tokens/s. The engine utilizes "--moe-direct-io + prefetch" for sequential streaming, maintaining bit-exactness to stock, and has minimal major page faults.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 18:41 UTC

收录当时偏移:UTC+02026年10月10日 00:00 UTC

发布
2026年10月9日 18:41
收录
2026年10月10日 00:00
来源类型
开发者社区
档位
社区
信源状态
同步延迟

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

暂无对比
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!)

About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading.

Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload

The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence.

Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops.

I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners.

And if you've tried running a 68GB MoE on a 16GB RAM machine with stock llama.cpp, you probably know how painful it gets.

I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around.

Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using mlock. If you've only got 16GB RAM, you're obviously not doing that. Depending on the setup, you either run into OOM issues or end up with terrible performance.

So I went back to my original idea and started implementing it properly as an optional feature inside llama.cpp:

The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock.

Engine / mode Decode speed Major page faults per token SSD I/O Output Stock llama.cpp (mmap) 1.41–2.12 tok/s 1,140–1,565 208–312 MB/token Coherent Our engine (blocking, demand-only) 0.73 tok/s 0 ~206 MB/token Bit-exact Our engine (--moe-direct-io + prefetch) 20.14–21.13 tok/s ~0 (+1 across 32 tokens!) Sequential streaming Bit-exact to stock The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency.

Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock llama.cpp on the same machine.

And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock.

That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project.

Right now, the slots start empty (-1), so the first request on a new topic can start around 3.5–4.5 tok/s before ramping up to 20+ tok/s as the hot working set settles.

We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch.

Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle.

We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains.

Right now, confidence-gated speculation (min-p 0.8) or suffix prompt lookup seems more promising for this kind of setup.

Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me.

Happy to answer questions or get into the io_uring and slot-remapping implementation details if anyone's interested.

来源·reddit.com