Skip to content
RCreddit.com·

Inference Engines will become a series of one-offs

AI summary

Inference engines are predicted to become a series of one-off implementations, such as ninfer, dwarfstar, Splash, llamAmpere, and gufo. This is because tasks like optimizing "tok/s go up" for a specific hardware/model combination are fully specified, making them ideal for 100% autonomous AI implementations with trivial correctness tests, eliminating human bottlenecks. A separate growth dimension involves projects like Freetoken and BeeLlama, which frontrun general inference engine features, though these are less model- or hardware-specific.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 29, 2026, 17:21 UTC

IngestedOffset at this time: UTC+0Sep 30, 2026, 04:00 UTC

Published
Sep 29, 2026, 17:21
Ingested
Sep 30, 2026, 04:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.

We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.

For better or worse, the list will continue to grow

They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.

But new ones will take their place...

THESIS

One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower

A few axioms you probably accept:

- the more general a codebase, the harder it is to cleanly fit new features in over time. This reduces the pace of innovation. The smaller, the faster

- AI coding is getting better and cheaper. Thus, the barrier to creating an inference engine is dropping

- many coding tasks are difficult to completely give to AI (or a human) because they are not fully specified. But "Make tok/s go up" in an inference engine fork for one hardware/model combo is fully specified, and is therefore a great candidate for 100% autonomous implementations to be perfectly fine in terms of quality (as long as correctness tests are included, which is trivial). No human bottleneck.

String these axioms together, and I arrive at

- general engines, like llama.cpp, will perpetually lag behind these one-offs in development speed, and therefore token speed

- none of the one-offs will be able to maintain generality and dev speed over time

- one-off inference engines for specific hardware/model combinations will continue to proliferate, and be loved

IMPLICATIONS

- This thesis brings up an interesting question: What elements of inference engines WILL remain in common?

Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.

Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?

One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.

There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.

- Maybe we'll see more 'half-general' inference engines that just target one hardware platform. So still general on the dimension of models, but not on hardware. Splash could be an example.

- Nobody wants to continuously scan github/reddit/x for the best inference engine for their model/rig. Some will just have their agent custom make one. But I think a larger number will not do that. So, hardware-specific communities will form. Think r/appleM2Max32gbLLM and r/4090And64gbRamLLM , etc (however that actually ends up organizing. exaggerating a bit on the names.)

ALTERNATIVE FUTURES

Source·reddit.com