Running decision model locally on an RTX 4090 to find out which one is the fastest
A user benchmarked four open decision models—Laya, d1, Clef-Flash, and Lev—on an RTX 4090 GPU to compare their speed and accuracy in identifying centipede names from Wikipedia articles. Laya, using Laya-BF16.gguf and llama.cpp, was the fastest with a 3.9 ms per word latency. Lev, despite being 13x slower at 51.0 ms and running on its own PyTorch server, achieved the highest accuracy in catching centipede names. Laya was ultimately chosen as the overall fastest and most adaptable model.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 8, 2026, 17:04 UTC
IngestedOffset at this time: UTC+0Oct 9, 2026, 02:00 UTC
- Published
- Oct 8, 2026, 17:04
- Ingested
- Oct 9, 2026, 02:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back
{"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}}
model weights engine per word (p50) words in 32s accuracy centipede names caught wrong picks Laya Laya-BF16.gguf llama.cpp b11495 3.9 ms 7,980 97.4% 70% 98 d1 3B d1-3B-AD-Q4_K_M.gguf llama.cpp b11495 6.0 ms 5,306 96.5% 51% 51 Clef-Flash 9B Clef-Flash-Q8_0.gguf llama.cpp b11495 24.4 ms 1,292 97.2% 36% 2 Lev 4B interfaze-ai/lev, bf16 lev serve (PyTorch) 51.0 ms 626 98.9% 83% 4 laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick
- engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default
- d1: our own AD-Q4_K_M quant ( atomic.chat ), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110
- Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up)