Jeeves. Reasoning improves Jev-like decision models
Jeeves is a reasoning Jev-style classifier that utilizes a diffusion drafter and is trained with SFT and CISPO. It demonstrates significant improvements across various benchmarks compared to Kev-9B Jev models. For instance, Jeeves achieved an "overall Test" score of 0.889, a "Transfer overall" score of 0.800, and a "JevBench overall" score of 0.935. It also showed strong performance in specific tasks like QNLI (0.925), SciQ (0.991), and MMLU (0.900), indicating enhanced decision-making capabilities.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 29, 2026, 11:13 UTC
IngestedOffset at this time: UTC+0Sep 29, 2026, 12:00 UTC
- Published
- Sep 29, 2026, 11:13
- Ingested
- Sep 29, 2026, 12:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
- Basis
- Running about 2.2× the median of this source's recent listed items
- Triggering item
- Jeeves. Reasoning improves Jev-like decision models
- Metric comparison
- 185 vs median 84.5 (20 baseline samples)
- Detected
- 09/29, 19:00
A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.
Acknowledgements
Inspired by Kev .
Highlights
- A 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, with a block-4 diffusion drafter and the full training code and train/dev/test data.
- Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench's public tiers (0.935 vs 0.866 for Jev).
- Supports yes/no (noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API.
- About 0.3 s per request without thinking and a 3.3 s median with it on one H100. Can be sped up by truncating chain length.
- Runs on CUDA (Hopper for the FP8 kernel).
Problem
Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides.
This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public).
Results
Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.
bench Kev-9B Jev Jeeves
Test overall (out-of-domain and held-out, item-weighted) 0.822 0.857 0.889
Transfer overall (MMLU-Pro and buried state) 0.579 0.800 0.746
JevBench overall (231 public items) 0.715* 0.866 0.935
QNLI 0.925 0.925 0.913
SciQ 0.963 0.988 0.991
TweetEval offensive 0.775 0.813 0.813
PAWS 0.763 0.788 0.875
MMLU 0.738 0.900 0.793
Emotion 0.600 0.588 0.647