Skip to content
HNHacker News·

Jeeves. Reasoning improves Jev-like decision models

AI summary

Jeeves is a reasoning Jev-style classifier that utilizes a diffusion drafter and is trained with SFT and CISPO. It demonstrates significant improvements across various benchmarks compared to Kev-9B Jev models. For instance, Jeeves achieved an "overall Test" score of 0.889, a "Transfer overall" score of 0.800, and a "JevBench overall" score of 0.935. It also showed strong performance in specific tasks like QNLI (0.925), SciQ (0.991), and MMLU (0.900), indicating enhanced decision-making capabilities.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 29, 2026, 11:13 UTC

IngestedOffset at this time: UTC+0Sep 29, 2026, 12:00 UTC

Published
Sep 29, 2026, 11:13
Ingested
Sep 29, 2026, 12:00
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Breakout verdict
Basis
Running about 2.2× the median of this source's recent listed items
Metric comparison
185 vs median 84.5 (20 baseline samples)
Detected
09/29, 19:00

A reasoning Jev-style classifier with a diffusion drafter, trained with SFT and CISPO.

Acknowledgements

Inspired by Kev .

Highlights

- A 9B Jev-like model (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, with a block-4 diffusion drafter and the full training code and train/dev/test data.

- Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench's public tiers (0.935 vs 0.866 for Jev).

- Supports yes/no (noul), multiple-choice (choice), and rating (score) questions in the same request, through a Jev-compatible API.

- About 0.3 s per request without thinking and a 3.3 s median with it on one H100. Can be sped up by truncating chain length.

- Runs on CUDA (Hopper for the FP8 kernel).

Problem

Jev-like models give calibrated decision probabilities, but at low accuracy. A lot of pipelines therefore rely on a reasoning model as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to reason before it decides.

This results in better performance on out of domain tasks, and outperforms Jev in JevBench hard (public).

Results

Accuracy with thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.

bench Kev-9B Jev Jeeves

Test overall (out-of-domain and held-out, item-weighted) 0.822 0.857 0.889

Transfer overall (MMLU-Pro and buried state) 0.579 0.800 0.746

JevBench overall (231 public items) 0.715* 0.866 0.935

QNLI 0.925 0.925 0.913

SciQ 0.963 0.988 0.991

TweetEval offensive 0.775 0.813 0.813

PAWS 0.763 0.788 0.875

MMLU 0.738 0.900 0.793

Emotion 0.600 0.588 0.647

Source·Hacker News·github.com