I benchmarked Jev aginst gpt-5.6-luna!
A recent benchmark compared TypeSafe's Jev model against gpt-5.6-luna, highlighting Jev's unique design as a non-text-generating model that provides probabilistic answers to typed questions. Jev demonstrated strong reasoning capabilities, outperforming gpt-5.6-luna on benchmarks like LogiQA (0.77 vs 0.59), WinoGrande (0.89 vs 0.66), ARC-Challenge (0.97 vs 0.87), and MMLU (0.94 vs 0.87). Furthermore, Jev achieved a 0.75 score on 120 new math word problems generated with Fable 5.1, comparable to its 0.72 on GSM8K, while gpt-5.6-luna scored 0.17 on the same problems with reasoning off.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月19日 08:25 UTC
收录当时偏移:UTC+02026年9月19日 14:00 UTC
- 发布
- 2026年9月19日 08:25
- 收录
- 2026年9月19日 14:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
I got access to TypeSafe's Jev a few days ago. It's an odd kind of model that doesn't generate text at all. You send it some content plus typed questions (yes/no, pick one of these options, rate this on a scale) and it gives you back probabilities.
Setup: 49 tasks, about 8,200 items, all from public labelled datasets (SST-2, MMLU, ARC, MS MARCO, XNLI, SQuAD, Banking77 and so on), plus some synthetic tests where code computes the right answer. Both models got the same question wording. The baseline was gpt-5.6-luna with reasoning off, and again with reasoning on low.
~105ms median server time vs 700-800ms about $0.04 per 1,000 items vs $0.16-0.19 calibration error roughly half the baseline's, so the probabilities actually mean something
The part that surprised me was reasoning. LogiQA 0.77 vs 0.59, WinoGrande 0.89 vs 0.66, ARC-Challenge 0.97 vs 0.87, MMLU 0.94 vs 0.87. I assumed it had memorised the benchmarks, so I generated 120 new math word problems with a script using Fable 5.1. It got 0.75 on those, about the same as its 0.72 on GSM8K. Luna with reasoning off got 0.17 on the same problems.
Reranking was also a clear win (NFCorpus nDCG@10 0.73 vs 0.63), and Luna needed 1-3 seconds per query there.
Where it lost: counting things in a list (0.87 vs 0.99 once Luna could reason) picking 1 of 77 intents (Banking77, 0.81 vs 0.87) asking a question and then its negation did not give probabilities that add up to 1, off by about 0.3 on average
in a separate private test with ~100 long documents in one request, it gave high relevance scores to a few completely unrelated ones. None of the clean benchmarks show that.
because I know someone will ask. These are public benchmarks, so contamination is possible for any model, and the new math set is my only control for that. The baseline was deliberately held to no/low reasoning. 200 items per task means anything under about 0.05 is noise. Only 7 of the leads are clearly outside the error bars. And it can't write text, call tools or explain itself, so it doesn't replace your LLM. It replaces the small classifier, reranker and "is this relevant" calls around it.
Repo has the task builders, runner, scorer, my raw per-item results and the full table. Python stdlib only. You can plug in any OpenAI-compatible model as another baseline with one argument, along with the full results: