I benchmarked Jev against gpt-5.6-luna!
A recent benchmark compared Jev, a non-generative model from TypeSafe, against gpt-5.6-luna. Jev, which answers typed questions with probabilities, demonstrated superior reasoning capabilities across several benchmarks. It scored 0.77 on LogiQA, 0.89 on WinoGrande, 0.97 on ARC-Challenge, and 0.94 on MMLU, outperforming gpt-5.6-luna in all instances. Furthermore, Jev achieved 0.75 on 120 new math word problems generated with Fable 5.1, comparable to its 0.72 on GSM8K, while Luna scored only 0.17 on the same problems with reasoning off.
This report uniquely highlights Jev's reasoning capabilities, showing it outperforms gpt-5.6-luna on standard benchmarks and new math problems, unlike other models that might only excel on memorized datasets.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 19, 2026, 11:00 UTC
- Ingested
- Sep 19, 2026, 11:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →