I benchmarked Jev against gpt-5.6-luna!
A recent benchmark compared Jev, a non-generative model from TypeSafe, against gpt-5.6-luna. Jev, which answers typed questions with probabilities, demonstrated superior reasoning capabilities across several benchmarks. It scored 0.77 on LogiQA, 0.89 on WinoGrande, 0.97 on ARC-Challenge, and 0.94 on MMLU, outperforming gpt-5.6-luna in all instances. Furthermore, Jev achieved 0.75 on 120 new math word problems generated with Fable 5.1, comparable to its 0.72 on GSM8K, while Luna scored only 0.17 on the same problems with reasoning off.
This report uniquely highlights Jev's reasoning capabilities, showing it outperforms gpt-5.6-luna on standard benchmarks and new math problems, unlike other models that might only excel on memorized datasets.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月19日 11:00 UTC
- 收录
- 2026年9月19日 11:00
- 来源类型
- 开发者社区
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →