跳到正文
RCreddit.com·
暂不在当前实时榜单

I ran GPT-6 Luna Max on MathArena's harness

AI 摘要

A developer tested GPT-6 Luna Max on MathArena's harness, noting its strong performance on Riemann Bench and considering it an underrated model. The model correctly answered 16 out of 19 finite or discrete questions, 14 out of 20 analysis and probability questions, and 9 out of 18 geometry, algebra, and topology questions. The developer abandoned testing another model, xhigh, after it performed worse than Luna Max.

为什么是这条

This report uniquely offers specific performance metrics for GPT-6 Luna Max across different mathematical domains, unlike general claims of strong performance.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月27日 05:00 UTC

收录
2026年9月27日 05:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com