跳到正文
RCreddit.com·

Astra is a quiet force. Artificial Analysis has updated their benchmark twice in 4 days to reflect its real strength

AI 摘要

Astra 正在展现出强大的实力,促使 Artificial Analysis 在四天内两次更新其基准测试,以准确反映其真实性能。类似的情况也曾发生在 Sol 身上,当时 Arena 不得不重组其编码基准测试,并增加了全栈基准测试,同时更新了其网络开发基准测试,以在 Astra 发布后更好地反映实际性能。这表明 OpenAI 可能并未完全针对基准测试进行优化。

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月8日 16:36 UTC

收录当时偏移:UTC+02026年9月9日 15:00 UTC

发布
2026年9月8日 16:36
收录
2026年9月9日 15:00
来源类型
开发者社区
档位
社区
信源状态
同步延迟

档位是按信源手工设定的编辑判断,不是逐条打分。

正文

Something similar happened with Sol and Arena had to restructure their coding benchmark and added a full-stack bench, plus updated the webdev bench to reflect real world performance after Astra released. Looks OpenAI is only lab not benchmaxxing

来源·reddit.com