RCreddit.com·
Astra is a quiet force. Artificial Analysis has updated their benchmark twice in 4 days to reflect its real strength
Astra 正在展现出强大的实力,促使 Artificial Analysis 在四天内两次更新其基准测试,以准确反映其真实性能。类似的情况也曾发生在 Sol 身上,当时 Arena 不得不重组其编码基准测试,并增加了全栈基准测试,同时更新了其网络开发基准测试,以在 Astra 发布后更好地反映实际性能。这表明 OpenAI 可能并未完全针对基准测试进行优化。
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月8日 16:36 UTC
收录当时偏移:UTC+02026年9月9日 15:00 UTC
- 发布
- 2026年9月8日 16:36
- 收录
- 2026年9月9日 15:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 同步延迟
档位是按信源手工设定的编辑判断,不是逐条打分。
正文
Something similar happened with Sol and Arena had to restructure their coding benchmark and added a full-stack bench, plus updated the webdev bench to reflect real world performance after Astra released. Looks OpenAI is only lab not benchmaxxing