跳到正文
RCreddit.com·
暂不在当前实时榜单

5.5 is IN-SANE. It almost broke our benchmark. I thought it was a bug.

AI 摘要

Claude Opus 5.5 has achieved the top position on a writing benchmark, significantly outperforming other models. It scored #1 with 2600 Elo, costing $3.43 and taking 17 minutes per script. Other models, such as the #2 xhigh with 2399 Elo and the #3 high with 2342 Elo, showed lower performance and varying costs and times. The benchmark also included adaptive, medium, and low-tier models, with the lowest-ranked at #20 with 2144 Elo.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月27日 08:00 UTC

收录
2026年9月27日 08:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com