跳到正文
RCreddit.com·
暂不在当前实时榜单

Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time

AI 摘要

MindTrial benchmarks show Sonnet 5.5 and Opus 5.5 significantly improve performance on a 98-task suite. Sonnet 5.5 jumped from 72/98 to 94/98, reducing request time by 74.9% and output tokens by 67.4%. Opus 5.5 achieved 96/98 with a request time of 0:58:35, compared to Opus 5's 88/98 and 3:40:24. Both new models passed all 39 text tasks, with Sonnet 5.5 showing a notable improvement in visual tasks from 12/26 to 25/26.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 10:00 UTC

收录
2026年10月4日 10:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com