Skip to content
RCreddit.com·
Not on the current live radar

Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time

AI summary

MindTrial benchmarks show Sonnet 5.5 and Opus 5.5 significantly improve performance on a 98-task suite. Sonnet 5.5 jumped from 72/98 to 94/98, reducing request time by 74.9% and output tokens by 67.4%. Opus 5.5 achieved 96/98 with a request time of 0:58:35, compared to Opus 5's 88/98 and 3:40:24. Both new models passed all 39 text tasks, with Sonnet 5.5 showing a notable improvement in visual tasks from 12/26 to 25/26.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 10:00 UTC

Ingested
Oct 4, 2026, 10:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com