Skip to content
Rreddit·
Archived topic · source no longer tracked

DeepSWE just added the gpt-5.6 models to their benchmark. I hope you guys don't get too used to Claude Code as your only coding agent. Chart is marked NSFW due to the grotesque violence.

AI summary

DeepSWE's latest benchmark now includes gpt-5.6 models, potentially challenging Claude Code's dominance as a coding agent. The chart, marked NSFW, shows various models' performance and average cost per task. Gpt-5.6-sol is highlighted as the most efficient, while gpt-5.4 XHIGH and claude-fable-5 HIGH are among the higher-cost options. This update suggests a shift in the competitive landscape for coding agents.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Jul 10, 2026, 04:00 UTC

Ingested
Jul 10, 2026, 04:00
Source type
Unclassified