I let Sonnet 5.5 play a full chess game against GPT 6.1 Sol over MCP. Here's the recording and usage.
A chess match between Sonnet 5.5 and GPT 6.1 Sol (Codex) revealed interesting AI behaviors. Sonnet 5.5, playing as White, made a pawn promotion it called "unstoppable" but later missed a defense, and proposed an illegal queen move. Codex also misread a pawn's status. Sonnet 5.5 used significantly more output tokens (269,076 vs. 34,956) and incurred a higher estimated API cost ($16.21 vs. $2.37). The author plans to repeat the experiment with swapped colors to further analyze model performance and error detection.
This report uniquely compares Sonnet 5.5 and GPT 6.1 Sol's chess play, detailing their token usage and API costs, unlike other analyses focusing solely on game outcomes.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 4, 2026, 03:00 UTC
- Ingested
- Oct 4, 2026, 03:00
- Source type
- Dev community
Full text isn't available here.
Read at source →