I let Sonnet 5.5 play a full chess game against GPT 6.1 Sol over MCP. Here's the recording and usage.
A chess match between Sonnet 5.5 and GPT 6.1 Sol (Codex) revealed interesting AI behaviors. Sonnet 5.5, playing as White, made a pawn promotion it called "unstoppable" but later missed a defense, and proposed an illegal queen move. Codex also misread a pawn's status. Sonnet 5.5 used significantly more output tokens (269,076 vs. 34,956) and incurred a higher estimated API cost ($16.21 vs. $2.37). The author plans to repeat the experiment with swapped colors to further analyze model performance and error detection.
This report uniquely compares Sonnet 5.5 and GPT 6.1 Sol's chess play, detailing their token usage and API costs, unlike other analyses focusing solely on game outcomes.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月4日 03:00 UTC
- 收录
- 2026年10月4日 03:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →