Claude Code Beats Codex in a Negotiation Competition
Heat trend
Collecting trend data
The percentage is based on available heat signal, not comment count or independent people.
Claude Code and Codex, prominent coding agents, are increasingly used for language-based tasks like negotiation. In a recent competition, Claude Code demonstrated superior negotiation skills against Codex across multiple games. Claude Code won two out of three initial games, committing 50,000 troops and securing Baltic ports. In later rounds, Claude Code (Opus 4.8) played as Napoleon against Codex (GPT-5.5) as Poniatowski, further showcasing its capabilities in complex negotiation scenarios.
People are now using Claude Code and Codex, two of the leading coding agents, to do almost everything, including tasks that have more to do with language than coding, such as negotiation.
For example, OpenAI recently highlighted a use case where Codex negotiated with customer service to get a refund on behalf of a user .
But can you really trust an agent to represent your best interests? And if so, which agent should you trust?
Put them in a negotiation competition with carefully designed cases, information gaps, conflicting interests, and systematic, objective evaluation.
I purchased The Negotiation Challenge: How to Win Negotiation Competitions and created an agent negotiation competition (link in the blog) based on one of its original cases, the Battle of Nations, which was designed based on the 1813 German War of Liberation.
In this negotiation Napoleon and Poland need to reach a deal on the following issues.
- How many Baltic seaports Napoleon will hand to Poland (more ↑: Napoleon - per port, constant / Poniatowski ++→ +)
- Whether Poniatowski receives the baton of an Imperial Marshal (yes: Napoleon -tiny / Poniatowski + small; mildly positive-sum)
The objective score is calculated from the final agreement reached by the parties. Each negotiable issue is assigned a point value in advance, based on how important that issue is to each side. After the negotiation ends, the agreed terms are converted into points according to the scoring sheet.
Games 1–3: Claude Code (Opus 4.8) as Poniatowski, Codex (GPT-5.5) as Napoleon
Game Claude Code Objective Codex Objective Objective Winner Troops Committed Days Held Baltic Ports Ceded Poland Restored Marshal Title Marriage to Pauline Rounds (12 Max) G1 69.41 23.08 Claude Code 50,000 3 4 yes yes no 4 G2 62.23 28.85 Claude Code 50,000 3 3 yes yes no 5 G3 45.99 58.11 Codex 50,000 4 2 yes yes no 5 Games 4–6: Claude Code (Opus 4.8) as Napoleon, Codex (GPT-5.5) as Poniatowski
Game Claude Code Objective Codex Objective Objective Winner Troops Committed Days Held Baltic Ports Ceded Poland Restored Marshal Title Marriage to Pauline Rounds (12 Max) G4 66.22 41.97 Claude Code 60,000 4 2 yes yes no 4 G5* 66.22 41.97 Claude Code 60,000 4 2 yes yes no 4 G6 66.22 41.97 Claude Code 60,000 4 2 yes yes no 4 Game 7: Claude Code (Opus 4.8) as Poniatowski, Codex (GPT-5.6 Sol) as Napoleon
Game Claude Code Objective Codex Objective Objective Winner Troops Committed Days Held Baltic Ports Ceded Poland Restored Marshal Title Marriage to Pauline Rounds (20 Max) G7 46.02 34.62 Claude Code 70,000 3 4 yes yes yes 4 Game 8: Claude Code (Opus 4.8) as Napoleon, Codex (GPT-5.6 Sol) as Poniatowski
Game Claude Code Objective Codex Objective Objective Winner Troops Committed Days Held Baltic Ports Ceded Poland Restored Marshal Title Marriage to Pauline Rounds (20 Max) G8 74.32 38.53 Claude Code 60,000 4 2 yes yes yes 4
Despite a fully competitive setting (which Codex fully understood), Codex placed too much weight on reaching an agreement quickly and too little on continuing to extract value.
Every game had capacity for more than 10 rounds, yet all of them closed at Round 4/5
The closing rationales repeatedly relied on 5 distinct high-frequency keywords: "meets the hard constraints," "safe," "complete," "acceptable," and "signable."
Examining the agents' records, I found that Claude Code usually showed longer and more structured plans, whereas Codex's visible pre-negotiation notes often did little more than summarize the private brief.
Game Claude Code Role Codex Role Target Red Lines Chip Valuation Decision Tree Disclosure Strategy BATNA Management Pre-Sign Check G1 Poniatowski Napoleon ✓ / — ✓ / — ✓ / — △ / — — / — △ / — — / — G2 Poniatowski Napoleon △ / ✓ ✓ / ✓ ✓ / — △ / — △ / △ △ / — — / — G3 Poniatowski Napoleon △ / ✓ ✓ / ✓ ✓ / △ △ / △ ✓ / △ △ / — — / — G4 Napoleon Poniatowski ✓ / ✓ ✓ / ✓ ✓ / △ — / — — / △ △ / — — / — G5 Napoleon Poniatowski ✓ / ✓ ✓ / ✓ ✓ / ✓ — / — — / — △ / — — / — G6 Napoleon Poniatowski ✓ / — ✓ / ✓ ✓ / ✓ — / — △ / — — / △ — / — G7 (GPT-5.6 Sol) Poniatowski Napoleon △ / — ✓ / — ✓ / — ✓ / — ✓ / — ✓ / — △ / — G8 (GPT-5.6 Sol) Napoleon Poniatowski ✓ / — ✓ / ✓ ✓ / — △ / — ✓ / — ✓ / — — / — Table 2: Pre-negotiation plan lengths (English characters, brief-received → first own action, opponent content excluded)
Game Codex Role Codex Plan Claude Code Role Claude Code Plan G1 Napoleon 143 Poniatowski 1,379 G2 Napoleon 300 Poniatowski 1,291 G3 Napoleon 139 Poniatowski 1,632 G4 Poniatowski 510 Napoleon 869 G5 Poniatowski 797 Napoleon 1,193 G6 Poniatowski 252 Napoleon 765 G7 (GPT-5.6 Sol) Napoleon 0 Poniatowski 1,879 G8 (GPT-5.6 Sol) Poniatowski 311 Napoleon 981
This habit likely came from the assistant's refuse-but-offer-alternative template ("I can't do X, but I can offer Y"), which post-training rewards in every helpful chatbot.
In this case, Codex did not treat the other party's arguments as moves made by an interested party; it absorbed them as neutral facts and let them set prices.
https://preview.redd.it/kynuw9ypzcnh1.png?width=2336&format=png&auto=webp&s=641204fec4c43a66cebf0c68aace1672db903dc0
The agent hired to negotiate spent the majority of its classified vocabulary narrating the machinery: whether the API was up, when to poll next, how its self-built notification loop was doing.