Session-Bench v1: what 12 coding harnesses preserve after the work is done. Same bug fix in each, session rebuilt from the files alone
Session-Bench v1, a follow-up to the v0.4 post, evaluates 12 coding harnesses by measuring what they preserve after a bug fix, with sessions rebuilt from files. DeepSeek Harness scored highest at 96.9, followed by Pi at 96.4 and Copilot CLI at 96.0. OpenCode, Kimi Code, OpenClaw, and Antigravity preserved 0% of events stored once, while Pi and Hermes preserved 100%. The full scorecard and replay instructions are available on jazzyalex.github.io.
Unlike the earlier v0.4, Session-Bench v1 runs the harnesses directly and measures their output, providing a more empirical comparison of 12 coding tools.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月9日 17:33 UTC
收录当时偏移:UTC+02026年10月9日 20:00 UTC
- 发布
- 2026年10月9日 17:33
- 收录
- 2026年10月9日 20:00
- 来源类型
- 开发者社区
- 档位
- 社区
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
Disclosure first: I maintain Session-Bench and Agent Sessions (a local session browser). This is the follow-up to the v0.4 post here in August. That version was twenty pass/fail gates, answered mostly from my own session corpus. v1 runs the harnesses and measures.
- One small task, two prompts: run a check that fails, take a correction, edit one file, run the check again. Four tool calls are scored.
- An outside witness records what really happened: the harness's live output stream, a ledger the check script writes, and file hashes.
- Then a decoder reads only the session files the harness left on disk, and 31 facts are compared with the witness: prompts, replies, tool calls, results, exit codes, the file before and after, model, token counts, timestamps, and whether the files open with standard tools.
- 100 points. Every score recomputes from the published session files with one command per run.
# Harness Score Bytes on disk Events stored once 1 DeepSeek Harness 96.9 193 KB 63% 2 Pi 96.4 19 KB 100% 3 Copilot CLI 96.0 182 KB 43% 4 OpenCode 94.4 235 KB 0% 5 Kimi Code 91.2 215 KB 0% 6 Claude Code 89.3 113 KB 81% 7 OpenClaw 88.1 697 KB 0% 8 Codex (CLI and Desktop) 87.5 185 KB 7% 9 Hermes 86.4 36 KB 100% 10 Claude Desktop 85.6 469 KB 75% 11 Antigravity 82.7 236 KB 0% 12 Cursor CLI 78.1 136 KB 28% What I think is worth discussing
- Nobody loses the conversation. Eleven of twelve get full marks on fidelity and causality: every prompt, reply, tool call, result and exit code comes back. The 19-point spread is made elsewhere.
- The same fix takes 19 KB in Pi and 697 KB in OpenClaw, and 3% of that 697 KB is the session. If you paste raw session files into another model for a summary or a handover, that is roughly 5,000 tokens against 174,000 for the same work.
- Four harnesses never write an event only once. OpenCode writes each message part again one to five times in an event table; Kimi writes each prompt three times. Anything that counts messages from the raw file over-counts.
- Only four of twelve record token usage that adds up to a stated total (DeepSeek Harness, Copilot CLI, OpenCode, Codex). Cursor CLI stores no token count in the session at all.
- Antigravity and Cursor CLI store protobuf inside SQLite with no published schema, so sqlite3 opens the file and cannot read the session.
- One synthetic task, three runs, one machine. Nothing about long sessions, compaction, sub-agents or crashes.
- The harnesses did not run the same model. This grades the files, not the model or the agent.
- Some scoring rules were written after the first captures, and three rules were waived or corrected after results were known. The report lists each one. Example: I first scored OpenCode at 97.4 and first place by leaving its event table out of the reading; an outside review caught that, and it is 94.4 and fourth.
- Reviews were done by separate AI agent sessions plus four rounds of outside model review, not by a second person.
- Is "token counts that add up to a stated total" the right bar for cost tracking, or is per-reply usage enough for you?
- Which harness or surface should be in the next run? Cursor Desktop is the one I could not score yet.
Scorecard, charts and the replay instructions: https://jazzyalex.github.io/agent-sessions/bench/v1/?campaign=reddit&ref=r-chatgptcoding-bench-update