ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]
A recent study on agent task solving, detailed in a paper and a Hugging Face blog post, evaluated 9 models across 507 stateful workflows. The research highlights a significant difference between discovery (solving a task at least once) and repeatability (solving it 20/20 times). For instance, Kimi-K3 achieved the broadest coverage, solving 93.89% of tasks at least once, but only 13.41% consistently. In contrast, Claude Opus 5 discovered fewer tasks (79.09%) but repeated far more (47.53%), indicating that ranking models based on these different metrics yields nearly reversed leaderboards.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Oct 9, 2026, 00:50 UTC
IngestedOffset at this time: UTC+0Oct 9, 2026, 08:00 UTC
- Published
- Oct 9, 2026, 00:50
- Ingested
- Oct 9, 2026, 08:00
- Source type
- Dev community
- Tier
- Community
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom.
We wanted to know how much of a single agent success rate survives repetition, and whether "the agent finished the task" means the backend/database actually ended up in the correct state.
- Thinkingbox-Bench includes 507 policy conditioned business workflows across 5 domains (retail, travel/hospitality, auto insurance, neobank internal IT, consulting IT/HR)
- Each task is run in 20 independently executed attempts, each starting from an identical clean backend: 10,140 trials per model
- Grading compares the terminal backend state and side effects against the required end state. Any trajectory that produces the right outcome passes; wrong, missing or extra effects fail. 477 of 507 tasks are graded on state alone; 30 also check a narrow property of the final response
- all-20 — fraction of tasks solved on every one of the 20 attempts. This is an observed count on a fixed trial budget, not an estimator
Discovery and repeatability rank models very differently. The attached figure plots all three metrics for 9 models, and the orange-to-green spread is the whole point. Kimi-K3 has the broadest coverage we measured, solving 93.89% of tasks at least once (476/507), but only 13.41% (68/507) on all 20. Claude Opus 5 discovers fewer (79.09%) and repeats far more (47.53%, or 241 tasks). Qwen3.8-27B: 89.35% at least once, 7.50% every time. Ranking by pass@20 and ranking by all-20 give you nearly reversed leaderboards.
Failures often look clean. In a retrospective ablation over 121,680 valid trials across 12 models, 79,853 failed the executable checks. 67.24% of those failures still terminated cleanly, invoked a state changing tool, and ended without a final tool error, so a completion-style proxy would have scored them as done. Of those clean terminating failures, the state checks found (categories overlap):
- 20/20 on our trial budget is an observed count, not a guarantee of future reliability
- The simulated user is a fixed LLM; that is a source of variance we discuss in the appendix
You can run any of the 507 tasks against your own model through HF OpenEnv environment, which is the fastest way to disagree with us using your own numbers.
We report both the observed all-20 count and a plug-in pass k estimate in the appendix. They answer different questions and can differ substantially: one records what happened in a fixed 20-attempt campaign, while the other estimates repeated success under additional assumptions. Which would you want emphasized on a reliability leaderboard or should both be shown?