跳到正文
RCreddit.com·

ThinkingBox: Solving an agent task once vs. solving it 20/20: 507 stateful workflows graded on terminal database state [R]

AI 摘要

A recent study on agent task solving, detailed in a paper and a Hugging Face blog post, evaluated 9 models across 507 stateful workflows. The research highlights a significant difference between discovery (solving a task at least once) and repeatability (solving it 20/20 times). For instance, Kimi-K3 achieved the broadest coverage, solving 93.89% of tasks at least once, but only 13.41% consistently. In contrast, Claude Opus 5 discovered fewer tasks (79.09%) but repeated far more (47.53%), indicating that ranking models based on these different metrics yields nearly reversed leaderboards.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 00:50 UTC

收录当时偏移:UTC+02026年10月9日 08:00 UTC

发布
2026年10月9日 00:50
收录
2026年10月9日 08:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

Disclosure: I'm one of the authors (Microsoft). The paper, code, dataset are public and ThinkingBox is on Hugging Face OpenEnv as well. Raw evaluation trajectories are not released Links at the bottom.

We wanted to know how much of a single agent success rate survives repetition, and whether "the agent finished the task" means the backend/database actually ended up in the correct state.

- Thinkingbox-Bench includes 507 policy conditioned business workflows across 5 domains (retail, travel/hospitality, auto insurance, neobank internal IT, consulting IT/HR)

- Each task is run in 20 independently executed attempts, each starting from an identical clean backend: 10,140 trials per model

- Grading compares the terminal backend state and side effects against the required end state. Any trajectory that produces the right outcome passes; wrong, missing or extra effects fail. 477 of 507 tasks are graded on state alone; 30 also check a narrow property of the final response

- all-20 — fraction of tasks solved on every one of the 20 attempts. This is an observed count on a fixed trial budget, not an estimator

Discovery and repeatability rank models very differently. The attached figure plots all three metrics for 9 models, and the orange-to-green spread is the whole point. Kimi-K3 has the broadest coverage we measured, solving 93.89% of tasks at least once (476/507), but only 13.41% (68/507) on all 20. Claude Opus 5 discovers fewer (79.09%) and repeats far more (47.53%, or 241 tasks). Qwen3.8-27B: 89.35% at least once, 7.50% every time. Ranking by pass@20 and ranking by all-20 give you nearly reversed leaderboards.

Failures often look clean. In a retrospective ablation over 121,680 valid trials across 12 models, 79,853 failed the executable checks. 67.24% of those failures still terminated cleanly, invoked a state changing tool, and ended without a final tool error, so a completion-style proxy would have scored them as done. Of those clean terminating failures, the state checks found (categories overlap):

- 20/20 on our trial budget is an observed count, not a guarantee of future reliability

- The simulated user is a fixed LLM; that is a source of variance we discuss in the appendix

You can run any of the 507 tasks against your own model through HF OpenEnv environment, which is the fastest way to disagree with us using your own numbers.

We report both the observed all-20 count and a plug-in pass k estimate in the appendix. They answer different questions and can differ substantially: one records what happened in a fixed 20-attempt campaign, while the other estimates repeated success under additional assumptions. Which would you want emphasized on a reliability leaderboard or should both be shown?

来源·reddit.com