Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
FrontierHarness Eval 对 9 种不同的测试工具(harness)进行了评估,所有测试均使用相同的模型,结果显示每次通过的成本差异高达 17 倍。这项评估涵盖了 15 次通过,如果将失败的尝试也计算在内,每个任务的成本将增至 3.24 美元。所有评估都在 Runta 代理运行时上进行,每次运行都从一个黄金检查点(golden checkpoint)进行全新恢复,确保了 vCPU、内存、磁盘大小、磁盘内容和内存状态完全一致,以保证测试环境的统一性。
Pass rate
Median cost per successful task
Median cost per task
Median cache hit rate per successful task
Median time per successful task
Beyond the numbers
- 01 OpenCode: failures excluded.
It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task.
- 02 Cache hit rate is not cost.
A cached 300-turn failure can still burn more than a short cache miss.
- 03 Quality and cost can diverge.
Claude Code passes 19 tasks, but reaches $18.34 in cost per task.
Run your harness on Runta.
If you want to test your own harness on Runta, we’ll give you $100 in credits to get started.
Get $100 in credits Start free trial
Tested harnesses
Codex
v0.148.0
DeepSeek Harness
v0.1.0-rc.8
Claude Code
v2.1.237
Pi
v0.84.2