Back
Hhackernews·shiqimei
24
·3 hr ago·Other · Official API

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

View original
Model release

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

AI summary

FrontierHarness Eval conducted an evaluation across 9 harnesses using the same model, revealing a 17x variation in cost per pass. The evaluation, which covered 15 passes, showed that the cost per task increased to $3.24 when failed attempts were included. The tests were performed on Runta agent runtimes, with each run being a fresh restore from a golden checkpoint, ensuring identical vCPU, memory, disk size, disk contents, and memory state for consistency.

Pass rate

Median cost per successful task

Median cost per task

Median cache hit rate per successful task

Median time per successful task

Beyond the numbers

- 01 OpenCode: failures excluded.

It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task.

- 02 Cache hit rate is not cost.

A cached 300-turn failure can still burn more than a short cache miss.

- 03 Quality and cost can diverge.

Claude Code passes 19 tasks, but reaches $18.34 in cost per task.

Run your harness on Runta.

If you want to test your own harness on Runta, we’ll give you $100 in credits to get started.

Get $100 in credits Start free trial

Tested harnesses

Codex

v0.148.0

DeepSeek Harness

v0.1.0-rc.8

Claude Code

v2.1.237

Pi

v0.84.2

Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x · BuzzRadr