返回
RCreddit.com

Coding benchmarks that are quickly showcasing deep capability

时间与来源
发布
09/06 12:23
收录
09/06 16:00
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。

While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence, and complete capability in Software Engineering :

1. Program-Bench

Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior (without access to decompilers or internet). Link: https://programbench.com/

- GPT-6 Astra: 5.5%

- Fable 5.1: 7%

- Kimi K3: 2%

- Qwen3.8 27b: 0%

- GPT 5.6 Sol: 1.5%

- GLM 5.3: 1.5%

- GPT 5.6 Luna: 0%

2. SRE-Bench

Can AI agents work out what a real-world binary does without its source code?

Link: https://www.vals.ai/benchmarks/srebench

Sure nobody is reading assembly code in daily work, it is hard. The ability to understand a compiled program is insane ability.

- GPT-6 Astra: 88%

- GPT-5.6 Sol: 55.9%

- Claude Opus 5 (max): 12.5%

3. Code Migration

Can language models reimplement working programs in another language?

Link: https://www.vals.ai/benchmarks/code-migration

- GPT-6 Astra: 67.7%

- Fable 5.1: 54.6%

- GLM 5.3: 44.2%

- GLM 5.3 Flash: 20.5%