跳到正文
RCreddit.com·
暂不在当前实时榜单

SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]

AI 摘要

SWE-Race is a new benchmark for coding agents, featuring 188 real concurrency bugs like race conditions and deadlocks, sourced from approximately 100 Python projects. Each task is graded using the project's own tests in an isolated container, preventing agents from accessing historical fixes. GLM-5.3 Flash achieved an 85% score with one attempt, and 82% with multiple attempts, comparable to GPT-5.6 Luna's 81%. The benchmark follows the DeepSWE protocol, and feedback on models to test next is encouraged.

为什么是这条

This benchmark is unique in its focus on real-world concurrency bugs from 100 Python projects, unlike other benchmarks that might use synthetic or less complex issues.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月6日 14:00 UTC

收录
2026年10月6日 14:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com