SWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results from three models [P]
SWE-Race is a new benchmark for coding agents, featuring 188 real concurrency bugs like race conditions and deadlocks, sourced from approximately 100 Python projects. Each task is graded using the project's own tests in an isolated container, preventing agents from accessing historical fixes. GLM-5.3 Flash achieved an 85% score with one attempt, and 82% with multiple attempts, comparable to GPT-5.6 Luna's 81%. The benchmark follows the DeepSWE protocol, and feedback on models to test next is encouraged.
This benchmark is unique in its focus on real-world concurrency bugs from 100 Python projects, unlike other benchmarks that might use synthetic or less complex issues.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 6, 2026, 14:00 UTC
- Ingested
- Oct 6, 2026, 14:00
- Source type
- Dev community
Full text isn't available here.
Read at source →