·
Archived topic · source no longer tracked
Orca-Bench: How Ready Are Language Model Agents for Oncall?
ORCA-bench is a new benchmark designed to evaluate language model agents in a production-fidelity oncall setting for root cause analysis (RCA). It uses a live OpenTelemetry-instrumented microservice system with six days of metrics, logs, and traces, and 1,079 RCA tasks. Expert SREs curate ground-truth symptoms. The best agents achieved only 25.3% RCA Accuracy on Medium-difficulty tasks and 10.0% on Hard tasks, even with Claude Fable 5. This indicates a significant gap before these agents can be safely entrusted with production reliability.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Jul 31, 2026, 21:00 UTC
- Ingested
- Jul 31, 2026, 21:00
- Source type
- Unclassified