跳到正文
·
Archived topic · 归档话题,来源已停止追踪

Orca-Bench: How Ready Are Language Model Agents for Oncall?

AI 摘要

ORCA-bench is a new benchmark designed to evaluate language model agents in a production-fidelity oncall setting for root cause analysis (RCA). It uses a live OpenTelemetry-instrumented microservice system with six days of metrics, logs, and traces, and 1,079 RCA tasks. Expert SREs curate ground-truth symptoms. The best agents achieved only 25.3% RCA Accuracy on Medium-difficulty tasks and 10.0% on Hard tasks, even with Claude Fable 5. This indicates a significant gap before these agents can be safely entrusted with production reliability.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年7月31日 21:00 UTC

收录
2026年7月31日 21:00
来源类型
未分类