跳到正文
RCreddit.com·
暂不在当前实时榜单

Why decontamination reports can't fix benchmark contamination, and what an evaluator has to do instead [D]

AI 摘要

OpenAI stopped reporting SWE-bench Verified in February, recommending other labs do the same, as frontier models could reproduce human-written fixes or problem details for some tasks. Progress had slowed, and the benchmark's creator retired it. The author, who built a similar system, acknowledges limitations like the benchmark's quality, the potential for repeated submissions to compromise hidden test sets, and the risk of leaked labels, prioritizing closing the gap related to repeated submissions.

为什么是这条

This post uniquely details the challenges of benchmark contamination, unlike typical reports that focus on decontamination, and highlights the specific limitations of current evaluation methods.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月20日 23:01 UTC

收录
2026年9月20日 23:01
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com