跳到正文
HChuggingface.co·

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

AI 摘要

The UK AI Security Institute (AISI) is utilizing EvalEval's infrastructure to openly share evaluation results, enhancing the reproducibility and verifiability of evaluation science. These results encompass six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4, alongside data from Cyber CTFs and The Last Ones cyber evaluations. This release supports AISI's paper, "How Inference Compute Shapes Frontier LLM Evaluation," which investigates the impact of inference-time compute and evaluation protocols on benchmark performance.

为什么是这条

This release is the first time AISI has openly shared evaluation results using EvalEval's infrastructure, unlike previous evaluations that lacked such verifiable data.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月22日 00:00 UTC

收录当时偏移:UTC+02026年9月22日 18:01 UTC

发布
2026年9月22日 00:00
收录
2026年9月22日 18:01
来源类型
官方发布
档位
当事方
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

→ 平稳
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

The EvalEval Coalition is thrilled to share that the UK AI Security Institute (AISI) is using EvalEval's infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.

AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025 , and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema . This next phase of the collaboration puts that shared infrastructure into practice.

Why reproducible evaluation reporting matters

As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them. Running the evaluations again may itself be prohibitively expensive.

EvalEval's mission is to improve this ecosystem through a shared reporting schema, Every Eval Ever , and an open platform, Evaluation Cards , that brings evaluation results and the information needed to interpret them into a common structure.

This builds naturally on AISI's work to make evaluation more efficient through OptStop , more statistically rigorous through HiBayES , and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.

What AISI is sharing

Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis. In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context, and configuration information for the five benchmarks in the paper's main experiment:

- HealthBench

- FrontierMath

- Humanity's Last Exam

- SWE-Bench Pro

- Terminal-Bench 2.0

These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models. The data accompany AISI's paper, How Inference Compute Shapes Frontier LLM Evaluation , which studies how benchmark performance depends on inference-time compute and evaluation protocol.

Performance on Humanity's Last Exam changes with evaluation protocol and inference compute. Each curve shows the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task. When models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased.

When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem. Where other reports lack these details, releases like AISI's provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance. As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.

AISI's Terminal-Bench 2.0 results alongside other reported evaluations for the same models, under different evaluation setups.

We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.

Contribute to the shared mission

- Model developers: Report verified evaluation results .

- Evaluation developers: Report benchmarks and run data using the Every Eval Ever schema .

- Evaluation, governance, and policy researchers: Explore Evaluation Cards by benchmark or model, or use it to examine the state of evaluation reporting as a whole.

About the EvalEval Coalition

The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem. Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis.