UK AISI and EvalEval push to make frontier LLM benchmark results reproducible

UK AISI and EvalEval push to make frontier LLM benchmark results reproducible

The UK AI Security Institute (AISI) and the EvalEval Coalition have announced a new phase of their collaboration, which began at a joint workshop alongside NeurIPS 2025 and has already fed into EvalEval's Every Eval Ever (EEE) reporting schema. In this phase, AISI is releasing publicly reported evaluation methods, results, context and configuration information through EvalEval's Evaluation Cards platform, where appropriate. The release covers the five benchmarks used in the main experiment of AISI's new paper, How Inference Compute Shapes Frontier LLM Evaluation: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results span six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The release also includes results from two related cyber evaluations, Cyber CTFs and The Last Ones, which use a different, partially overlapping set of models. The underlying paper examines how benchmark performance changes with inference-time compute and with the evaluation protocol used. As an example, on Humanity's Last Exam the source describes curves showing the cumulative share of attempted tasks solved within a given token count, using the earliest observed success per task, and notes that when models received correctness feedback from an oracle after each attempt, they continued to solve additional tasks as token use increased. The stated purpose of the release is to let researchers and practitioners examine individual studies more closely, compare findings across the wider evaluation ecosystem, and use AISI's verified, documented results as reference points where other reports lack setup details. The blog post frames this as building on AISI's existing work to make evaluation more efficient (OptStop), more statistically rigorous (HiBayES), and more standardised in areas like transcript analysis and capability elicitation. It closes with a call to action: model developers are asked to report verified evaluation results, evaluation developers to report benchmarks and run data using the EEE schema, and evaluation, governance and policy researchers to explore Evaluation Cards by benchmark or model. No specific numeric scores for any of the six models on the five benchmarks are given in the text itself.

Key facts

  • AISI is releasing publicly reported evaluation methods, results, context and configuration data through EvalEval's Evaluation Cards platform.
  • The release covers five benchmarks from AISI's paper How Inference Compute Shapes Frontier LLM Evaluation: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
  • Results span six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4, plus two related cyber evaluations (Cyber CTFs, The Last Ones) using a partially overlapping model set.
  • The paper finds that benchmark performance, illustrated with Humanity's Last Exam, depends on the evaluation protocol and on how much inference-time compute (token count) models are given, with oracle correctness feedback letting models solve more tasks as token use rises.
  • The collaboration builds on AISI tools including OptStop (evaluation efficiency) and HiBayES (statistical rigor), and on EvalEval's Every Eval Ever schema, which AISI's feedback helped shape.

Why it matters

AI evaluations are increasingly cited as evidence about model and system performance, but the source says results are reported across many formats and outlets, often without enough detail to reproduce them, and rerunning evaluations can itself be prohibitively expensive. By publishing verified methods, configuration and context alongside results, AISI and EvalEval aim to close that reproducibility gap and give the wider field common reference points.

Who it affects

The release is aimed at model developers, evaluation developers, and evaluation, governance and policy researchers. AISI, a research organisation inside the UK government's Department for Science, Innovation and Technology tasked with giving governments a scientific understanding of advanced AI risks, is the source of the released results; EvalEval, a research community building infrastructure for the evaluation ecosystem, hosts them.

How to use it

AISI's verified results, methods and configuration data for the five named benchmarks and six models are available through EvalEval's Evaluation Cards platform. The source invites model developers to report their own verified evaluation results, evaluation developers to submit benchmarks and run data using the Every Eval Ever schema, and researchers to browse Evaluation Cards by benchmark or model.

How solid is it

The account comes directly from a joint AISI/EvalEval announcement describing their own release and paper, so the description of what is being published is authoritative for that release. The source does not give specific numeric scores for any model on any benchmark, nor a publication date for the blog post or the paper, or a byline for the post.

Risks and caveats

The two cyber evaluations included in the release, Cyber CTFs and The Last Ones, use a different and only partially overlapping set of models than the five main benchmarks, which the source flags but does not fully detail. No performance figures are given here to independently assess the paper's inference-compute findings, and the exact dates of the NeurIPS 2025 workshop where the collaboration began are not specified beyond the year.

“We are excited about this adoption and look forward to further standardising and sharing evaluations with AISI and other AI evaluation organisations.”

— EvalEval Coalition