Can Shared Evaluation Methods Improve AI Benchmark Reproducibility?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can Shared Evaluation Methods Improve AI Benchmark Reproducibility? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, which pair scores with verification, context and configuration details. The release covers five benchmarks across six frontier models, plus two cyber evaluations using a partly different model set.

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, a format that pairs scores with verification, evaluation context and configuration information so readers can see how each result was produced. The release covers results for five benchmarks across six frontier models, along with two cyber evaluations that use a different, partly overlapping model set. The records accompany AISI’s paper How Inference Compute Shapes Frontier LLM Evaluation.

The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones, but those runs use a model set that only partly overlaps with the main experiment, so the same model list should not be assumed to apply to them.

EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI’s public reporting is available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.

The associated paper examines how model scores depend on inference-time compute and evaluation protocol. For Humanity’s Last Exam, the analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased — a finding that shows the same benchmark can yield different results under different conditions.

At a glance
announcementWhen: announced alongside AISI’s paper on inf…
The developmentThe UK AI Security Institute has begun publishing evaluation results through EvalEval’s Evaluation Cards, releasing records covering five benchmarks across six frontier models alongside setup details needed to interpret the runs.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Setup Details Change Score Comparisons

Benchmark scores are often cited as if they measure the same thing across models, but the AISI paper illustrates that different evaluation protocols can produce different results for the same benchmark. In the Humanity’s Last Exam analysis, outcomes shifted with inference compute and with whether models received correctness feedback between attempts. A score reported without those conditions leaves readers unsure what performance it represents.

Publishing results with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that uses evaluations as evidence about advanced AI capabilities. The records do not by themselves settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see.

From Shared Schema to Public Records

The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that shared infrastructure to publicly reported AISI methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information. Together, these efforts address a practical reporting problem: results published across formats and outlets may omit details needed to interpret or reproduce a run, while repeating costly evaluations may not be feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Coverage and Reproduction Limits

The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. It says publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of all AISI evaluation work.

The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare records consistently.

Broader Adoption of Every Eval Ever

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection.

Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of records that contributors publish. No further release date or adoption milestone was specified.

Key Questions

What is AISI publishing through EvalEval?

Selected AI benchmark results in the form of Evaluation Cards, which combine scores with verification, evaluation context and configuration details. The release accompanies AISI’s paper on inference-time compute and evaluation protocols.

Which models and benchmarks are covered?

The main experiment covers HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0 across Claude Opus 4, 4.5 and 4.6, and GPT-5, 5.2 and 5.4. Two cyber evaluations — Cyber CTFs and The Last Ones — use a different, partly overlapping model set that has not been fully enumerated.

Why does evaluation protocol matter for benchmark scores?

According to AISI’s paper, scores can shift depending on inference-time compute and feedback conditions. For example, in Humanity’s Last Exam runs where models received correctness feedback after each attempt, they solved additional tasks as token use increased.

Does the release include all of AISI’s evaluation work?

No. The announcement says publicly reported methods and findings are available where appropriate; it does not claim every AISI evaluation or underlying transcript is included, and it does not specify how many records exist.

Can other organisations contribute to the shared format?

Yes. EvalEval says model developers can submit verified results and evaluation developers can report benchmarks and run data using the Every Eval Ever schema, though no adoption milestone has been announced.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Build vs Buy a Prebuilt AI Workstation

An analysis of the rising costs and benefits of building or buying prebuilt AI workstations amid 2026 component shortages and AI boom.

Five AI Managers Faced an Impostor—and Held the Line

Five frontier AI models rejected fake-CEO pressure in Firmulate’s live company test, showing integrity can be tested before deployment—and before a real breach.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX’s recent $60 billion all-stock purchase of AI coding firm Cursor is a strategic move, offering growth and competitive advantages amid rising AI valuations.

Bitcoin Battles Unfold in Live Warzone Visualization

A new web-based visualization transforms Bitcoin trading data into an immersive, cinematic battlefield, highlighting market dynamics without trading advice.