A Guide To Reproducing AI Benchmark Results With UK AISI And EvalEval
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: A Guide To Reproducing AI Benchmark Results With UK AISI And EvalEval on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected results from five benchmarks and two cyber evaluations through EvalEval’s Evaluation Cards, which pair results with verification and evaluation details. The accompanying paper shows that inference-time compute and correctness feedback can affect measured performance; the release does not claim to include all AISI evaluations or transcripts.

The UK AI Security Institute (AISI) is publishing selected results from its AI evaluations through EvalEval’s Evaluation Cards, adding verification, context and configuration details to help readers understand how the results were produced, as the original analysis details. The release covers five benchmarks across six frontier models, as well as two cyber evaluations with a separate, partly overlapping model set.

The five benchmarks in the main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The model results listed for those benchmarks cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results for Cyber CTFs and The Last Ones. Those cyber evaluations use a different set of models, and the announcement does not provide a complete list for them.

The records accompany AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how evaluation results vary with inference-time compute and protocol, a topic also covered in published defense LLM benchmark results. EvalEval says the cards combine results with evaluation context and configuration information. They draw on a shared format for benchmark metadata, evaluation-run data and model metadata, making individual records easier to inspect alongside other reported runs.

The release describes selected publicly reported methods and findings being made available “where appropriate.” It does not say that every AISI evaluation or every underlying transcript is included. The announcement also does not establish that outside researchers have independently reproduced the results.

At a glance
reportWhen: Release announced; no specific publicat…
The developmentAISI is sharing selected AI evaluation results using EvalEval’s Evaluation Cards, alongside a paper examining how inference-time compute and evaluation protocols shape scores.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Conditions Matter

Benchmark scores are often cited as if they were directly comparable, but the protocol and compute used can change what a score represents. AISI’s analysis of Humanity’s Last Exam tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.

That finding makes the reporting details relevant to researchers, model developers and policymakers who use evaluations as evidence about AI capabilities. Pairing results with setup information can help readers see whether two apparently similar scores came from comparable conditions. The cards do not determine which benchmark or protocol is best, and they cannot by themselves settle questions about reproducibility. They make some of the conditions behind published results more visible.

A Shared Format for Evaluations

The release builds on work between AISI and EvalEval that began at a joint workshop held alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current cards apply that infrastructure to selected AISI methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring results together with benchmark and model information. These efforts address a reporting problem: evaluation results appear across different formats, sometimes without details needed to interpret a run, while repeating costly evaluations may not be practical.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

What the Release Leaves Open

The announcement does not specify how many records or transcripts are available, which setup fields appear on every card, or whether external researchers have reproduced the findings. It also does not enumerate the model set used for Cyber CTFs and The Last Ones, explain how disagreements between results from different protocols would be handled, or give a record-by-record publication schedule. The available information should be treated as a selected release rather than a complete archive of AISI evaluation work.

Wider Use of EEE

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under the EEE schema, model developers can submit verified results and evaluation developers can report benchmark and run data. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection.

No further release date or adoption milestone was specified. The usefulness of broader comparisons will depend on whether contributors publish records with consistent, sufficiently complete details.

Key Questions

Which benchmarks are covered in AISI’s main experiment?

The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Which models are listed for those five benchmarks?

The listed models are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The cyber evaluations use a different, partly overlapping model set.

What information do EvalEval’s Evaluation Cards provide?

EvalEval describes the cards as including evaluation results, verification, context and configuration information, organised alongside benchmark and model metadata.

Does this release include all AISI evaluations?

No such claim was made. The announcement says selected publicly reported methods and findings are being shared where appropriate; it does not describe a complete archive.

Why can inference-time compute affect benchmark scores?

AISI’s paper examines how scores change with token use and evaluation protocol. Its Humanity’s Last Exam analysis reports that runs with correctness feedback after each attempt went on to solve additional tasks as token use increased.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

OpenEuroLLM. The third path.

European project OpenEuroLLM faces resource challenges amid ambitious goals for multilingual LLMs, highlighting limits of pan-European AI collaboration.

Why Dataset Culture Matters More Than People Think

Outstanding dataset culture shapes decision-making quality and trust, but understanding its true impact will reveal why it matters more than you think.

Breaking Internal Barriers To AI Adoption

Organizations are overcoming internal resistance to AI by focusing on organizational change, partnerships, and workforce engagement, not just technology.

Which AI Tools Will Dominate In 2026?

Analyzing which AI tools are projected to dominate in 2026 based on industry trends, capabilities, and expert insights.