🔍 Read the full analysis: A Guide To Reproducing AI Benchmark Results With UK AISI And EvalEval on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected results from five benchmarks and two cyber evaluations through EvalEval’s Evaluation Cards, which pair results with verification and evaluation details. The accompanying paper shows that inference-time compute and correctness feedback can affect measured performance; the release does not claim to include all AISI evaluations or transcripts.
The UK AI Security Institute (AISI) is publishing selected results from its AI evaluations through EvalEval’s Evaluation Cards, adding verification, context and configuration details to help readers understand how the results were produced, as the original analysis details. The release covers five benchmarks across six frontier models, as well as two cyber evaluations with a separate, partly overlapping model set.
The five benchmarks in the main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The model results listed for those benchmarks cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results for Cyber CTFs and The Last Ones. Those cyber evaluations use a different set of models, and the announcement does not provide a complete list for them.
The records accompany AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how evaluation results vary with inference-time compute and protocol, a topic also covered in published defense LLM benchmark results. EvalEval says the cards combine results with evaluation context and configuration information. They draw on a shared format for benchmark metadata, evaluation-run data and model metadata, making individual records easier to inspect alongside other reported runs.
The release describes selected publicly reported methods and findings being made available “where appropriate.” It does not say that every AISI evaluation or every underlying transcript is included. The announcement also does not establish that outside researchers have independently reproduced the results.
Why Evaluation Conditions Matter
Benchmark scores are often cited as if they were directly comparable, but the protocol and compute used can change what a score represents. AISI’s analysis of Humanity’s Last Exam tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.
That finding makes the reporting details relevant to researchers, model developers and policymakers who use evaluations as evidence about AI capabilities. Pairing results with setup information can help readers see whether two apparently similar scores came from comparable conditions. The cards do not determine which benchmark or protocol is best, and they cannot by themselves settle questions about reproducibility. They make some of the conditions behind published results more visible.
The release builds on work between AISI and EvalEval that began at a joint workshop held alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current cards apply that infrastructure to selected AISI methods and findings.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in transcript analysis and capability elicitation. EvalEval’s Evaluation Cards bring results together with benchmark and model information. These efforts address a reporting problem: evaluation results appear across different formats, sometimes without details needed to interpret a run, while repeating costly evaluations may not be practical.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
What the Release Leaves Open
The announcement does not specify how many records or transcripts are available, which setup fields appear on every card, or whether external researchers have reproduced the findings. It also does not enumerate the model set used for Cyber CTFs and The Last Ones, explain how disagreements between results from different protocols would be handled, or give a record-by-record publication schedule. The available information should be treated as a selected release rather than a complete archive of AISI evaluation work.
Wider Use of EEE
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. Under the EEE schema, model developers can submit verified results and evaluation developers can report benchmark and run data. Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection.
No further release date or adoption milestone was specified. The usefulness of broader comparisons will depend on whether contributors publish records with consistent, sufficiently complete details.
Key Questions
Which benchmarks are covered in AISI’s main experiment?
The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
Which models are listed for those five benchmarks?
The listed models are Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The cyber evaluations use a different, partly overlapping model set.
What information do EvalEval’s Evaluation Cards provide?
EvalEval describes the cards as including evaluation results, verification, context and configuration information, organised alongside benchmark and model metadata.
Does this release include all AISI evaluations?
No such claim was made. The announcement says selected publicly reported methods and findings are being shared where appropriate; it does not describe a complete archive.
Why can inference-time compute affect benchmark scores?
AISI’s paper examines how scores change with token use and evaluation protocol. Its Humanity’s Last Exam analysis reports that runs with correctness feedback after each attempt went on to solve additional tasks as token use increased.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
