TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The UK AI Security Institute is using EvalEval’s Evaluation Cards to publish evaluation results alongside verification, context and configuration details. The release covers five benchmarks across six frontier models, plus two cyber evaluations with a different, partly overlapping model set.
The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, adding verification, context and configuration details that help readers interpret how the results were produced. The release accompanies AISI’s paper on inference-time compute and evaluation protocols, and includes results for five benchmarks across six frontier models.
The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones. Those use a different set of models that overlaps only partly with the main experiment, so the same model list should not be assumed to apply to them.
The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The paper examines how scores depend on inference-time compute and evaluation protocol. For Humanity’s Last Exam, the reported analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased.
EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI’s public reporting is available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.
Evaluation reporting · UK AISI × EvalEval
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Selected evaluation results now travel with verification, context and configuration details—helping readers understand what a score means and how it was produced.
“Evaluation results, with the context needed to inspect them.”
The Evaluation Cards approach01 / The release
Results linked to the conditions behind them
The UK AI Security Institute is using EvalEval’s open Evaluation Cards to publish selected findings from its paper, How Inference Compute Shapes Frontier LLM Evaluation. The cards bring results together with benchmark metadata, evaluation-run data and model information.
Five evaluation suites
The paper’s main experiment reports results on:
- HealthBench
- FrontierMath
- Humanity’s Last Exam
- SWE-Bench Pro
- Terminal-Bench 2.0
Six frontier systems
Two more evaluations
Cyber CTFs and The Last Ones are also included. They use a different, partly overlapping model set, so the six-model list above should not be assumed to apply.
02 / Why setup matters
A score reflects more than a model
Benchmark results can shift with inference-time compute and evaluation protocol. Without those conditions, similar-looking scores may describe meaningfully different runs.
Humanity’s Last Exam
AISI tracks the cumulative share of attempted tasks solved within a token count, counting each task’s earliest observed success. In runs with correctness feedback from an oracle after each attempt, models went on to solve additional tasks as token use increased.
Conceptual illustration of the reported relationship; no numerical values shown.
From result to interpretable record
Evaluation Cards are intended to help readers inspect the ingredients behind a result and compare it with other documented runs.
Result
Reported score or finding
Context
Benchmark and task details
Setup
Run and configuration data
Inspection
Compare with care
03 / Shared infrastructure
From a common schema to public records
AISI and EvalEval’s collaboration grew from a joint workshop alongside NeurIPS 2025. EvalEval says Institute feedback helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to selected public AISI methods and findings.
Shared language
EEE organizes benchmark, run and model metadata in a common format.
Documented runs
Evaluation Cards connect results with context and configuration information.
Public access
Selected AISI records are made available where appropriate.
Wider use
Developers can submit verified results and report runs through the schema.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”EvalEval Coalition
Related AISI work: OptStop on evaluation efficiency, HiBayES on statistical rigor, and standardisation in transcript analysis and capability elicitation.
Paper: “How Inference Compute Shapes Frontier LLM Evaluation” — UK AI Security Institute04 / Coverage and limits
More inspectable does not mean fully reproduced
The cards support inspection and may help researchers identify differences between runs. The announcement does not establish complete coverage or independent replication.
The release does not specify the total number of records or transcripts available.
It does not confirm every setup field is present for every benchmark.
Independent reproduction of these results is not reported.
Not every AISI evaluation or underlying transcript is claimed to be included.
The cyber model set, record-by-record release dates and disagreement process are not specified.
A common schema cannot by itself make different protocols directly comparable.
05 / What comes next
Broader adoption will shape the value
EvalEval expects to continue standardising and sharing evaluations with AISI and other organisations. Model developers can submit verified results; evaluation developers can report benchmarks and run data using EEE. Researchers in evaluation, governance and policy can explore cards by benchmark or model. Wider participation could improve cross-study comparisons if records are consistent and complete.
Useful reporting step: setup details belong with the evidence.Assessment
The strongest test will be broad, consistent coverage and whether independent researchers can reproduce results from published records. Persistent gaps in critical setup details would limit how much confidence the cards support.
06 / Key questions
What readers should know
What has AISI announced?
Selected public evaluation results are being published through EvalEval’s Evaluation Cards with verification, context and configuration information.
Which main benchmarks are included?
HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
Which models are covered?
The main experiment includes Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Cyber evaluations use a different, partly overlapping set.
Does this prove the results can be reproduced?
No independent replications are reported, and the announcement does not say every necessary detail is available for every evaluation. The records are intended to support inspection and reproduction.
Why include setup information?
Inference compute and evaluation protocol can affect scores. Documented conditions help readers interpret results and spot mismatched comparisons.
Where can the schema go next?
Broader adoption by model and evaluation developers could make comparisons easier, depending on the consistency and completeness of published records.
Setup Details Change Score Comparisons
Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis illustrates why: outcomes shifted with inference compute and with whether models received correctness feedback between attempts. A score without those conditions can leave readers unsure what performance it represents.
Publishing results with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that uses evaluations as evidence about advanced AI capabilities. The records do not by themselves settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see.
The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that shared infrastructure to publicly reported AISI methods and findings.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information. Together, the efforts address a practical reporting problem: results published across formats and outlets may omit details needed to interpret or reproduce a run, while repeating costly evaluations may not be feasible.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
Coverage and Reproduction Limits
The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. It says publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of all AISI evaluation work.
The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare the records consistently.
Broader Adoption of EEE
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema.
Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of records contributors publish. No further release date or adoption milestone was specified.
Where I land
I see the release as a useful reporting step because it puts selected AISI results alongside information readers need to understand how they were generated. In particular, the paper’s analysis shows that inference compute and feedback conditions can affect measured performance, making setup details part of the evidence rather than an administrative extra.
The strongest counterargument is that a common schema cannot make unlike evaluations directly comparable, and these records may cover only part of AISI’s work. I would judge the effort more strongly if the collection showed broad, consistent coverage and independent researchers could reproduce results from the published records. Evidence of missing critical setup details or persistent gaps across contributors would make me more cautious about how much confidence the cards support.
Key Questions
What has AISI announced?
AISI is using EvalEval’s Evaluation Cards to make selected public evaluation results available with verification, context and configuration information.
Which benchmarks are included in the main experiment?
The records cover HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
Which models are covered?
The five main-experiment benchmarks include results for Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping set of models.
Does the release prove that the results can be reproduced?
The records provide information intended to support inspection and reproduction. The announcement does not report independent replications or say that every necessary detail is available for every evaluation.
Why include evaluation setup information?
Scores can change with factors such as inference compute and evaluation protocol. Setup details help readers interpret scores and spot when comparisons involve different conditions.
Source: Hugging Face
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
