AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The UK AI Security Institute is using EvalEval’s Evaluation Cards to publish evaluation results alongside verification, context and configuration details. The release covers five benchmarks across six frontier models, plus two cyber evaluations with a different, partly overlapping model set.

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, adding verification, context and configuration details that help readers interpret how the results were produced. The release accompanies AISI’s paper on inference-time compute and evaluation protocols, and includes results for five benchmarks across six frontier models.

The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI has also shared results from two cyber evaluations, Cyber CTFs and The Last Ones. Those use a different set of models that overlaps only partly with the main experiment, so the same model list should not be assumed to apply to them.

The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation. The paper examines how scores depend on inference-time compute and evaluation protocol. For Humanity’s Last Exam, the reported analysis tracks the cumulative share of attempted tasks solved within a given token count, using each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they went on to solve additional tasks as token use increased.

EvalEval describes the released records as including verified results, evaluation context and configuration information. The platform organizes benchmark metadata, evaluation-run data and model metadata into a common format. AISI’s public reporting is available where appropriate; the announcement does not claim that every AISI evaluation or every underlying transcript is included.

At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Evaluation reporting · UK AISI × EvalEval

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Selected evaluation results now travel with verification, context and configuration details—helping readers understand what a score means and how it was produced.

“Evaluation results, with the context needed to inspect them.”

The Evaluation Cards approach
5Main benchmarks
6Frontier models
5Core benchmarks
6Models in main study
2Cyber evaluations
EEEShared reporting schema

01 / The release

Results linked to the conditions behind them

The UK AI Security Institute is using EvalEval’s open Evaluation Cards to publish selected findings from its paper, How Inference Compute Shapes Frontier LLM Evaluation. The cards bring results together with benchmark metadata, evaluation-run data and model information.

Main experiment · Benchmarks

Five evaluation suites

The paper’s main experiment reports results on:

  • HealthBench
  • FrontierMath
  • Humanity’s Last Exam
  • SWE-Bench Pro
  • Terminal-Bench 2.0
Main experiment · Models

Six frontier systems

Claude Opus 4Claude Opus 4.5Claude Opus 4.6 GPT-5GPT-5.2GPT-5.4
Additional · Cyber

Two more evaluations

Cyber CTFs and The Last Ones are also included. They use a different, partly overlapping model set, so the six-model list above should not be assumed to apply.

02 / Why setup matters

A score reflects more than a model

Benchmark results can shift with inference-time compute and evaluation protocol. Without those conditions, similar-looking scores may describe meaningfully different runs.

Humanity’s Last Exam

AISI tracks the cumulative share of attempted tasks solved within a token count, counting each task’s earliest observed success. In runs with correctness feedback from an oracle after each attempt, models went on to solve additional tasks as token use increased.

More tokens
Oracle feedback

Conceptual illustration of the reported relationship; no numerical values shown.

From result to interpretable record

Evaluation Cards are intended to help readers inspect the ingredients behind a result and compare it with other documented runs.

1

Result

Reported score or finding

2

Context

Benchmark and task details

3

Setup

Run and configuration data

4

Inspection

Compare with care

03 / Shared infrastructure

From a common schema to public records

AISI and EvalEval’s collaboration grew from a joint workshop alongside NeurIPS 2025. EvalEval says Institute feedback helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that infrastructure to selected public AISI methods and findings.

1

Shared language

EEE organizes benchmark, run and model metadata in a common format.

2

Documented runs

Evaluation Cards connect results with context and configuration information.

3

Public access

Selected AISI records are made available where appropriate.

4

Wider use

Developers can submit verified results and report runs through the schema.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
EvalEval Coalition

Related AISI work: OptStop on evaluation efficiency, HiBayES on statistical rigor, and standardisation in transcript analysis and capability elicitation.

Paper: “How Inference Compute Shapes Frontier LLM Evaluation” — UK AI Security Institute

04 / Coverage and limits

More inspectable does not mean fully reproduced

The cards support inspection and may help researchers identify differences between runs. The announcement does not establish complete coverage or independent replication.

01

The release does not specify the total number of records or transcripts available.

02

It does not confirm every setup field is present for every benchmark.

03

Independent reproduction of these results is not reported.

04

Not every AISI evaluation or underlying transcript is claimed to be included.

05

The cyber model set, record-by-record release dates and disagreement process are not specified.

06

A common schema cannot by itself make different protocols directly comparable.

05 / What comes next

Broader adoption will shape the value

EvalEval expects to continue standardising and sharing evaluations with AISI and other organisations. Model developers can submit verified results; evaluation developers can report benchmarks and run data using EEE. Researchers in evaluation, governance and policy can explore cards by benchmark or model. Wider participation could improve cross-study comparisons if records are consistent and complete.

Useful reporting step: setup details belong with the evidence.
Assessment

The strongest test will be broad, consistent coverage and whether independent researchers can reproduce results from published records. Persistent gaps in critical setup details would limit how much confidence the cards support.

06 / Key questions

What readers should know

What has AISI announced?

Selected public evaluation results are being published through EvalEval’s Evaluation Cards with verification, context and configuration information.

Which main benchmarks are included?

HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Which models are covered?

The main experiment includes Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. Cyber evaluations use a different, partly overlapping set.

Does this prove the results can be reproduced?

No independent replications are reported, and the announcement does not say every necessary detail is available for every evaluation. The records are intended to support inspection and reproduction.

Why include setup information?

Inference compute and evaluation protocol can affect scores. Documented conditions help readers interpret results and spot mismatched comparisons.

Where can the schema go next?

Broader adoption by model and evaluation developers could make comparisons easier, depending on the consistency and completeness of published records.

Setup Details Change Score Comparisons

Benchmark scores are often cited as if they measure the same thing across models, but different evaluation protocols can produce different results. The AISI paper’s Humanity’s Last Exam analysis illustrates why: outcomes shifted with inference compute and with whether models received correctness feedback between attempts. A score without those conditions can leave readers unsure what performance it represents.

Publishing results with their setup information gives researchers and practitioners a way to inspect individual evaluations and compare them with other reported runs. It can also help identify when superficially similar scores came from meaningfully different conditions. That matters for research, model development and policy work that uses evaluations as evidence about advanced AI capabilities. The records do not by themselves settle which benchmark or protocol is best, but they make some of the conditions behind a result easier to see.

From Shared Schema to Public Records

The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from the Institute helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. The current release applies that shared infrastructure to publicly reported AISI methods and findings.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in areas such as transcript analysis and capability elicitation. EvalEval’s related project, Evaluation Cards, combines evaluation results with benchmark and model information. Together, the efforts address a practical reporting problem: results published across formats and outlets may omit details needed to interpret or reproduce a run, while repeating costly evaluations may not be feasible.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Coverage and Reproduction Limits

The announcement does not specify how many records or transcripts are available, which individual setup fields are present for every benchmark, or whether outside researchers have independently reproduced the results. It says publicly reported methods and findings are being made available where appropriate, so the release should not be read as a complete archive of all AISI evaluation work.

The cyber evaluations use a different, partly overlapping model set, and the announcement does not enumerate that set. It also does not give a release date for each record or describe a process for resolving disagreements between results reported under different protocols. Those details would help readers judge the current coverage and compare the records consistently.

Broader Adoption of EEE

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmarks and run data using the schema.

Researchers in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make cross-study comparisons easier, though its value will depend on the consistency and completeness of records contributors publish. No further release date or adoption milestone was specified.

Where I land

I see the release as a useful reporting step because it puts selected AISI results alongside information readers need to understand how they were generated. In particular, the paper’s analysis shows that inference compute and feedback conditions can affect measured performance, making setup details part of the evidence rather than an administrative extra.

The strongest counterargument is that a common schema cannot make unlike evaluations directly comparable, and these records may cover only part of AISI’s work. I would judge the effort more strongly if the collection showed broad, consistent coverage and independent researchers could reproduce results from the published records. Evidence of missing critical setup details or persistent gaps across contributors would make me more cautious about how much confidence the cards support.

Key Questions

What has AISI announced?

AISI is using EvalEval’s Evaluation Cards to make selected public evaluation results available with verification, context and configuration information.

Which benchmarks are included in the main experiment?

The records cover HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Which models are covered?

The five main-experiment benchmarks include results for Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping set of models.

Does the release prove that the results can be reproduced?

The records provide information intended to support inspection and reproduction. The announcement does not report independent replications or say that every necessary detail is available for every evaluation.

Why include evaluation setup information?

Scores can change with factors such as inference compute and evaluation protocol. Setup details help readers interpret scores and spot when comparisons involve different conditions.

Source: Hugging Face

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Salesforce Launches Agentforce 360: Agents Go Enterprise-Grade

AIThis post was created with the assistance of artificial intelligence (AI).Date: October…

The real reason talented developers aren’t getting hired

AIThis post was created with the assistance of artificial intelligence (AI). Buying…

HBM Ate the Fab

AIThis post was created with the assistance of artificial intelligence (AI).Part 2…

Free White Paper: NVIDIA Alpamayo and the New Era of Reasoning-Based Autonomy

AIThis post was created with the assistance of artificial intelligence (AI).How “open”…