AI NewsModels & agentsAnnouncement

UK AI Security Institute publishes benchmark cards for five frontier tests, in a common schema anyone can add to

The UK AI Security Institute published verified evaluation data for six frontier models on five benchmarks through Hugging Face's EvalEval platform on 22 September, in a shared Every Eval Ever schema anyone can add to.

AI News

Editorial3 min read

LinkedInX

Why it mattersA team weighing model choice by benchmark score has been reading numbers that were produced under different conditions and reported without the conditions, and a shared card format is the piece that lets two numbers actually be compared.

Two benchmark scores from two different labs, on the same model, on the same test, often disagree, because the labs ran the test with different inference-time compute, different sampling settings, or a different version of the benchmark. On 22 September the UK AI Security Institute (AISI) put its verified evaluation results for six frontier models onto the EvalEval platform, in a shared schema anyone else can publish results in, so a reader can see which conditions produced a number.

What was published

AISI released Evaluation Cards for six models on five benchmarks: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4, against HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0. It also released results on two cyber evaluations, Cyber CTFs and The Last Ones, on a partly overlapping set of models. The cards live at evalcards.evalevalai.com and are the data behind an AISI paper titled "How Inference Compute Shapes Frontier LLM Evaluation".

What an Evaluation Card actually contains

The EvalEval Coalition writes that its mission is "to improve this ecosystem through a shared reporting schema, Every Eval Ever, and an open platform, Evaluation Cards, that brings evaluation results and the information needed to interpret them into a common structure." A card carries three parts: the benchmark metadata (which version, which tasks, which scoring rule), the evaluation-run data (the model output, the seed, the sampling parameters, the inference-compute budget), and the model metadata. A reader who sees two scores that disagree can open both cards and find the setting that changed.

The schema, and why it exists

Every Eval Ever, or EEE, is a schema and a public repository at github.com/evaleval/every_eval_ever. Any lab or vendor can format its own results the same way and publish them, so the cards do not depend on AISI, EvalEval, or Hugging Face running every test itself. AISI is the first named contributor with production evaluations at this scale, and the release is meant to encourage a second contributor to add cards in the same shape.

What the AISI paper says the cards let you see

The paper, "How Inference Compute Shapes Frontier LLM Evaluation", is on arXiv at 2606.17930. Its argument is that reasoning-model scores rise or fall with the compute the evaluator lets them use, and a benchmark result quoted without that budget is a claim without a unit. A card records the compute budget so two runs against the same benchmark can be compared honestly, and two labs quoting different scores can trace the disagreement to a specific number rather than argue about it.

Public benchmarks decide procurement in a lot of places, and until now a vendor score has been a claim with the settings left out. AISI publishing cards for six real models on five real benchmarks, in a schema built to accept more contributors, gives a reader a way to check the settings before quoting the number. The schema is public and the platform is open, so the cost for other labs of joining is low, and the cost of staying out is that a card-carrying score from someone else can be checked while their own cannot.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX