Start here

What an AI benchmark measures and what it leaves out

An AI benchmark is a public, fixed set of questions or tasks, with a scoring rule, that many models are run on so their results can be listed in one table. A score tells you how a model did on that test, under the settings chosen by whoever ran it. It leaves out your documents, your users' wording, your rules, your costs and your response times. A 2025 review of 445 benchmarks found that fewer than 1 in 10 used complete real-world tasks, and that 16.0% used statistics to compare results. Use a benchmark to pick which models to try, then decide with an eval built from your own cases.

Published September 30, 2026. Editorial.

Key takeaways

  • An AI benchmark has three parts: a fixed set of items chosen by its authors, a rule for marking each answer, and a public table of model scores called a leaderboard.
  • A 2025 review of 445 LLM benchmarks found that less than 10% used complete real-world tasks and that 16.0% used uncertainty estimates or statistical tests to compare results.
  • The BetterBench paper scored 24 benchmarks against 46 practices: 14 of the 24 did not repeat runs or report uncertainty, and MMLU had the lowest weighted average at 5.5.
  • A benchmark is written once for all models with public questions. Your own eval is written for one product from its real cases, and the cases stay private.
  • Use benchmark scores to choose two to four models to try, then run those models on your own cases before you decide.

In 2020 seven researchers published a test called MMLU: 15,908 multiple-choice questions across 57 subjects, from elementary mathematics to law [1]. Each question has four options, so guessing scores 25%. People recruited through Amazon Mechanical Turk, a website where people are paid to do small online tasks, with no special training, scored 34.5% [1]. The authors of a newer test, first published in January 2025, write that large language models (LLMs, the AI models behind chat assistants) "now achieve over 90% accuracy on popular benchmarks like MMLU" [2].

A line such as "90% on MMLU" in a model announcement is a true statement about one set of exam questions. Whether that model will answer your customers correctly from your own documents is a second question, and the test was built before your product existed. This page explains what a benchmark is, how its score is made, and where the score stops being useful to a buyer. It is the first page of our guide to AI benchmarks vs your own evals.

What an AI benchmark is

An AI benchmark is a public, fixed set of questions or tasks that many models are scored on, so that their results can be listed in one table. It has three parts.

  • The items. The questions or tasks, chosen once by the benchmark's authors. MMLU's test set has 14,079 questions [1]. The original SWE-bench, a coding benchmark, has 2,294 tasks taken from real problem reports in 12 software projects written in the Python programming language [3].
  • The scoring rule. A fixed way to decide whether each answer counts. For MMLU the rule is whether the model chose the right option out of four. For SWE-bench the project site reports "% Resolved", which it defines as "the percentage of task instances solved" [3].
  • The published table. A public ranking of models by score, called a leaderboard.

Everyone runs the same items under the same rule, so two models can be compared directly. The cost of that design is the subject of the sections below: the items are fixed, they are public, and they were chosen by people who have never seen your product.

Who creates AI benchmarks

The benchmarks named on this page come from three kinds of author.

  • Research groups. MMLU was written by seven researchers and presented at a machine learning conference in 2021 [1].
  • Organisations and companies that work on AI testing. Humanity's Last Exam, a 2,500-question test, lists the Center for AI Safety and Scale AI as its authors' affiliations [2].
  • Model vendors. The SWE-bench project site records, in a news item dated August 2024, that its human-filtered version called Verified was made with OpenAI, a company whose models are scored on it [3].

The maker decides what goes into the test, so find out who built a benchmark and whether they also sell a model or a testing service.

Researchers at Stanford scored 24 AI benchmarks against 46 good practices in a paper called BetterBench, accepted at the NeurIPS 2024 conference [4]. They found that "14 out of 24 benchmarks did not perform multiple evaluations of the same model or report statistical significance or uncertainty of results". They also found that 17 of the 24 gave no easy script for repeating the results in the benchmark's first paper [4]. MMLU received the lowest weighted average in their assessment, 5.5. A science benchmark called GPQA received 11.0, where the highest possible score per category is 15 [4]. Those figures rate how each benchmark was designed and documented, by the Stanford team's own criteria.

How a benchmark score is produced

A score starts when someone runs a model on every item and applies the scoring rule. Often that is the company that built the model. They choose how many attempts the model gets, whether it may use tools such as web search, and which program runs the model through the task. Each choice can move the number, and how to read a model vendor's benchmark claims lists the questions to ask about each one.

Two more things affect the number. First, a score comes from a sample of questions, so it has a margin of error, and most benchmark papers leave it out. A review of 445 LLM benchmarks by 29 expert reviewers, presented at the NeurIPS 2025 conference, found that "16.0% used uncertainty estimates or statistical tests to compare the results" [5]. When no margin of error is published, a reader cannot tell whether a gap of one or two points between two models is larger than the change from running the same model twice.

Second, the benchmark's own answer sheet can contain mistakes. MMLU, GPQA and ARC-AGI explained gives the published error counts for the knowledge tests, and SWE-bench explained does the same for the coding test.

The gap between the test and the skill in its name

A benchmark has a name such as "language understanding" or "software engineering", and a score on it is easy to read as a score for the whole skill. A 2021 paper by Raji and four co-authors, accepted at NeurIPS, argues that this reading claims more than the test supports [6]. A benchmark is a finite and specific set of data, they write, and "the claims that are justified through these benchmark datasets extend far beyond the tasks they are initially designed for" [6]. Their paper is an argument about two older benchmarks and reports no experiment.

The 2025 review of 445 benchmarks measured the same concern [5]. Researchers call the concern construct validity: whether a test measures the thing it is named after, or in the review's words, "having measures that represent what matters to the phenomenon". The reviewers reported the following.

What the reviewers checked Share found What it means for a reader
The paper defined what the benchmark measures 78.2% of papers The other papers name a skill and leave it undefined
The definition is contested 47.8% of definitions given Experts disagree on what the skill is
The paper gave evidence that the benchmark measures what it claims 53.4% of papers The other papers assert the claim
The benchmark reused data that was easy to get 27.0% of benchmarks Questions were picked for convenience
The benchmark used complete real-world tasks Less than 10% of benchmarks Most tests use short questions in place of whole jobs

These percentages are the reviewers' own classification of benchmark papers from academic conferences [5]. The last row is the one a buyer should remember.

What a benchmark leaves out

A public test is written for every model and every reader, so it leaves out what is particular to you. Five such things decide whether a model works in a product.

  1. Your documents and data. The model has to answer from your price list, your policies and your records.
  2. Your users' wording. Real questions are short, misspelled and missing context. Exam questions are complete.
  3. Your definition of a correct answer. A reply can be accurate and still break a rule of yours, such as promising a refund that your policy does not allow.
  4. Cost and response time. Two models with the same score can differ in price per task and in how long the user waits.
  5. What happens when the model is replaced. A benchmark score is for one version on one date.

Researchers at Princeton made the fourth point in their 2024 paper AI Agents That Matter. They describe "a narrow focus on accuracy without attention to other metrics" in tests of AI agents, which are programs in which a model works in steps and uses tools. They also write that "the benchmarking needs of model and downstream developers have been conflated" [7], meaning treated as one. A downstream developer is a team that builds a product using someone else's model, which is the position of most buyers. Stanford's AI Index 2026 report puts the general point in one sentence: "strong benchmark performance does not always translate to real-world utility" [8].

How a benchmark differs from your own eval

An eval is a test of an AI system: a set of inputs, the system's outputs, and a rule that marks each output. A benchmark is an eval written once for all models. Your own eval is written for one product, from that product's real cases, with your definition of a correct answer. What is an AI eval explains the idea from the beginning.

OpenAI's documentation makes the same distinction. It names "Industry benchmarks for comparing models in isolation" as one type of evaluation and "Specific tests you implement to measure your LLM application's performance" as another [9]. Its instruction for the second type is "Make tests reflect model capability in real-world distributions" [9]. In plain words, build the tests from the mix of inputs your product receives in real use. That is a model vendor telling its own customers to test on their own cases.

Question Public benchmark Your own eval
Who chose the questions The benchmark's authors, once, for all models Your team, from your product's real cases
Are the questions public In most cases yes, so they can end up in the text a model is trained on They stay with you, so no model has seen them
What the score tells you How a model did on that test, under the settings the runner chose How a model did on your task, with your instructions and your data
When it goes out of date When the leading models all score near the maximum When your product or your users change, and you add cases

Can you trust a benchmark score

Trust a benchmark score as a measurement of that test, on that date, with those settings. Three known problems limit what the number can tell you beyond that.

  1. The questions may have been in the model's training text. Benchmark contamination covers the evidence.
  2. Every strong model may score near the top. The AI Index calls this saturation, "where models reach scores so high that a test can no longer distinguish between them" [8]. Its example is MMLU-Pro, a harder version of MMLU released in 2024: as of early 2026 the 15 leading models all scored above 87% [8]. Benchmark saturation and Goodhart's law explains why this keeps happening.
  3. The company that ran the test chose the settings. The AI Index notes that "third-party evaluations have documented cases where models perform more poorly in independent testing compared to developer-reported results" [8].

Why vendors quote benchmarks

A benchmark is the one number that exists for every model on the day it is released, so it is the number that goes into the announcement. The BetterBench authors observed that model announcements report MMLU and GPQA results side by side "without explicitly acknowledging their limitations or quality differences" [4].

The benchmarks in those announcements change as older tests stop separating the models. The AI Index reports that top accuracy on Humanity's Last Exam went "from under 10% to 38.3%" in a single year [8]. Go to the benchmark's own paper for the basic facts: the AI Index gives that exam 2,700 questions, and the exam's paper says 2,500 [2] [8].

How to use a benchmark number

Use these five steps each time you are shown a score.

  1. Open the benchmark's own paper or site and write down three things: how many items it has, what an item is, and how an answer is marked.
  2. Ask who ran the model and with which settings.
  3. Ask how close the items are to your task. A coding score says little about a support assistant.
  4. Use the scores to pick two to four models to try.
  5. Run those models on your own cases. How to choose an AI model with your own evals is the method.

If the model will act through tools instead of only writing text, read agent benchmarks explained first. If a supplier shows you its own test results in place of a public benchmark, how to evaluate a vendor's eval suite gives the checks for that case.

Where Reveneau fits

At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. We apply the same order to model choice. A public benchmark score has one use in our work: it helps decide which models to try. The decision itself comes from evals written for the client's product, with the client's definition of a correct answer, and those evals stay in the project so they can be run again when a vendor releases a new model. Our guide to eval-driven development gives the method for code.

We grade the judged checks in our suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. If you are choosing a model for a product, our AI development service starts with the evals.

Common questions

What is an AI benchmark in plain words?

An AI benchmark is a public, fixed set of questions or tasks, with a rule for marking answers, that many models are scored on so the results can be listed in one table. MMLU is an example: its paper describes 15,908 multiple-choice questions across 57 subjects, each with four options, so a model that guesses scores 25%.

How is a public AI benchmark different from an eval I write for my own product?

A public AI benchmark is written once by outside authors for all models, and its questions are usually published. An eval for your own product is written by your team from your real cases, with your definition of a correct answer, and the cases stay private. OpenAI's documentation separates the two in the same way, listing industry benchmarks and tests you implement for your application as different types.

Who creates AI benchmarks?

AI benchmarks are created by research groups, by organisations and companies that work on AI testing, and by model vendors. Humanity's Last Exam lists the Center for AI Safety and Scale AI as its authors' affiliations, and the SWE-bench project site records that its Verified version was made with OpenAI. Check who built a benchmark and whether they also sell a model.

Can I trust the benchmark scores in a model announcement?

You can trust a benchmark score as a measurement of that one test, on that date, with the settings the runner chose. Stanford's AI Index 2026 notes that third-party evaluations have documented cases where models performed more poorly in independent testing than in developer-reported results. Treat an announced score as a claim to check, and decide with a test on your own cases.

Why do AI vendors quote benchmark scores?

AI vendors quote benchmark scores because a benchmark is the one number that exists for every model on the day it is released, so buyers can compare models. The BetterBench paper observed that announcements report MMLU and GPQA results side by side without acknowledging their limitations or quality differences. A quoted score is useful for choosing which models to try.

How is a benchmark score calculated?

A benchmark score is calculated by running the model on every item and applying the benchmark's marking rule, then reporting the share that passed. For MMLU the rule is whether the model picked the right option out of four. For SWE-bench the project site reports the percentage of task instances solved. The party running the test chooses settings such as attempts and tools.

What does a benchmark score leave out of a buying decision?

A benchmark score leaves out everything particular to your product: your documents, your users' wording, your rules for a correct answer, cost per task, response time, and behaviour after the model is replaced. A 2025 review of 445 LLM benchmarks found that less than 10% used complete real-world tasks, so most scores come from short questions in place of whole jobs.

How much does a gap of one or two points between two models mean?

A gap of one or two benchmark points between two models often cannot be interpreted, because most benchmarks publish no margin of error. The 2025 review of 445 benchmarks found that 16.0% used uncertainty estimates or statistical tests to compare results. Without that figure, a reader cannot tell whether the gap is larger than the change from running the same model twice.

What is construct validity in an AI benchmark?

Construct validity is whether a benchmark measures the skill it is named after. The 2025 review of 445 benchmarks describes it as having measures that represent what matters to the phenomenon. The reviewers found that 53.4% of papers presented evidence for the construct validity of their benchmark, and that 47.8% of the definitions papers gave for the measured skill are contested.

How reliable are AI benchmarks themselves?

The reliability of AI benchmarks differs from one benchmark to the next. Stanford researchers scored 24 benchmarks against 46 practices in the BetterBench paper. They found that 14 of the 24 did not run multiple evaluations or report uncertainty, and that 17 gave no easy script to repeat the first published results. MMLU had the lowest weighted average, 5.5, and GPQA scored 11.0.

Do benchmarks matter if I am buying a finished AI product and not a model?

Benchmarks matter less for a finished AI product than for a model, because the product adds its own instructions, data and rules to the model. The Princeton paper AI Agents That Matter says the needs of model developers and downstream developers have been treated as one in benchmarking. Ask the supplier for results on cases like yours instead of the model's public scores.

What should I do after I see a benchmark score?

After you see a benchmark score, open the benchmark's own paper and note how many items it has, what an item is and how answers are marked. Then ask who ran the model and with which settings. The AI Index 2026 gives 2,700 questions for Humanity's Last Exam while the exam's paper gives 2,500, so check basic facts at the source before relying on them.

References