Why a score can mislead

Benchmark saturation and Goodhart's law: when a score stops separating models

A benchmark is saturated when every strong model scores so close to the top that the gaps between them are smaller than the benchmark's own errors or its margin of error. Stanford's AI Index 2026 reports the top 15 models on the MMLU-Pro knowledge test all above 87%, with just over 4 percentage points between first and fifteenth. Goodhart's law, in Marilyn Strathern's wording, explains the pattern: "When a measure becomes a target, it ceases to be a good measure." The same law applies to the eval set you build for your own product, so keep part of it unseen by the people who tune against it and add new cases on a schedule.

Published September 30, 2026. Editorial.

Key takeaways

  • The AI Index 2026 defines benchmark saturation as the state where models reach scores so high that a test can no longer distinguish between them.
  • On MMLU, models score over 90% while one study estimates that 6.49% of the questions contain errors, so the flawed share is larger than half of the distance left to a perfect score.
  • On a 500-task benchmark a score of 76.8% has a 95% margin of 3.70 points either side, so two models 3 points apart are not separated by that benchmark alone.
  • Goodhart's law in its short form, "When a measure becomes a target, it ceases to be a good measure", is Marilyn Strathern's 1997 restatement of an idea from the economist Charles Goodhart.
  • Your own eval set can be overfitted too: split it into a working part and a held-out part (cases the tuners never read) before tuning starts, and add new cases from real use on a schedule.

When the GLUE language benchmark was new, solving it was judged to be beyond what the methods of the time could do. The Dynabench paper of 2021 records what followed: "GLUE saturated within a year and its successor, SuperGLUE, already has models rather than humans at the top of its leaderboard" [2]. Saturated means that models had reached scores too high for the test to separate them. A leaderboard is the public table that ranks models by score.

Stanford's AI Index 2026 reports that the cycle has shortened: "Evaluations intended to be challenging for years are saturated in months, compressing the window in which benchmarks remain useful for tracking progress" [1]. This page explains the two ideas behind that sentence. Saturation is the state in which every strong model scores near the top of a test. Goodhart's law explains why a published score tends to reach that state. Both apply to the eval set you build for your own product, and the last sections say what to do about it.

What benchmark saturation means

The AI Index gives the definition: "Benchmark saturation, where models reach scores so high that a test can no longer distinguish between them, remains a concern" [1].

A benchmark compares models well while their scores are spread out. Once the strong models all score within a few points of each other, the gaps that remain are small enough to come from two causes unrelated to skill. One is mistakes in the benchmark's own questions and answers. The other is chance, because a benchmark is a limited sample of questions. A gap that either cause could produce is no evidence that one model is better than another. The next two sections take those causes in turn, and what an AI benchmark measures covers how a score is produced in the first place.

When the gap is smaller than the benchmark's own errors

MMLU is a multiple-choice knowledge test from 2020, with 15,908 questions across 57 subjects [5]. The paper that introduced a harder test, Humanity's Last Exam, states that language models "now achieve over 90% accuracy on popular benchmarks like MMLU" [3].

In 2024 Gema and 15 co-authors re-checked 5,700 MMLU questions by hand, 100 from each subject. Their conclusion: "We estimate that 6.49% of MMLU questions contain errors" [4]. In the Virology subject, 57% of the questions they analysed contained errors [4].

Put those two figures side by side. A model above 90% has under 10 points left to gain. An estimated 6.49 points of the test are questions with errors, and 6.49 divided by 10 is 0.649, so the flawed share is larger than half of the distance left to a perfect score. When an answer key is wrong, a model that answers correctly is marked wrong. Among models above 90%, a difference of one or two points can therefore reflect which model agrees with a faulty key. The 6.49% is an estimate from a sample: 5,700 of 15,908 questions is 35.8% of the test.

The AI Index reports a wider review of nine benchmarks, which it credits to Truong and co-authors (2025): "A review found invalid question rates ranging from 2% on MMLU Math to 42% on GSM8K" [1]. GSM8K is a set of grade school maths problems. Those two numbers need a caution. The chart that shows them in the report labels its axis "Precision@50", which measures the questions that the review method flagged. An invalid rate would measure the whole benchmark. We quote the report's sentence as written and treat the question of what the 42% counts as unsettled.

When the gap is smaller than the margin of error

A benchmark is a sample of questions, so every score has a margin of error that depends on how many questions there are. Evan Miller's paper on reporting eval results, written at Anthropic, recommends reporting a standard error with every score [11]. A standard error is a number that says how far a score would be expected to move if the test were run on a different sample of questions.

SWE-bench Verified is a coding benchmark of 500 tasks [9]. The AI Index says that in February 2026 the leading model was at 76.8% (an approximate figure in the report), with several others "grouped between 70% and 76%" [1]. Here is the simplest version of the calculation for the leader's score, treating the 500 tasks as independent and using one run:

Step Working Result
Score as a fraction 76.8 divided by 100 0.768
Standard error Square root of (0.768 x 0.232 / 500) 0.0189, which is 1.89 points
95% margin, the range expected to hold the result in 95 of 100 repeats 1.96 x 1.89 3.70 points
Interval 76.8 minus 3.70, 76.8 plus 3.70 73.1% to 80.5%

A model at 74% and a model at 76.8% are inside the same interval, so their order could change on a different sample of tasks. How many test cases an LLM eval needs explains the method.

The sources also disagree with each other at this level. OpenAI, writing in the same month, gave the best published score on SWE-bench Verified as 80.9% [8]. That is 4.1 points above the AI Index figure. This page reports both and does not choose between them.

Benchmarks where the sources report these signs

Benchmark Released What the source reports
GLUE Year not given in the paper read It "saturated within a year". By 2021 its successor, SuperGLUE, had models above humans at the top [2]
MMLU 2020 [5] Models score over 90% [3]. An estimated 6.49% of questions contain errors [4]
MMLU-Pro 2024 [1] As of early 2026 the top 15 models all score above 87%, and first to fifteenth are "just over 4 percentage points" apart [1]
SWE-bench Verified 13 August 2024 [9] By OpenAI's account the top score rose 6.0 points in 6 months, from 74.9% to 80.9%. OpenAI stopped reporting it on 23 February 2026 [8]
Humanity's Last Exam 2025 [3] Top accuracy "went from under 10% to 38.3%" in a single year [1]

The AI Index text gives no table of how many years each benchmark took to saturate, and this page does not estimate one.

The same pattern shows in rankings built from votes. Arena is a site that ranks models by people's votes between pairs of answers. On its leaderboard as of March 2026, the AI Index lists six companies in the top tier with ratings from 1,503 down to 1,424, and the top four are within 22 points of each other (1,503 minus 1,481) [1]. The report says the narrow gaps are "shifting competitive pressure toward cost, reliability, and domain-specific performance" [1].

What Goodhart's law says and where the wording comes from

Goodhart's law is named after the economist Charles Goodhart, who was Chief Adviser to the Bank of England. A University of Cambridge page gives the statement, from page 96 of the book "Monetary Theory and Practice": "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes" [6].

The shorter sentence that is usually quoted has a different author. The same Cambridge page says that Professor Marilyn Strathern "re-stated Goodhart's Law more succinctly and more generally": "When a measure becomes a target, it ceases to be a good measure" [6]. The page places the sentence in Strathern's 1997 article "'Improving ratings': audit in the British University system", in the journal European Review, volume 5, pages 305 to 321 [7].

We read both wordings on the Cambridge page. Strathern's article is behind a paywall. Its public abstract describes the subject, which is the growth of procedures for evaluating performance, and does not contain the sentence [7]. Goodhart's original text was not opened for this page.

In plain words: a number works as a measure while nobody is rewarded for moving it. Once people are rewarded for the number, they find ways to raise it that leave the underlying quality unchanged. The number then tells you less about the quality.

How Goodhart's law applies to AI benchmarks

Benchmark scores appear in model announcements and sales material, so for a model vendor the score is a target. The sources describe three ways a score can rise while the skill it is named after stays the same.

Exposure during training. OpenAI wrote on 23 February 2026 that gains on SWE-bench Verified "increasingly reflect how much the model was exposed to the benchmark at training time", and added: "we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too" [8]. This is a dated case of a vendor that stopped reporting a benchmark it helped to build. It rests on OpenAI's own audit. Benchmark contamination covers the evidence.

Building systems against the test. An agent is a program in which a model works in steps and uses tools. Kapoor and co-authors at Princeton found that "many agent benchmarks have inadequate holdout sets, and sometimes none at all", and that agents built against them "overfit to the benchmark in various ways" [10]. A holdout set is a group of test items kept back and never used while the system is built. To overfit is to score well on the items seen and worse on new ones.

Adapting to the ranking platform. The AI Index, citing Singh and co-authors (2025), says that standing on the Arena leaderboard "may partly reflect adaptation to the platform rather than general capability alone" [1]. The platform's operator disputes several of that study's claims, and both sides are set out in how to read a model vendor's benchmark claims.

The AI Index states the result in one line: "strong benchmark performance does not always translate to real-world utility" [1].

Why new benchmarks keep appearing and how long one lasts

In the sources read for this page, each saturated benchmark was followed by a harder one. GLUE was followed by SuperGLUE [2]. MMLU was followed by MMLU-Pro in 2024 [1] and by Humanity's Last Exam, a set of 2,500 questions whose authors say that scores over 90% on MMLU limit how well the abilities of the leading large language models can be measured [3]. Scores on the successor then rise as well: the top result on Humanity's Last Exam more than tripled in one year, from under 10% to 38.3% [1].

The sources give no fixed period after which a benchmark stops being useful. Two dated cases: GLUE saturated within a year [2], and SWE-bench Verified ran for 18 months and 10 days between its release on 13 August 2024 [9] and OpenAI's decision to stop reporting it [8]. For your own decisions, stop relying on a benchmark when one of three things is true:

  1. The models you are comparing are inside each other's margin of error.
  2. The questions were public before the models' training cut-off dates, the dates after which a model has no training text.
  3. An audit has found errors in the benchmark as large as the gaps between the models.

The same law applies to your own eval set

An eval set built from your own cases can be overfitted in the same way. Each time a team changes a prompt (the written instructions given to the model), runs the set again and keeps the change because the score went up, the prompt is fitted more closely to those specific cases.

An invented illustration: a team tunes a support assistant's prompt against 200 saved cases and reaches 96% on them. On 100 other cases that the team never looked at, the same prompt scores 81%. The 15-point gap shows that part of the 96% came from fitting the 200 cases.

Four practices keep your set useful as a measure:

  1. Split the set before tuning starts. Keep a working part for tuning and a held-out part that stays unread by the people who change the prompt.
  2. Run the held-out part only at decision points, such as before a release, and record both scores. A growing gap between them is the sign of overfitting.
  3. Refresh on a schedule. Add new cases from real use, and retire cases that every version passes. How to build an LLM eval dataset gives the method.
  4. Read the failures as well as the average. A score can rise while one type of failure gets worse.

When evals give false confidence covers this for evals of AI-written code, how many tests AI-generated code needs covers the size of a suite, and choosing a model with your own evals uses the same set to compare models. The hub page puts saturation next to the other ways a score misleads.

How Reveneau applies this

Reveneau is an AI software development consultancy, and Goodhart's law applies directly to our method. All of our code is written by AI, and every change must pass an eval suite before release. For the AI that writes the code, that suite is a target.

We handle this in the way this page recommends. We write the suite from the client's specification before any code exists, so the checks describe what was asked for and cannot be shaped around code that is already there. When the specification changes, we write new checks from it. For an AI feature, we keep part of the case set unseen while prompts are tuned. The judged checks in the suite are graded by Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader.

Reveneau, as a company, takes responsibility for the whole project through production and after release. To discuss an eval set for your product, contact us or read about AI development at Reveneau.

Common questions

What is benchmark saturation?

Benchmark saturation is the state in which the strong models all score so close to the top of a test that the test can no longer tell them apart. Stanford's AI Index 2026 gives MMLU-Pro as an example: as of early 2026 the top 15 models all score above 87%, and first and fifteenth are just over 4 percentage points apart. Gaps that small can come from errors in the test or from chance.

What is Goodhart's law?

Goodhart's law says that a number stops being a reliable measure once people are rewarded for moving it. A University of Cambridge page gives the economist Charles Goodhart's own statement: any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes. In AI, a benchmark score that appears in model announcements is a number under that pressure.

Who wrote the sentence about a measure becoming a target?

Marilyn Strathern wrote the short form, "When a measure becomes a target, it ceases to be a good measure", as a restatement of Charles Goodhart's idea. A University of Cambridge page credits it to Strathern's 1997 article on audit in British universities, in European Review, volume 5. We read the wording on that Cambridge page, because the article itself is behind a paywall and its public abstract does not include the sentence.

Why do new AI benchmarks keep appearing?

New AI benchmarks keep appearing because the old ones stop separating the leading models. The authors of Humanity's Last Exam built its 2,500 questions because models score over 90% on benchmarks like MMLU. The Dynabench paper records that GLUE saturated within a year and was followed by SuperGLUE. Scores on each successor then rise too: the AI Index reports Humanity's Last Exam going from under 10% to 38.3% in one year.

How often should a benchmark be replaced?

A benchmark should be replaced when it stops separating the models being compared, and the sources give no fixed period for that. GLUE saturated within a year, according to the Dynabench paper. SWE-bench Verified ran for 18 months and 10 days from its release on 13 August 2024 until OpenAI stopped reporting it on 23 February 2026. Check the gaps against the margin of error instead of counting months.

Can my own eval set be overfitted?

Your own eval set can be overfitted whenever the same cases are used both to tune a prompt and to judge it. Each change that is kept because the score went up fits the prompt more closely to those cases. Kapoor and co-authors at Princeton found the same fault in public agent benchmarks, many of which keep too few test items back, and some of which keep none.

How do I stop my team overfitting to our eval set?

Stop overfitting by splitting the eval set before any tuning starts. The team tunes against a working part and never reads the held-out part, which is run only at decision points such as a release. Record both scores each time. In the invented illustration on this page, 96% on the working cases against 81% on the held-out cases shows a 15-point gap that came from fitting.

How close can two benchmark scores be before the gap stops meaning anything?

Two benchmark scores stop being separable when they are inside each other's margin of error, which depends on the number of questions. On a 500-task benchmark such as SWE-bench Verified, a score of 76.8% has a standard error of 1.89 points and a 95% margin of 3.70 points either side, in the simplest calculation. A model at 74% is inside that interval of 73.1% to 80.5%.

Is a saturated benchmark still useful for anything?

A saturated benchmark, one where the strong models all score near the top, is still useful as a minimum check: a model that scores far below the group on it has a problem worth looking into. A saturated benchmark is weak evidence for ranking the models at the top. The AI Index says the narrow gaps between leading models are moving competition toward cost, reliability and performance in a specific field, which are things you can measure on your own cases.

Has a vendor ever stopped reporting a benchmark it used to quote?

OpenAI stopped reporting SWE-bench Verified on 23 February 2026, a benchmark it had helped to build, and recommended that other model developers stop too. By OpenAI's own account the top score had risen only from 74.9% to 80.9% in six months, and gains increasingly reflected how much a model had been exposed to the benchmark during training. The decision rests on OpenAI's own audit.

How much work is it to keep an eval set useful over time?

Keeping an eval set useful takes a recurring task with an owner: add new cases from real use on a fixed schedule and retire the cases every version passes. The work per case is writing down the correct answer. The cost of skipping it is the one Kapoor and co-authors describe for public agent benchmarks, where systems overfit to a test with too few items kept back.

Does benchmark saturation matter if I only buy a model and never train one?

Benchmark saturation matters to a buyer because it removes the public score as a way to choose between the leading models. On Arena's public ranking table in March 2026 the AI Index lists the top four companies within 22 rating points of each other. When public scores are that close, the model that is best for you is decided by cost, speed and results on your own cases.

References