AI benchmarks vs your own evals: how to read a model score / Why a score can mislead
Benchmark contamination: when the test questions were in the training data
Benchmark contamination happens when a benchmark's public questions and answers end up in the text a model was trained on, so the model can score well by repeating what it has read. The measured effect is real and uneven. Scale AI's GSM1k study found accuracy drops of up to 8% on 1,205 newly written maths problems for some model families, and little to no sign of the problem in Gemini, GPT and Claude models. On SWE-bench, one study saw a score fall from 12.47% to 3.97% after leaked and weakly tested tasks were removed. A buyer protects against contamination by testing on private cases from their own business, which no model has read.
Published September 30, 2026. Editorial.
Key takeaways
- Contamination means test questions or their answers reached a model before the test, usually through public text collected for training. The score then partly measures memory.
- GSM1k (Scale AI, 2024) scored models on 1,205 new maths problems and found accuracy drops of up to 8% for some model families, while Gemini, GPT and Claude models showed little to no sign of overfitting, meaning a higher score on seen questions.
- On SWE-bench, the SWE-Bench+ study found the solution inside the issue text for 32.67% of the fixes that one test setup passed, and that setup's score fell from 12.47% to 3.97% once flawed tasks were removed.
- LiveBench limits contamination by adding and updating questions monthly and keeping 1 in 6 questions unpublished at any time.
- Cases from your own business that were never published are the one test set you can be sure no model was trained on.
On 23 February 2026 OpenAI said it had stopped reporting scores on SWE-bench Verified, a coding benchmark it helped to create [7]. One reason it gave was that every leading model it tested could reproduce, for certain tasks, the original human-written fix or exact details of the task text. OpenAI's conclusion was that gains on the benchmark "increasingly reflect how much the model was exposed to the benchmark at training time" [7].
The name for that problem is benchmark contamination. This page explains what it is, sets out what six studies measured, and says what a buyer can do about it. Our position: a score on public questions is weaker evidence than a score on questions the model could never have read, and the questions from your own business are the easiest ones of that kind to get.
What benchmark contamination means
A benchmark is a fixed, public set of questions or tasks with known answers, and many models are scored on it. A large language model (LLM) learns from a large collection of text gathered up to a certain date. That collection is called the training data, and the date is called the training cut-off. If a benchmark's questions and answers were on the public internet before that date, they can be inside the collection. The model can then answer a test question by repeating text it has already read.
A survey by Xu and co-authors at University College Dublin defines the problem this way: it occurs when language models "inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance during the evaluation phase" [1]. The survey names it Benchmark Data Contamination. Other papers call the same thing data leakage [5] or test set contamination [8].
The authors of the GSM1k study use a wider definition: "data closely resembling benchmark questions leaks into the training data" [2]. Under that definition a reworded copy of a question, with the same answer, counts too.
A second type of leak happens inside the test itself: the answer is written in the question the model is given. The SWE-bench findings below include both types.
Why a public benchmark is exposed to this
A benchmark has to be published so that other people can run it, and publication puts it where training text is collected. SWE-bench is built from issues in public code projects (an issue is a written report of a fault or a request). One study of it found that "over 94% of the issues were created before LLM's knowledge cutoff dates" [5].
A buyer cannot check this directly when a vendor keeps a model's training data private. Deng and co-authors at Yale wrote that the problem "is especially critical for closed-source models and certain open-source models where training data transparency is lacking" [4]. The evidence therefore comes from experiments that test a model without access to its training data.
What six studies measured
| Study | What was done | What was found |
|---|---|---|
| GSM1k, Zhang and co-authors at Scale AI, 2024 [2] | People wrote 1,205 new maths problems and each model was scored on the new set and on the public set, GSM8k | Accuracy fell by up to 8% for some model families. The leading models showed "minimal signs of overfitting" |
| GSM-Symbolic, Mirzadeh and co-authors, 2024 [3] | The same questions were rebuilt as 100 templates with 50 versions each, so names and numbers could change | Every model scored lower when only the numbers changed |
| Deng and co-authors, 2023 [4] | One wrong answer option was hidden in questions from MMLU, a multiple-choice knowledge benchmark, and the model was asked to fill it in | ChatGPT matched the hidden option exactly 52% of the time and GPT-4 57% |
| SWE-Bench+, Aleithan and co-authors, 2024 [5] | The authors read by hand the tasks that one setup, SWE-Agent with GPT-4, had passed | 32.67% of the passing fixes had the solution in the issue text. The score fell from 12.47% to 3.97% |
| The SWE-Bench Illusion, Liang, Garg and Moghaddam, 2025 [6] | Models were asked to name the faulty file from the issue text alone | Up to 76% correct on SWE-bench tasks and up to 53% on tasks from other code projects |
| OpenAI's audit, 23 February 2026 [7] | GPT-5 was used to probe three models for memorised material | All three could reproduce the reference fix or exact task details for certain tasks |
For studies [4], [5] and [6], this page takes its figures from each paper's abstract. Every "up to" figure is the best or worst case across the models tested.
GSM1k: new questions of the same kind
GSM8k is a public set of grade school arithmetic problems. In 2024 a team at Scale AI wrote a matching private set called GSM1k: 1,205 problems "requiring only elementary mathematical reasoning", created "solely with human annotators" [2]. They then compared each model's score on the public set with its score on the new one.
The largest fall was 8%. The paper reports "accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting across almost all model sizes" [2]. Overfitting means a model scores higher on the questions it has seen than on new questions of the same kind. The paper names the families: "Phi, Mistral and some models in the Llama family seem to be overfitting GSM8k, while models such as Gemini, GPT, and Claude show little to no signs of overfitting" [2].
The authors also measured how readily each model could produce a GSM8k example word for word. Models that did so more readily lost more on the new set, and the paper reports the strength of that link as Spearman's r squared of 0.36, on a scale where 1 is a perfect link [2].
Two facts limit what this study shows. The authors state that "all models broadly demonstrate generalization to novel math problems" [2], so the fall was partial for every model. And Scale AI sells evaluation services, which a reader should know when judging the result. The authors kept GSM1k private "to prevent a similar problem of data contamination occurring in the future" [2].
GSM-Symbolic: the same questions with different numbers
GSM-Symbolic tests whether a model's score depends on the exact names and numbers it may have seen. Mirzadeh and five co-authors rebuilt maths questions as templates in which names and numbers can be swapped, "using 100 templates and generating 50 samples per template, resulting in 5000 total examples for each benchmark" [3].
Three results are in the paper. First, "the performance of all models declines when only the numerical values in the question are altered" [3]. Second, a score moves between versions of the same question set: for one model, Gemma2-9B, "the gap between the worst performance and the best performance is more than 12%" [3]. Third, adding one sentence that looks relevant and has no effect on the answer lowered scores by up to 65% across the models tested [3].
The authors offer an explanation and label it a hypothesis: that the models "replicate reasoning steps from their training data" [3]. The models tested date from 2024.
The SWE-bench findings: the code was already public
SWE-bench asks a model to fix real issues in public code projects. Three sources report leaks in it.
SWE-Bench+ looked at the tasks passed by one setup: SWE-Agent, a program that lets a model work on code in steps, running the GPT-4 model. In 32.67% of those passing fixes, "the solutions were directly provided in the issue report or the comments" [5]. A further 31.08% passed because the project's tests were too weak to tell a correct fix from a wrong one. With both groups removed, the score went from 12.47% to 3.97% [5]. The working: 12.47 minus 3.97 is 8.50 points, and 3.97 divided by 12.47 is 0.318, so 31.8% of the original score was left. Both shares are of one agent's successful fixes only.
The SWE-Bench Illusion measured memory more directly. Given only the issue text, with no view of the project's files, models named the file that contained the fault with "up to 76% accuracy" [6]. On code projects outside SWE-bench the figure was up to 53%, which is 23 points lower. Word-for-word overlap with the reference fix, the human-written fix stored as the correct answer (the paper's measure is "consecutive 5-gram accuracy") reached up to 35% on SWE-bench and up to 18% on other benchmarks [6]. The authors describe the gap as "pointing to possible data contamination or memorization" [6], and the word "possible" is theirs.
OpenAI's audit used GPT-5 to probe GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview. By OpenAI's account, all of them could reproduce the reference fix or exact wording from the task for certain tasks, and the first sign was that GPT-5.2 solved 31 tasks OpenAI had judged to be almost impossible to solve [7]. This is OpenAI's audit of a benchmark it co-built, and the post gives no contamination rate per model.
How to tell whether a model has seen a test
Researchers use four checks. A buyer can run the last two.
- Compare dates. If the benchmark was public before the model's training cut-off, exposure is possible [5]. Ask the vendor for the cut-off date.
- Ask for something only memory could supply. Deng and co-authors hid a wrong answer option in MMLU questions and asked the model to fill it in. Exact matches of 52% and 57% show the models had seen those questions [4]. How many points that added to their scores is outside what the study measured.
- Change the details. Swap the names and numbers in a question and score the model again, as GSM-Symbolic did [3]. A model that understood the question keeps its score.
- Write new questions of the same kind. Compare the score on the public set with the score on the new set, as GSM1k did [2].
What LiveBench and private test sets do about it
LiveBench is a benchmark designed to limit contamination by replacing its questions. Its questions are "based on recently-released math competitions, arXiv papers, news articles, and datasets" and are "added and updated on a monthly basis" [8]. New questions stay unpublished for one month, "so that the public leaderboard always has 1/6 questions that are private" [8]. A leaderboard is the public table that ranks models by score, and 1 in 6 is 16.7%. A program scores each answer against a known correct answer, with no language model acting as judge. The April 2025 version of the paper describes 18 tasks in 6 categories, with the top models below 70% accuracy [8]. The current LiveBench table was not read for this page.
Other benchmark owners keep part of the test unpublished. ARC-AGI-2, a reasoning benchmark made of puzzle tasks, has 120 public evaluation tasks, 120 semi-private ones and 120 private ones, and its owner states: ""Private" means these tasks have not been exposed to third-parties" [9].
A private set gives stronger protection against contamination than a public one. Its cost is that nobody outside the owner can read the questions and check them for errors, so the reader has to trust the owner. The page on benchmark saturation, the point where every strong model scores near the top, covers what happens when a benchmark's own errors are larger than the gaps between models.
What a buyer should do
Read a public benchmark score as an upper estimate of what the model can do on new work, and read the settings behind a vendor's number before comparing two of them. Then collect evidence that contamination cannot affect:
- Use your own cases. Your support tickets, contracts and internal documents were never published, so no model was trained on them. How to build an LLM eval dataset gives the method.
- Keep the cases private. Store them with your other confidential business data.
- Add cases written after the model's cut-off date. A question about last month's product change cannot be in training data gathered before it.
- Vary the details. Change names, amounts and dates in a sample of cases and check that the score holds.
How to choose an AI model with your own evals puts these steps in order, and the hub page sets out what public scores are still good for. A private set can mislead in its own ways, which when evals give false confidence describes. If a supplier shows you its own test results, how to evaluate a vendor's eval suite lists what to ask.
Where Reveneau fits
Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass an eval suite before release. We write that suite from the client's specification before any code exists. Contamination is one reason that order matters: checks written from one client's specification were in no model's training data, so a pass on them cannot come from memory of a public test.
When a build includes an AI feature, we choose the model by running the candidates on cases taken from the client's own task. A public benchmark score is our reason to include a model in that comparison, and the client's cases decide it. Reveneau, as a company, takes responsibility for the whole project through production and after release.
To have a model chosen on your own cases, contact us or read about AI development at Reveneau.
Common questions
What is benchmark contamination?
Benchmark contamination is what happens when the questions or answers of a public test reach a model before the test, usually because they were on the public internet and were collected into the model's training data. A survey from University College Dublin describes it as models incorporating benchmark information from their training data, which makes their measured performance inaccurate or unreliable. The score then partly reflects memory.
What is data leakage in an AI benchmark?
Data leakage in an AI benchmark is any route by which the answer reaches the model outside its own reasoning. One route is training data that contains the test questions. A second route is the test itself: the SWE-Bench+ study found that in 32.67% of the fixes one test setup passed, the solution was written in the issue text or the comments that the model was given.
How do I know if a model has seen the test questions?
You can test for exposure in two ways without access to the training data. Change the names and numbers in a sample of questions and see whether the score holds, which is the GSM-Symbolic method. Or write new questions of the same kind and compare the two scores, which is the GSM1k method. Researchers also compare the benchmark's publication date with the model's training cut-off date.
What is GSM-Symbolic?
GSM-Symbolic is a maths benchmark published in October 2024 that rebuilds existing questions as templates so that names and numbers can be changed. The authors used 100 templates with 50 versions each, which gives 5,000 examples. Every model they tested scored lower when only the numbers were changed, and one added sentence with no effect on the answer lowered scores by up to 65% in the worst case.
What did the GSM1k study find about contamination?
The GSM1k study, from Scale AI in 2024, found accuracy drops of up to 8% when models moved from the public GSM8k maths set to 1,205 newly written problems. The paper names Phi, Mistral and some Llama models as showing overfitting, and says Gemini, GPT and Claude models showed little to no sign of it. All models still solved new problems, so the loss was partial.
What is LiveBench?
LiveBench is a benchmark built to limit contamination by replacing its questions over time. Its paper says questions come from recently released maths competitions, arXiv papers, news articles and datasets, and are added and updated monthly. New questions are held back for one month, so 1 in 6 questions is private at any time. Answers are scored by a program against a known correct answer.
Are private test sets better than public ones?
Private test sets give stronger protection against contamination, because questions that were never published cannot be in training data. ARC-AGI-2 keeps 120 private tasks for this reason, and Scale AI did not release GSM1k. The cost of a private set is that outsiders cannot read the questions and check them for errors. A set built from your own unpublished business cases gives you both privacy and the ability to inspect it.
Does contamination affect every model by the same amount?
Contamination does not affect every model by the same amount, so no single discount can be applied to all scores. In the GSM1k study the largest accuracy drop was 8% and it was concentrated in some model families, while the paper says the leading models show minimal signs of overfitting. The same paper found that models more able to reproduce a public example word for word tended to lose more on the new questions.
How much can contamination raise a benchmark score?
The measured size of contamination depends on the benchmark and the model. On grade school maths, GSM1k found drops of up to 8%. On SWE-bench, the SWE-Bench+ authors removed tasks with leaked solutions and weak tests, and one test setup's score fell from 12.47% to 3.97%, which left 31.8% of the original score. That second figure covers one setup with one model and includes weak tests as well as leaks.
How much work is it to protect my own model choice from contamination?
Protecting a model choice from contamination takes a set of cases from your own business and the discipline to keep them unpublished. The cases already exist in your support tickets, contracts or internal documents. The added work is writing down the correct answer for each one and, for a sample, changing names, amounts and dates to confirm the score holds, as the GSM-Symbolic authors did with their 100 templates.
Does benchmark contamination matter if I only use a model on my own documents?
Benchmark contamination matters to you at the moment you choose a model, because the public scores you compare may be raised by exposure to the test. OpenAI stopped reporting SWE-bench Verified on 23 February 2026 and said gains on it increasingly reflected exposure at training time. Once you test candidates on your own unpublished documents, contamination of public benchmarks no longer affects your decision.
What should I do when a vendor quotes a score on a public benchmark?
When a vendor quotes a public benchmark score, ask for the model's training cut-off date and compare it with the date the benchmark was published. The SWE-Bench+ study found over 94% of SWE-bench issues were created before the tested models' knowledge cut-off dates. Then treat the score as an upper estimate and ask the vendor to run the model on a small set of your own unpublished cases.
References
- [1] Xu, Guan, Greene and Kechadi (University College Dublin), Benchmark Data Contamination of Large Language Models: A Survey, arXiv (6 June 2024): the definition of benchmark data contamination. Abstract read.
- [2] Zhang, Da, Lee and co-authors (Scale AI), A Careful Examination of Large Language Model Performance on Grade School Arithmetic, arXiv (1 May 2024, latest version 22 November 2024): 1,205 human-written problems, accuracy drops of up to 8%, the model families named, Spearman's r squared of 0.36, and the decision to keep GSM1k private.
- [3] Mirzadeh, Alizadeh, Shahrokhi, Tuzel, Bengio and Farajtabar, GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, arXiv (7 October 2024, latest version 27 August 2025): 100 templates with 50 samples each, lower scores when only numbers change, a gap of more than 12% for Gemma2-9B, drops of up to 65% from one added clause, and the authors' hypothesis.
- [4] Deng, Zhao, Tang, Gerstein and Cohan (Yale), Investigating Data Contamination in Modern Benchmarks for Large Language Models, arXiv (16 November 2023, latest version 3 April 2024): exact match rates of 52% for ChatGPT and 57% for GPT-4 on hidden MMLU answer options. Abstract read.
- [5] Aleithan, Xue, Mohajer, Nnorom, Uddin and Wang, SWE-Bench+: Enhanced Coding Benchmark for LLMs, arXiv (9 October 2024): 32.67% of passing fixes had the solution in the issue text, 31.08% passed on weak tests, the score fell from 12.47% to 3.97%, and over 94% of issues predate the models' knowledge cut-off dates. Abstract read.
- [6] Liang, Garg and Moghaddam, The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason, arXiv (14 June 2025, latest version 1 December 2025): up to 76% accuracy naming the faulty file on SWE-bench against up to 53% elsewhere, and up to 35% against up to 18% for word-for-word overlap. Abstract read.
- [7] OpenAI, Why SWE-bench Verified no longer measures frontier coding capabilities (23 February 2026), read from the Internet Archive snapshot of 15 September 2026: OpenAI stopped reporting the benchmark, the three models probed, the 31 tasks, and OpenAI's conclusion about exposure at training time.
- [8] White, Dooley, Roberts and co-authors, LiveBench: A Challenging, Contamination-Limited LLM Benchmark, arXiv (27 June 2024, latest version 18 April 2025): question sources, monthly updates, 1 in 6 questions private, automatic scoring with no LLM judge, 18 tasks in 6 categories, top models below 70%.
- [9] ARC Prize Foundation, ARC-AGI-2 benchmark page (read 30 September 2026): 120 public evaluation, 120 semi-private and 120 private tasks, and the foundation's definition of private.
Related reading
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
More in Why a score can mislead
Benchmark saturation and Goodhart's law: when a score stops separating models
A benchmark is saturated when every strong model scores so close to the top that the gaps between them are smaller than the benchmark's own errors or its margin of error. Stanford's AI Index 2026 reports the top 15 models on the MMLU-Pro knowledge test all above 87%, with just over 4 percentage points between first and fifteenth. Goodhart's law, in Marilyn Strathern's wording, explains the pattern: "When a measure becomes a target, it ceases to be a good measure." The same law applies to the eval set you build for your own product, so keep part of it unseen by the people who tune against it and add new cases on a schedule.
How to read a model vendor's benchmark claims
A benchmark number in a model announcement depends on settings the vendor chose: the version of the benchmark, the program that ran the model, the tools allowed, the thinking effort, the number of attempts and the scoring rule. By OpenAI's own account, GPT-4 scored 2.7% and 28.3% on the same coding benchmark under two different surrounding programs. In Anthropic's system card for Claude Opus 5.5, one benchmark is reported at 81.8 under partial scoring and 48.7 under strict scoring for the same runs. This page gives twelve questions to ask before you compare two published scores, and explains what a rating built from votes, such as Arena's, measures.