Why a score can mislead

Benchmark contamination: when the test questions were in the training data

Benchmark contamination happens when a benchmark's public questions and answers end up in the text a model was trained on, so the model can score well by repeating what it has read. The measured effect is real and uneven. Scale AI's GSM1k study found accuracy drops of up to 8% on 1,205 newly written maths problems for some model families, and little to no sign of the problem in Gemini, GPT and Claude models. On SWE-bench, one study saw a score fall from 12.47% to 3.97% after leaked and weakly tested tasks were removed. A buyer protects against contamination by testing on private cases from their own business, which no model has read.

Published September 30, 2026. Editorial.

Key takeaways

  • Contamination means test questions or their answers reached a model before the test, usually through public text collected for training. The score then partly measures memory.
  • GSM1k (Scale AI, 2024) scored models on 1,205 new maths problems and found accuracy drops of up to 8% for some model families, while Gemini, GPT and Claude models showed little to no sign of overfitting, meaning a higher score on seen questions.
  • On SWE-bench, the SWE-Bench+ study found the solution inside the issue text for 32.67% of the fixes that one test setup passed, and that setup's score fell from 12.47% to 3.97% once flawed tasks were removed.
  • LiveBench limits contamination by adding and updating questions monthly and keeping 1 in 6 questions unpublished at any time.
  • Cases from your own business that were never published are the one test set you can be sure no model was trained on.

On 23 February 2026 OpenAI said it had stopped reporting scores on SWE-bench Verified, a coding benchmark it helped to create [7]. One reason it gave was that every leading model it tested could reproduce, for certain tasks, the original human-written fix or exact details of the task text. OpenAI's conclusion was that gains on the benchmark "increasingly reflect how much the model was exposed to the benchmark at training time" [7].

The name for that problem is benchmark contamination. This page explains what it is, sets out what six studies measured, and says what a buyer can do about it. Our position: a score on public questions is weaker evidence than a score on questions the model could never have read, and the questions from your own business are the easiest ones of that kind to get.

What benchmark contamination means

A benchmark is a fixed, public set of questions or tasks with known answers, and many models are scored on it. A large language model (LLM) learns from a large collection of text gathered up to a certain date. That collection is called the training data, and the date is called the training cut-off. If a benchmark's questions and answers were on the public internet before that date, they can be inside the collection. The model can then answer a test question by repeating text it has already read.

A survey by Xu and co-authors at University College Dublin defines the problem this way: it occurs when language models "inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance during the evaluation phase" [1]. The survey names it Benchmark Data Contamination. Other papers call the same thing data leakage [5] or test set contamination [8].

The authors of the GSM1k study use a wider definition: "data closely resembling benchmark questions leaks into the training data" [2]. Under that definition a reworded copy of a question, with the same answer, counts too.

A second type of leak happens inside the test itself: the answer is written in the question the model is given. The SWE-bench findings below include both types.

Why a public benchmark is exposed to this

A benchmark has to be published so that other people can run it, and publication puts it where training text is collected. SWE-bench is built from issues in public code projects (an issue is a written report of a fault or a request). One study of it found that "over 94% of the issues were created before LLM's knowledge cutoff dates" [5].

A buyer cannot check this directly when a vendor keeps a model's training data private. Deng and co-authors at Yale wrote that the problem "is especially critical for closed-source models and certain open-source models where training data transparency is lacking" [4]. The evidence therefore comes from experiments that test a model without access to its training data.

What six studies measured

Study What was done What was found
GSM1k, Zhang and co-authors at Scale AI, 2024 [2] People wrote 1,205 new maths problems and each model was scored on the new set and on the public set, GSM8k Accuracy fell by up to 8% for some model families. The leading models showed "minimal signs of overfitting"
GSM-Symbolic, Mirzadeh and co-authors, 2024 [3] The same questions were rebuilt as 100 templates with 50 versions each, so names and numbers could change Every model scored lower when only the numbers changed
Deng and co-authors, 2023 [4] One wrong answer option was hidden in questions from MMLU, a multiple-choice knowledge benchmark, and the model was asked to fill it in ChatGPT matched the hidden option exactly 52% of the time and GPT-4 57%
SWE-Bench+, Aleithan and co-authors, 2024 [5] The authors read by hand the tasks that one setup, SWE-Agent with GPT-4, had passed 32.67% of the passing fixes had the solution in the issue text. The score fell from 12.47% to 3.97%
The SWE-Bench Illusion, Liang, Garg and Moghaddam, 2025 [6] Models were asked to name the faulty file from the issue text alone Up to 76% correct on SWE-bench tasks and up to 53% on tasks from other code projects
OpenAI's audit, 23 February 2026 [7] GPT-5 was used to probe three models for memorised material All three could reproduce the reference fix or exact task details for certain tasks

For studies [4], [5] and [6], this page takes its figures from each paper's abstract. Every "up to" figure is the best or worst case across the models tested.

GSM1k: new questions of the same kind

GSM8k is a public set of grade school arithmetic problems. In 2024 a team at Scale AI wrote a matching private set called GSM1k: 1,205 problems "requiring only elementary mathematical reasoning", created "solely with human annotators" [2]. They then compared each model's score on the public set with its score on the new one.

The largest fall was 8%. The paper reports "accuracy drops of up to 8%, with several families of models showing evidence of systematic overfitting across almost all model sizes" [2]. Overfitting means a model scores higher on the questions it has seen than on new questions of the same kind. The paper names the families: "Phi, Mistral and some models in the Llama family seem to be overfitting GSM8k, while models such as Gemini, GPT, and Claude show little to no signs of overfitting" [2].

The authors also measured how readily each model could produce a GSM8k example word for word. Models that did so more readily lost more on the new set, and the paper reports the strength of that link as Spearman's r squared of 0.36, on a scale where 1 is a perfect link [2].

Two facts limit what this study shows. The authors state that "all models broadly demonstrate generalization to novel math problems" [2], so the fall was partial for every model. And Scale AI sells evaluation services, which a reader should know when judging the result. The authors kept GSM1k private "to prevent a similar problem of data contamination occurring in the future" [2].

GSM-Symbolic: the same questions with different numbers

GSM-Symbolic tests whether a model's score depends on the exact names and numbers it may have seen. Mirzadeh and five co-authors rebuilt maths questions as templates in which names and numbers can be swapped, "using 100 templates and generating 50 samples per template, resulting in 5000 total examples for each benchmark" [3].

Three results are in the paper. First, "the performance of all models declines when only the numerical values in the question are altered" [3]. Second, a score moves between versions of the same question set: for one model, Gemma2-9B, "the gap between the worst performance and the best performance is more than 12%" [3]. Third, adding one sentence that looks relevant and has no effect on the answer lowered scores by up to 65% across the models tested [3].

The authors offer an explanation and label it a hypothesis: that the models "replicate reasoning steps from their training data" [3]. The models tested date from 2024.

The SWE-bench findings: the code was already public

SWE-bench asks a model to fix real issues in public code projects. Three sources report leaks in it.

SWE-Bench+ looked at the tasks passed by one setup: SWE-Agent, a program that lets a model work on code in steps, running the GPT-4 model. In 32.67% of those passing fixes, "the solutions were directly provided in the issue report or the comments" [5]. A further 31.08% passed because the project's tests were too weak to tell a correct fix from a wrong one. With both groups removed, the score went from 12.47% to 3.97% [5]. The working: 12.47 minus 3.97 is 8.50 points, and 3.97 divided by 12.47 is 0.318, so 31.8% of the original score was left. Both shares are of one agent's successful fixes only.

The SWE-Bench Illusion measured memory more directly. Given only the issue text, with no view of the project's files, models named the file that contained the fault with "up to 76% accuracy" [6]. On code projects outside SWE-bench the figure was up to 53%, which is 23 points lower. Word-for-word overlap with the reference fix, the human-written fix stored as the correct answer (the paper's measure is "consecutive 5-gram accuracy") reached up to 35% on SWE-bench and up to 18% on other benchmarks [6]. The authors describe the gap as "pointing to possible data contamination or memorization" [6], and the word "possible" is theirs.

OpenAI's audit used GPT-5 to probe GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview. By OpenAI's account, all of them could reproduce the reference fix or exact wording from the task for certain tasks, and the first sign was that GPT-5.2 solved 31 tasks OpenAI had judged to be almost impossible to solve [7]. This is OpenAI's audit of a benchmark it co-built, and the post gives no contamination rate per model.

How to tell whether a model has seen a test

Researchers use four checks. A buyer can run the last two.

  1. Compare dates. If the benchmark was public before the model's training cut-off, exposure is possible [5]. Ask the vendor for the cut-off date.
  2. Ask for something only memory could supply. Deng and co-authors hid a wrong answer option in MMLU questions and asked the model to fill it in. Exact matches of 52% and 57% show the models had seen those questions [4]. How many points that added to their scores is outside what the study measured.
  3. Change the details. Swap the names and numbers in a question and score the model again, as GSM-Symbolic did [3]. A model that understood the question keeps its score.
  4. Write new questions of the same kind. Compare the score on the public set with the score on the new set, as GSM1k did [2].

What LiveBench and private test sets do about it

LiveBench is a benchmark designed to limit contamination by replacing its questions. Its questions are "based on recently-released math competitions, arXiv papers, news articles, and datasets" and are "added and updated on a monthly basis" [8]. New questions stay unpublished for one month, "so that the public leaderboard always has 1/6 questions that are private" [8]. A leaderboard is the public table that ranks models by score, and 1 in 6 is 16.7%. A program scores each answer against a known correct answer, with no language model acting as judge. The April 2025 version of the paper describes 18 tasks in 6 categories, with the top models below 70% accuracy [8]. The current LiveBench table was not read for this page.

Other benchmark owners keep part of the test unpublished. ARC-AGI-2, a reasoning benchmark made of puzzle tasks, has 120 public evaluation tasks, 120 semi-private ones and 120 private ones, and its owner states: ""Private" means these tasks have not been exposed to third-parties" [9].

A private set gives stronger protection against contamination than a public one. Its cost is that nobody outside the owner can read the questions and check them for errors, so the reader has to trust the owner. The page on benchmark saturation, the point where every strong model scores near the top, covers what happens when a benchmark's own errors are larger than the gaps between models.

What a buyer should do

Read a public benchmark score as an upper estimate of what the model can do on new work, and read the settings behind a vendor's number before comparing two of them. Then collect evidence that contamination cannot affect:

  • Use your own cases. Your support tickets, contracts and internal documents were never published, so no model was trained on them. How to build an LLM eval dataset gives the method.
  • Keep the cases private. Store them with your other confidential business data.
  • Add cases written after the model's cut-off date. A question about last month's product change cannot be in training data gathered before it.
  • Vary the details. Change names, amounts and dates in a sample of cases and check that the score holds.

How to choose an AI model with your own evals puts these steps in order, and the hub page sets out what public scores are still good for. A private set can mislead in its own ways, which when evals give false confidence describes. If a supplier shows you its own test results, how to evaluate a vendor's eval suite lists what to ask.

Where Reveneau fits

Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass an eval suite before release. We write that suite from the client's specification before any code exists. Contamination is one reason that order matters: checks written from one client's specification were in no model's training data, so a pass on them cannot come from memory of a public test.

When a build includes an AI feature, we choose the model by running the candidates on cases taken from the client's own task. A public benchmark score is our reason to include a model in that comparison, and the client's cases decide it. Reveneau, as a company, takes responsibility for the whole project through production and after release.

To have a model chosen on your own cases, contact us or read about AI development at Reveneau.

Common questions

What is benchmark contamination?

Benchmark contamination is what happens when the questions or answers of a public test reach a model before the test, usually because they were on the public internet and were collected into the model's training data. A survey from University College Dublin describes it as models incorporating benchmark information from their training data, which makes their measured performance inaccurate or unreliable. The score then partly reflects memory.

What is data leakage in an AI benchmark?

Data leakage in an AI benchmark is any route by which the answer reaches the model outside its own reasoning. One route is training data that contains the test questions. A second route is the test itself: the SWE-Bench+ study found that in 32.67% of the fixes one test setup passed, the solution was written in the issue text or the comments that the model was given.

How do I know if a model has seen the test questions?

You can test for exposure in two ways without access to the training data. Change the names and numbers in a sample of questions and see whether the score holds, which is the GSM-Symbolic method. Or write new questions of the same kind and compare the two scores, which is the GSM1k method. Researchers also compare the benchmark's publication date with the model's training cut-off date.

What is GSM-Symbolic?

GSM-Symbolic is a maths benchmark published in October 2024 that rebuilds existing questions as templates so that names and numbers can be changed. The authors used 100 templates with 50 versions each, which gives 5,000 examples. Every model they tested scored lower when only the numbers were changed, and one added sentence with no effect on the answer lowered scores by up to 65% in the worst case.

What did the GSM1k study find about contamination?

The GSM1k study, from Scale AI in 2024, found accuracy drops of up to 8% when models moved from the public GSM8k maths set to 1,205 newly written problems. The paper names Phi, Mistral and some Llama models as showing overfitting, and says Gemini, GPT and Claude models showed little to no sign of it. All models still solved new problems, so the loss was partial.

What is LiveBench?

LiveBench is a benchmark built to limit contamination by replacing its questions over time. Its paper says questions come from recently released maths competitions, arXiv papers, news articles and datasets, and are added and updated monthly. New questions are held back for one month, so 1 in 6 questions is private at any time. Answers are scored by a program against a known correct answer.

Are private test sets better than public ones?

Private test sets give stronger protection against contamination, because questions that were never published cannot be in training data. ARC-AGI-2 keeps 120 private tasks for this reason, and Scale AI did not release GSM1k. The cost of a private set is that outsiders cannot read the questions and check them for errors. A set built from your own unpublished business cases gives you both privacy and the ability to inspect it.

Does contamination affect every model by the same amount?

Contamination does not affect every model by the same amount, so no single discount can be applied to all scores. In the GSM1k study the largest accuracy drop was 8% and it was concentrated in some model families, while the paper says the leading models show minimal signs of overfitting. The same paper found that models more able to reproduce a public example word for word tended to lose more on the new questions.

How much can contamination raise a benchmark score?

The measured size of contamination depends on the benchmark and the model. On grade school maths, GSM1k found drops of up to 8%. On SWE-bench, the SWE-Bench+ authors removed tasks with leaked solutions and weak tests, and one test setup's score fell from 12.47% to 3.97%, which left 31.8% of the original score. That second figure covers one setup with one model and includes weak tests as well as leaks.

How much work is it to protect my own model choice from contamination?

Protecting a model choice from contamination takes a set of cases from your own business and the discipline to keep them unpublished. The cases already exist in your support tickets, contracts or internal documents. The added work is writing down the correct answer for each one and, for a sample, changing names, amounts and dates to confirm the score holds, as the GSM-Symbolic authors did with their 100 templates.

Does benchmark contamination matter if I only use a model on my own documents?

Benchmark contamination matters to you at the moment you choose a model, because the public scores you compare may be raised by exposure to the test. OpenAI stopped reporting SWE-bench Verified on 23 February 2026 and said gains on it increasingly reflected exposure at training time. Once you test candidates on your own unpublished documents, contamination of public benchmarks no longer affects your decision.

What should I do when a vendor quotes a score on a public benchmark?

When a vendor quotes a public benchmark score, ask for the model's training cut-off date and compare it with the date the benchmark was published. The SWE-Bench+ study found over 94% of SWE-bench issues were created before the tested models' knowledge cut-off dates. Then treat the score as an upper estimate and ask the vendor to run the model on a small set of your own unpublished cases.

References