AI benchmarks vs your own evals: how to read a model score / Why a score can mislead
How to read a model vendor's benchmark claims
A benchmark number in a model announcement depends on settings the vendor chose: the version of the benchmark, the program that ran the model, the tools allowed, the thinking effort, the number of attempts and the scoring rule. By OpenAI's own account, GPT-4 scored 2.7% and 28.3% on the same coding benchmark under two different surrounding programs. In Anthropic's system card for Claude Opus 5.5, one benchmark is reported at 81.8 under partial scoring and 48.7 under strict scoring for the same runs. This page gives twelve questions to ask before you compare two published scores, and explains what a rating built from votes, such as Arena's, measures.
Published September 30, 2026. Editorial.
Key takeaways
- A published benchmark score is a model plus settings: the benchmark version, the surrounding program, the tools, the thinking effort, the number of attempts and the scoring rule.
- By OpenAI's own account, GPT-4 scored between 2.7% and 28.3% on SWE-bench Lite depending on the program that ran it, a gap of 25.6 points for one model.
- Anthropic's system card for Claude Opus 5.5 says competitor figures are taken from the other developers' own cards or from public ranking tables, so the columns of its table were produced by different companies.
- An Arena rating is computed from people's votes between two anonymous answers. It measures which answer voters preferred, and the AI Index notes that preferences may not align with correctness.
- On a 500-task benchmark a score of 71% has a 95% margin of 3.98 points either side, so a rival at 74% is inside the same interval.
OpenAI reported scores for one model, GPT-4, on one coding benchmark, SWE-bench Lite, inside two different surrounding programs. By OpenAI's own account the score was 2.7% with one program and 28.3% with the other [2]. The model and the tasks were the same, and the two scores were 25.6 points apart.
The surrounding program is called a scaffold: the code that hands the model its task, gives it tools and runs it step by step. It is one of several settings that whoever runs a benchmark has to choose. A benchmark number in a model announcement is therefore the result of a model plus a list of settings. This page names the settings, shows how one vendor's published document reports them, and ends with a list of questions to print.
Which settings change a benchmark score
| Setting | What a source shows |
|---|---|
| The scaffold | GPT-4 on SWE-bench Lite scored between 2.7% and 28.3% depending on the scaffold, by OpenAI's account [2] |
| The version of the benchmark | GPT-4o scored 16% on the original SWE-bench and 33.2% on the cleaned Verified set, by OpenAI's account [2] |
| The task subset | OSWorld results can cover 369 tasks or 361, on the original benchmark or the revised OSWorld-Verified [3] |
| The tools allowed | Humanity's Last Exam in Anthropic's card for Claude Opus 5.5: 64.4 with no tools and 67.7 with tools [1] |
| The scoring rule | OSWorld 2.0 in the same card: 81.8 under partial scoring and 48.7 under strict scoring, for the same runs [1] |
| The thinking effort (how much reasoning the model is allowed before it answers) | The same card reports its results at "max effort" and one benchmark at "xhigh effort" [1] |
| The number of runs | The same card averages five trials [1]. In a review of 24 benchmarks, 14 did not repeat runs or report uncertainty [10] |
Each score in that table is a company's own report on its own model. The scores are used here for one purpose: to show how far a setting can move a number.
What a model card is and what one card reports
A model card is the document a vendor publishes with a model release, listing test results and the conditions of each test. Anthropic calls its version a system card. One card was opened for this page: Anthropic's System Card for Claude Opus 5.5, dated 22 September 2026 [1]. No card from OpenAI or Google was read, so this section describes how one vendor reports its settings and makes no comparison between vendors.
The card states its settings in one note: "Unless otherwise noted, all Claude Opus 5.5 results use the following standard configuration: adaptive thinking at max effort, default sampling settings (temperature, top_p), averaged over five trials" [1]. In plain words, the model was allowed its largest amount of reasoning before each answer, the controls that set how much its output varies between runs were left at their defaults, and each score is the mean of five runs. The exception is in the same note: "Terminal-Bench 4.0 score is reported at xhigh effort" [1]. The card also says that context window sizes, meaning the amount of text the model can read at one time, "are evaluation dependent and do not exceed 1M tokens" [1]. A token is a piece of a word.
A reader can take three lessons from this card.
- The settings are in a note beside the table. Read the note before the numbers.
- The comparison columns come from other documents. The card says: "Competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards" [1]. A leaderboard is a public table that ranks models by score. Each column may therefore reflect a different scaffold, effort level and number of attempts, chosen by a different company.
- One cell can hold two results. On OSWorld 2.0 the card gives 81.8 and 48.7 for the same runs under two scoring rules, which is 33.1 points apart. On Humanity's Last Exam the row with tools is 3.3 points above the row with no tools [1].
The choice of benchmarks is a setting too. The card reports three SWE-bench variants, each as "an average over five trials", and Table 8.1.A of the card has no row for SWE-bench Verified [1]. SWE-bench explained covers the variants and why they differ.
Attempts, repeated runs and cost
Kapoor and co-authors at Princeton studied AI agents, which are programs in which a model works in steps and uses tools. They found that "simply calling the underlying model multiple times can increase accuracy" [4]. A score should therefore come with the number of attempts allowed per task. Agent reliability across repeated runs explains how to measure this.
The same paper argues that "AI agent evaluations must be cost-controlled" and that results should be reported with their cost in dollars [4]. For a sense of the amounts in mid-2024, the paper notes that "the authors of SWE-Agent capped each run of the agent at $4 USD" [4]. A higher score reached with a larger budget per task is a more expensive result, and an announcement that gives the score alone leaves that out.
An outsider often cannot repeat a published result. BetterBench, a Stanford review of 24 AI benchmarks, found that 14 "did not perform multiple evaluations of the same model or report statistical significance or uncertainty of results", and that 17 "do not provide easy-to-run scripts to replicate the results reported in the initial paper" [10]. In plain words, most of the 24 gave no margin of error, and most gave an outsider no ready way to repeat the published result.
The margin of error on a score
A benchmark is a sample of questions, so each score has a margin of error. Evan Miller's paper on the statistics of evals, written at Anthropic, names the habit it criticises: "Evals are commonly run and reported with a "highest number is best" mentality" [5]. Its worked example uses two made-up models and three tests, with gaps of plus 2.5, minus 3.1 and minus 2.7 points, to show that gaps of that size may be chance [5].
An invented illustration with exact arithmetic. Two models score 71% and 74% on a benchmark of 500 tasks. For the first model:
- Standard error, which is how far the score would be expected to move on a different sample of tasks: the square root of (0.71 x 0.29 / 500) is 0.0203, which is 2.03 points.
- 95% margin, which is the range either side of the score expected to hold the result in 95 of 100 repeats: 1.96 x 2.03 is 3.98 points.
- Interval: 71 minus 3.98 to 71 plus 3.98, which is 67.0% to 75.0%.
The second model's 74% is inside that interval, so this benchmark alone does not rank the two. Miller recommends a more exact comparison, which looks at the difference between the two models question by question [5]. A vendor that has the per-question results can supply it. Benchmark saturation covers what happens when every leading model is inside the same margin.
What Chatbot Arena, now called Arena, measures
Arena, formerly LMArena and before that Chatbot Arena [8], ranks models by people's votes. Its founders' paper describes the method: "a user can ask a question and get answers from two anonymous LLMs. Afterward, the user casts a vote for the model that delivers the preferred response" [6]. An LLM is a large language model. Voters write their own questions, and by January 2024 the site had received "over 240K votes" [6].
Each vote is a win for one model and a loss for the other. The paper fits Bradley-Terry coefficients to those records, which is a statistical method that turns wins and losses between pairs into one number per model [6]. Stanford's AI Index calls the resulting numbers "Arena Elo ratings" [9]. An Elo-style rating is a number of that type: it is computed from the results of one-against-one comparisons, and the model with the higher rating tended to win more of its comparisons.
The rating measures which answer voters preferred. The AI Index notes that "preferences may not align with correctness" [9], and the founders report that "the crowdsourced human votes are in good agreement with those of expert raters" [6]. The ratings at the top are close: as of March 2026 the AI Index lists the leading four companies between 1,503 and 1,481, a spread of 22 points [9].
The dispute about private testing on Arena
In April 2025 Singh and co-authors published "The Leaderboard Illusion". The paper says that "undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired", and it reports "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release" [7]. It estimates that Google and OpenAI received 19.2% and 20.4% of all the data on the platform [7]. Several of the authors work at Cohere, a model provider that competes on the same leaderboard, and the data shares are the authors' estimates.
Arena replied on 9 May 2025 that the paper "contains several incorrect claims" [8]. By Arena's account its policy on unreleased models had been public from 1 March 2024, "any model provider can submit as many public and private variants as they would like, as long as we have capacity for it", and the gain from private testing is 11 rating points after 50 tests and 3,000 votes, which Arena gives as an approximate figure. Arena's reply quotes the paper as claiming a gain of 100 or more points, and as putting open models at 8.8% of the leaderboard, against 40.9% in Arena's own published statistics. Those two figures from the paper were read in Arena's reply [8]. Arena also committed to mark a score as "provisional" until 2,000 new votes arrive after release, when more than 10 models were tested in parallel before release [8].
Both sides agree on the practice that matters to a buyer. A vendor can test several unreleased versions and release the one that scored best. Arena says so directly: "Chatbot Arena helps providers identify their best models, and that is a good thing" [8]. They disagree on how much that raises the published rating.
Whether a vendor's benchmark numbers are independent
A vendor's benchmark table is a company describing its own product. The AI Index reports that "third-party evaluations have documented cases where models perform more poorly in independent testing compared to developer-reported results" [9]. A result is independent when a party that gains nothing from the result ran the test and published the settings.
Apply the same check to the sources on this page. The OpenAI and Anthropic documents describe their own models, the author of the error-bars paper works at Anthropic, and the Arena study has authors at a competing provider. What an AI benchmark measures covers who makes benchmarks, and how to evaluate a vendor's eval suite covers the tests a supplier wrote itself.
Twelve questions to ask about a benchmark claim
Print this list and take it to the vendor.
- Which benchmark is this, which version, and how many tasks were run?
- Who ran the test: the vendor, the benchmark's owner or an outside party?
- Which program ran the model, and is it the one I would use?
- Which tools was the model allowed to use?
- How much thinking effort was set, and what does that setting cost?
- How many attempts were allowed per task?
- Is the score one run, the best run or an average, and of how many runs?
- Which scoring rule was used: partial credit, or strict pass and fail?
- Were the other models in the table run by the same party with the same settings?
- What did each task cost, and how long did it take?
- What is the margin of error, and is the gap to the next model larger than it?
- How many versions of the model were tested before this result was published?
Two further questions have their own pages: whether the model saw the questions during training, in benchmark contamination, and what the model scores on your work, in choosing a model with your own evals. The hub page links the benchmarks that vendors quote most.
Claim wording and what to check
| Claim wording | What to check | Questions |
|---|---|---|
| "Leads on benchmark X" | The size of the lead against the margin of error | 1, 11 |
| "X% on SWE-bench" | The variant, the scaffold and the number of attempts | 1, 3, 6 |
| "Beats model Y" | Who ran model Y, and with which settings | 2, 9 |
| "Ranked first on Arena" | The rating gap to second place, and the versions tested before release | 11, 12 |
| "With tools" | Which tools, and the score with no tools | 4 |
| "Average of five runs" | The spread between the runs, and the cost of each | 7, 10 |
Where Reveneau fits
Reveneau is an AI software development consultancy. We read a vendor's benchmark table with the questions on this page, and we make the decision on other evidence. When a build includes an AI feature, we run the candidate models on cases taken from the client's own task, with the same prompt and the same settings for each model, and we report quality, cost per task and response time together.
The same rule governs our own work. All of our code is written by AI, and every change must pass an eval suite, written from the specification before the code, before it is released. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release.
To compare models on your own cases, contact us or read about AI development at Reveneau.
Best for
- A buyer comparing two or three models from their announcements before running any test
- A team that a vendor has asked to accept a benchmark figure as proof of quality
- A reviewer checking a supplier's proposal that quotes a position in a public ranking
Avoid if
- You already have an eval set built from your own cases: run the models on it and decide from that
- The models differ mainly on price or response time, which a benchmark score does not report
Check before you decide
- The benchmark's name, version and number of tasks
- Who ran each model in the comparison table, and with which settings
- Whether the score is one run or an average, and of how many runs
- The margin of error, and whether the gap to the next model is larger
Common questions
How do I read a model announcement that quotes benchmark scores?
Read a model announcement by finding the note that states the test settings before you read the scores. In Anthropic's system card for Claude Opus 5.5, one note gives the thinking effort, the sampling settings and the number of trials for every result. Then check which benchmark version was used, who ran the comparison models, and whether the gap to the next model is larger than the margin of error.
What is a model card?
A model card is the document a vendor publishes with a model release, listing test results and the conditions of each test. Anthropic calls its version a system card. The card for Claude Opus 5.5, dated 22 September 2026, states that its results use maximum thinking effort and are averaged over five trials. A model card is the vendor's own report, so its figures are self-reported.
What is Chatbot Arena, also called LMArena?
Chatbot Arena, later renamed LMArena and now Arena, is a site that ranks language models by people's votes. A visitor asks a question, receives answers from two anonymous models, and votes for the answer they prefer. The founders' paper reports over 240K votes by January 2024. The ranking shows which answers voters preferred, which is a different thing from which answers were correct.
What is an Elo-style rating for AI models?
An Elo-style rating is one number per model, computed from the wins and losses in one-against-one comparisons. The model with the higher rating tended to win more of its comparisons. Arena's founders describe the statistical method they use as Bradley-Terry coefficients. As of March 2026 the AI Index lists the four leading companies between 1,503 and 1,481, a spread of 22 points.
Which settings change a benchmark score?
Seven settings change a benchmark score: the benchmark version, the task subset, the program that runs the model, the tools allowed, the thinking effort, the number of attempts or runs, and the scoring rule. By OpenAI's own account, the program alone moved GPT-4 from 2.7% to 28.3% on SWE-bench Lite. In Anthropic's card, the scoring rule alone separates 81.8 from 48.7 on OSWorld 2.0.
Are a vendor's benchmark numbers independent?
A vendor's benchmark numbers are self-reported, because the vendor ran the test on its own model and chose the settings. Stanford's AI Index reports that third-party evaluations have documented cases where models did worse in independent testing than in developer-reported results. A number is independent when a party that gains nothing from the result ran the test and published its settings.
What is a scaffold in a benchmark result?
A scaffold is the program around the model during a test: the code that hands the model its task, gives it tools and runs it step by step. Two scaffolds can give one model two different scores. OpenAI reported GPT-4 at 2.7% with an early scaffold and 28.3% with another on SWE-bench Lite, so a score is only comparable when the scaffold is named.
What margin of error should I expect on a benchmark score?
The margin of error on a benchmark score depends on the number of tasks. In the invented illustration on this page, a score of 71% on 500 tasks has a standard error of 2.03 points and a 95% margin of 3.98 points either side, giving an interval of 67.0% to 75.0%. Evan Miller's paper on error bars recommends reporting that figure with every score.
Can a vendor test several versions of a model and publish only the best score?
A vendor can test several unreleased versions on Arena and release the one that scored best, and Arena confirms this. The Leaderboard Illusion paper reports 27 private variants tested by Meta before one release. Arena's response says any provider may submit private variants and puts the gain at an approximate 11 rating points after 50 tests and 3,000 votes. The two sides disagree on the size of the effect.
How long does it take to check a vendor's benchmark claim?
Checking a vendor's benchmark claim takes one reading of the settings note and one message to the vendor with the twelve questions on this page. In Anthropic's card for Claude Opus 5.5 the settings for every result are stated in a single note, with one exception named in the same place. The answers the vendor cannot give are as informative as the ones it can.
Does a rating built from votes tell me whether a model's answers are correct?
A rating built from votes records which answers people preferred, so correctness has to be tested separately. The AI Index notes that preferences may not align with correctness, while Arena's founders report that crowd votes agree well with expert raters. For a product where a wrong answer costs money, test correctness on your own cases with known right answers.
What should I do after I have read a vendor's benchmark claims?
After reading the claims, send the vendor the questions it left unanswered, then run the two or three candidate models on your own cases with the same prompt and settings for each. Anthropic's card shows why: its competitor figures are taken from other developers' documents, so the models in that table were tested by different companies with their own settings. Your own run puts every candidate under one setup.
References
- [1] Anthropic, System Card: Claude Opus 5.5 (22 September 2026): the standard configuration note, the Terminal-Bench 4.0 exception, the context window statement, the source of competitor figures, and Anthropic's own figures for Humanity's Last Exam, OSWorld 2.0 and the SWE-bench variants. The only vendor card read for this page.
- [2] OpenAI, Introducing SWE-bench Verified (13 August 2024), read from the Internet Archive snapshot of 15 September 2026: GPT-4 between 2.7% and 28.3% on SWE-bench Lite depending on the scaffold, and GPT-4o at 16% on the original SWE-bench and 33.2% on Verified.
- [3] OSWorld authors, OSWorld project site (update note dated 28 July 2025, read 30 September 2026): the revised OSWorld-Verified version, and 8 tasks that may be excluded, leaving 361 of 369.
- [4] Kapoor, Stroebl, Siegel, Nadgir and Narayanan (Princeton University), AI Agents That Matter, arXiv (1 July 2024): repeated calls raise accuracy, agent evaluations must be cost-controlled, and the $4 cap per SWE-Agent run.
- [5] Evan Miller (Anthropic), Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations, arXiv (1 November 2024): the highest-number habit, the worked example with gaps of 2.5, 3.1 and 2.7 points, and the recommendation to compare models on question-level differences.
- [6] Chiang, Zheng, Sheng, Angelopoulos and co-authors (UC Berkeley), Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, arXiv (7 March 2024): how a vote is collected, over 240K votes by January 2024, Bradley-Terry coefficients, and agreement with expert raters.
- [7] Singh, Nan, Wang and co-authors (Cohere and universities), The Leaderboard Illusion, arXiv (29 April 2025, latest version 12 May 2025): private testing of multiple variants, 27 private variants tested by Meta, and estimated data shares of 19.2% and 20.4%.
- [8] Arena, Our Response to 'The Leaderboard Illusion' Writeup (9 May 2025): the claim of incorrect statements, the policy published on 1 March 2024, open submission of private variants, the +11 Elo estimate, the 100-point and 8.8% claims as Arena quotes them from the paper, Arena's 40.9% figure for open models, and the provisional marking commitment.
- [9] Stanford HAI, AI Index Report 2026, Chapter 2: Technical Performance (2026 edition): Arena Elo ratings as of March 2026, preferences and correctness, and the statement on independent testing against developer-reported results.
- [10] Reuel, Hardy, Smith, Lamparth, Hardy and Kochenderfer (Stanford), BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices, arXiv (20 November 2024): 14 of 24 benchmarks did not repeat runs or report uncertainty, and 17 of 24 gave no easy script to replicate results.
Related reading
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
Ask an AI vendor for its time to production, and skip the demo
Each of Wonderful's customer stories starts with the time it took to reach production. That time is the better proof, because it measures the vendor, while a demo only measures the model.
How to negotiate a software contract you can actually verify
Most build contracts describe effort, timeline, and payment, and leave the one hard question unanswered: on what evidence do you agree the thing is finished?
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
More in Why a score can mislead
Benchmark contamination: when the test questions were in the training data
Benchmark contamination happens when a benchmark's public questions and answers end up in the text a model was trained on, so the model can score well by repeating what it has read. The measured effect is real and uneven. Scale AI's GSM1k study found accuracy drops of up to 8% on 1,205 newly written maths problems for some model families, and little to no sign of the problem in Gemini, GPT and Claude models. On SWE-bench, one study saw a score fall from 12.47% to 3.97% after leaked and weakly tested tasks were removed. A buyer protects against contamination by testing on private cases from their own business, which no model has read.
Benchmark saturation and Goodhart's law: when a score stops separating models
A benchmark is saturated when every strong model scores so close to the top that the gaps between them are smaller than the benchmark's own errors or its margin of error. Stanford's AI Index 2026 reports the top 15 models on the MMLU-Pro knowledge test all above 87%, with just over 4 percentage points between first and fifteenth. Goodhart's law, in Marilyn Strathern's wording, explains the pattern: "When a measure becomes a target, it ceases to be a good measure." The same law applies to the eval set you build for your own product, so keep part of it unseen by the people who tune against it and add new cases on a schedule.