Why a score can mislead

How to read a model vendor's benchmark claims

A benchmark number in a model announcement depends on settings the vendor chose: the version of the benchmark, the program that ran the model, the tools allowed, the thinking effort, the number of attempts and the scoring rule. By OpenAI's own account, GPT-4 scored 2.7% and 28.3% on the same coding benchmark under two different surrounding programs. In Anthropic's system card for Claude Opus 5.5, one benchmark is reported at 81.8 under partial scoring and 48.7 under strict scoring for the same runs. This page gives twelve questions to ask before you compare two published scores, and explains what a rating built from votes, such as Arena's, measures.

Published September 30, 2026. Editorial.

Key takeaways

  • A published benchmark score is a model plus settings: the benchmark version, the surrounding program, the tools, the thinking effort, the number of attempts and the scoring rule.
  • By OpenAI's own account, GPT-4 scored between 2.7% and 28.3% on SWE-bench Lite depending on the program that ran it, a gap of 25.6 points for one model.
  • Anthropic's system card for Claude Opus 5.5 says competitor figures are taken from the other developers' own cards or from public ranking tables, so the columns of its table were produced by different companies.
  • An Arena rating is computed from people's votes between two anonymous answers. It measures which answer voters preferred, and the AI Index notes that preferences may not align with correctness.
  • On a 500-task benchmark a score of 71% has a 95% margin of 3.98 points either side, so a rival at 74% is inside the same interval.

OpenAI reported scores for one model, GPT-4, on one coding benchmark, SWE-bench Lite, inside two different surrounding programs. By OpenAI's own account the score was 2.7% with one program and 28.3% with the other [2]. The model and the tasks were the same, and the two scores were 25.6 points apart.

The surrounding program is called a scaffold: the code that hands the model its task, gives it tools and runs it step by step. It is one of several settings that whoever runs a benchmark has to choose. A benchmark number in a model announcement is therefore the result of a model plus a list of settings. This page names the settings, shows how one vendor's published document reports them, and ends with a list of questions to print.

Which settings change a benchmark score

Setting What a source shows
The scaffold GPT-4 on SWE-bench Lite scored between 2.7% and 28.3% depending on the scaffold, by OpenAI's account [2]
The version of the benchmark GPT-4o scored 16% on the original SWE-bench and 33.2% on the cleaned Verified set, by OpenAI's account [2]
The task subset OSWorld results can cover 369 tasks or 361, on the original benchmark or the revised OSWorld-Verified [3]
The tools allowed Humanity's Last Exam in Anthropic's card for Claude Opus 5.5: 64.4 with no tools and 67.7 with tools [1]
The scoring rule OSWorld 2.0 in the same card: 81.8 under partial scoring and 48.7 under strict scoring, for the same runs [1]
The thinking effort (how much reasoning the model is allowed before it answers) The same card reports its results at "max effort" and one benchmark at "xhigh effort" [1]
The number of runs The same card averages five trials [1]. In a review of 24 benchmarks, 14 did not repeat runs or report uncertainty [10]

Each score in that table is a company's own report on its own model. The scores are used here for one purpose: to show how far a setting can move a number.

What a model card is and what one card reports

A model card is the document a vendor publishes with a model release, listing test results and the conditions of each test. Anthropic calls its version a system card. One card was opened for this page: Anthropic's System Card for Claude Opus 5.5, dated 22 September 2026 [1]. No card from OpenAI or Google was read, so this section describes how one vendor reports its settings and makes no comparison between vendors.

The card states its settings in one note: "Unless otherwise noted, all Claude Opus 5.5 results use the following standard configuration: adaptive thinking at max effort, default sampling settings (temperature, top_p), averaged over five trials" [1]. In plain words, the model was allowed its largest amount of reasoning before each answer, the controls that set how much its output varies between runs were left at their defaults, and each score is the mean of five runs. The exception is in the same note: "Terminal-Bench 4.0 score is reported at xhigh effort" [1]. The card also says that context window sizes, meaning the amount of text the model can read at one time, "are evaluation dependent and do not exceed 1M tokens" [1]. A token is a piece of a word.

A reader can take three lessons from this card.

  1. The settings are in a note beside the table. Read the note before the numbers.
  2. The comparison columns come from other documents. The card says: "Competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards" [1]. A leaderboard is a public table that ranks models by score. Each column may therefore reflect a different scaffold, effort level and number of attempts, chosen by a different company.
  3. One cell can hold two results. On OSWorld 2.0 the card gives 81.8 and 48.7 for the same runs under two scoring rules, which is 33.1 points apart. On Humanity's Last Exam the row with tools is 3.3 points above the row with no tools [1].

The choice of benchmarks is a setting too. The card reports three SWE-bench variants, each as "an average over five trials", and Table 8.1.A of the card has no row for SWE-bench Verified [1]. SWE-bench explained covers the variants and why they differ.

Attempts, repeated runs and cost

Kapoor and co-authors at Princeton studied AI agents, which are programs in which a model works in steps and uses tools. They found that "simply calling the underlying model multiple times can increase accuracy" [4]. A score should therefore come with the number of attempts allowed per task. Agent reliability across repeated runs explains how to measure this.

The same paper argues that "AI agent evaluations must be cost-controlled" and that results should be reported with their cost in dollars [4]. For a sense of the amounts in mid-2024, the paper notes that "the authors of SWE-Agent capped each run of the agent at $4 USD" [4]. A higher score reached with a larger budget per task is a more expensive result, and an announcement that gives the score alone leaves that out.

An outsider often cannot repeat a published result. BetterBench, a Stanford review of 24 AI benchmarks, found that 14 "did not perform multiple evaluations of the same model or report statistical significance or uncertainty of results", and that 17 "do not provide easy-to-run scripts to replicate the results reported in the initial paper" [10]. In plain words, most of the 24 gave no margin of error, and most gave an outsider no ready way to repeat the published result.

The margin of error on a score

A benchmark is a sample of questions, so each score has a margin of error. Evan Miller's paper on the statistics of evals, written at Anthropic, names the habit it criticises: "Evals are commonly run and reported with a "highest number is best" mentality" [5]. Its worked example uses two made-up models and three tests, with gaps of plus 2.5, minus 3.1 and minus 2.7 points, to show that gaps of that size may be chance [5].

An invented illustration with exact arithmetic. Two models score 71% and 74% on a benchmark of 500 tasks. For the first model:

  • Standard error, which is how far the score would be expected to move on a different sample of tasks: the square root of (0.71 x 0.29 / 500) is 0.0203, which is 2.03 points.
  • 95% margin, which is the range either side of the score expected to hold the result in 95 of 100 repeats: 1.96 x 2.03 is 3.98 points.
  • Interval: 71 minus 3.98 to 71 plus 3.98, which is 67.0% to 75.0%.

The second model's 74% is inside that interval, so this benchmark alone does not rank the two. Miller recommends a more exact comparison, which looks at the difference between the two models question by question [5]. A vendor that has the per-question results can supply it. Benchmark saturation covers what happens when every leading model is inside the same margin.

What Chatbot Arena, now called Arena, measures

Arena, formerly LMArena and before that Chatbot Arena [8], ranks models by people's votes. Its founders' paper describes the method: "a user can ask a question and get answers from two anonymous LLMs. Afterward, the user casts a vote for the model that delivers the preferred response" [6]. An LLM is a large language model. Voters write their own questions, and by January 2024 the site had received "over 240K votes" [6].

Each vote is a win for one model and a loss for the other. The paper fits Bradley-Terry coefficients to those records, which is a statistical method that turns wins and losses between pairs into one number per model [6]. Stanford's AI Index calls the resulting numbers "Arena Elo ratings" [9]. An Elo-style rating is a number of that type: it is computed from the results of one-against-one comparisons, and the model with the higher rating tended to win more of its comparisons.

The rating measures which answer voters preferred. The AI Index notes that "preferences may not align with correctness" [9], and the founders report that "the crowdsourced human votes are in good agreement with those of expert raters" [6]. The ratings at the top are close: as of March 2026 the AI Index lists the leading four companies between 1,503 and 1,481, a spread of 22 points [9].

The dispute about private testing on Arena

In April 2025 Singh and co-authors published "The Leaderboard Illusion". The paper says that "undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired", and it reports "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release" [7]. It estimates that Google and OpenAI received 19.2% and 20.4% of all the data on the platform [7]. Several of the authors work at Cohere, a model provider that competes on the same leaderboard, and the data shares are the authors' estimates.

Arena replied on 9 May 2025 that the paper "contains several incorrect claims" [8]. By Arena's account its policy on unreleased models had been public from 1 March 2024, "any model provider can submit as many public and private variants as they would like, as long as we have capacity for it", and the gain from private testing is 11 rating points after 50 tests and 3,000 votes, which Arena gives as an approximate figure. Arena's reply quotes the paper as claiming a gain of 100 or more points, and as putting open models at 8.8% of the leaderboard, against 40.9% in Arena's own published statistics. Those two figures from the paper were read in Arena's reply [8]. Arena also committed to mark a score as "provisional" until 2,000 new votes arrive after release, when more than 10 models were tested in parallel before release [8].

Both sides agree on the practice that matters to a buyer. A vendor can test several unreleased versions and release the one that scored best. Arena says so directly: "Chatbot Arena helps providers identify their best models, and that is a good thing" [8]. They disagree on how much that raises the published rating.

Whether a vendor's benchmark numbers are independent

A vendor's benchmark table is a company describing its own product. The AI Index reports that "third-party evaluations have documented cases where models perform more poorly in independent testing compared to developer-reported results" [9]. A result is independent when a party that gains nothing from the result ran the test and published the settings.

Apply the same check to the sources on this page. The OpenAI and Anthropic documents describe their own models, the author of the error-bars paper works at Anthropic, and the Arena study has authors at a competing provider. What an AI benchmark measures covers who makes benchmarks, and how to evaluate a vendor's eval suite covers the tests a supplier wrote itself.

Twelve questions to ask about a benchmark claim

Print this list and take it to the vendor.

  1. Which benchmark is this, which version, and how many tasks were run?
  2. Who ran the test: the vendor, the benchmark's owner or an outside party?
  3. Which program ran the model, and is it the one I would use?
  4. Which tools was the model allowed to use?
  5. How much thinking effort was set, and what does that setting cost?
  6. How many attempts were allowed per task?
  7. Is the score one run, the best run or an average, and of how many runs?
  8. Which scoring rule was used: partial credit, or strict pass and fail?
  9. Were the other models in the table run by the same party with the same settings?
  10. What did each task cost, and how long did it take?
  11. What is the margin of error, and is the gap to the next model larger than it?
  12. How many versions of the model were tested before this result was published?

Two further questions have their own pages: whether the model saw the questions during training, in benchmark contamination, and what the model scores on your work, in choosing a model with your own evals. The hub page links the benchmarks that vendors quote most.

Claim wording and what to check

Claim wording What to check Questions
"Leads on benchmark X" The size of the lead against the margin of error 1, 11
"X% on SWE-bench" The variant, the scaffold and the number of attempts 1, 3, 6
"Beats model Y" Who ran model Y, and with which settings 2, 9
"Ranked first on Arena" The rating gap to second place, and the versions tested before release 11, 12
"With tools" Which tools, and the score with no tools 4
"Average of five runs" The spread between the runs, and the cost of each 7, 10

Where Reveneau fits

Reveneau is an AI software development consultancy. We read a vendor's benchmark table with the questions on this page, and we make the decision on other evidence. When a build includes an AI feature, we run the candidate models on cases taken from the client's own task, with the same prompt and the same settings for each model, and we report quality, cost per task and response time together.

The same rule governs our own work. All of our code is written by AI, and every change must pass an eval suite, written from the specification before the code, before it is released. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release.

To compare models on your own cases, contact us or read about AI development at Reveneau.

Best for

  • A buyer comparing two or three models from their announcements before running any test
  • A team that a vendor has asked to accept a benchmark figure as proof of quality
  • A reviewer checking a supplier's proposal that quotes a position in a public ranking

Avoid if

  • You already have an eval set built from your own cases: run the models on it and decide from that
  • The models differ mainly on price or response time, which a benchmark score does not report

Check before you decide

  • The benchmark's name, version and number of tasks
  • Who ran each model in the comparison table, and with which settings
  • Whether the score is one run or an average, and of how many runs
  • The margin of error, and whether the gap to the next model is larger

Common questions

How do I read a model announcement that quotes benchmark scores?

Read a model announcement by finding the note that states the test settings before you read the scores. In Anthropic's system card for Claude Opus 5.5, one note gives the thinking effort, the sampling settings and the number of trials for every result. Then check which benchmark version was used, who ran the comparison models, and whether the gap to the next model is larger than the margin of error.

What is a model card?

A model card is the document a vendor publishes with a model release, listing test results and the conditions of each test. Anthropic calls its version a system card. The card for Claude Opus 5.5, dated 22 September 2026, states that its results use maximum thinking effort and are averaged over five trials. A model card is the vendor's own report, so its figures are self-reported.

What is Chatbot Arena, also called LMArena?

Chatbot Arena, later renamed LMArena and now Arena, is a site that ranks language models by people's votes. A visitor asks a question, receives answers from two anonymous models, and votes for the answer they prefer. The founders' paper reports over 240K votes by January 2024. The ranking shows which answers voters preferred, which is a different thing from which answers were correct.

What is an Elo-style rating for AI models?

An Elo-style rating is one number per model, computed from the wins and losses in one-against-one comparisons. The model with the higher rating tended to win more of its comparisons. Arena's founders describe the statistical method they use as Bradley-Terry coefficients. As of March 2026 the AI Index lists the four leading companies between 1,503 and 1,481, a spread of 22 points.

Which settings change a benchmark score?

Seven settings change a benchmark score: the benchmark version, the task subset, the program that runs the model, the tools allowed, the thinking effort, the number of attempts or runs, and the scoring rule. By OpenAI's own account, the program alone moved GPT-4 from 2.7% to 28.3% on SWE-bench Lite. In Anthropic's card, the scoring rule alone separates 81.8 from 48.7 on OSWorld 2.0.

Are a vendor's benchmark numbers independent?

A vendor's benchmark numbers are self-reported, because the vendor ran the test on its own model and chose the settings. Stanford's AI Index reports that third-party evaluations have documented cases where models did worse in independent testing than in developer-reported results. A number is independent when a party that gains nothing from the result ran the test and published its settings.

What is a scaffold in a benchmark result?

A scaffold is the program around the model during a test: the code that hands the model its task, gives it tools and runs it step by step. Two scaffolds can give one model two different scores. OpenAI reported GPT-4 at 2.7% with an early scaffold and 28.3% with another on SWE-bench Lite, so a score is only comparable when the scaffold is named.

What margin of error should I expect on a benchmark score?

The margin of error on a benchmark score depends on the number of tasks. In the invented illustration on this page, a score of 71% on 500 tasks has a standard error of 2.03 points and a 95% margin of 3.98 points either side, giving an interval of 67.0% to 75.0%. Evan Miller's paper on error bars recommends reporting that figure with every score.

Can a vendor test several versions of a model and publish only the best score?

A vendor can test several unreleased versions on Arena and release the one that scored best, and Arena confirms this. The Leaderboard Illusion paper reports 27 private variants tested by Meta before one release. Arena's response says any provider may submit private variants and puts the gain at an approximate 11 rating points after 50 tests and 3,000 votes. The two sides disagree on the size of the effect.

How long does it take to check a vendor's benchmark claim?

Checking a vendor's benchmark claim takes one reading of the settings note and one message to the vendor with the twelve questions on this page. In Anthropic's card for Claude Opus 5.5 the settings for every result are stated in a single note, with one exception named in the same place. The answers the vendor cannot give are as informative as the ones it can.

Does a rating built from votes tell me whether a model's answers are correct?

A rating built from votes records which answers people preferred, so correctness has to be tested separately. The AI Index notes that preferences may not align with correctness, while Arena's founders report that crowd votes agree well with expert raters. For a product where a wrong answer costs money, test correctness on your own cases with known right answers.

What should I do after I have read a vendor's benchmark claims?

After reading the claims, send the vendor the questions it left unanswered, then run the two or three candidate models on your own cases with the same prompt and settings for each. Anthropic's card shows why: its competitor figures are taken from other developers' documents, so the models in that table were tested by different companies with their own settings. Your own run puts every candidate under one setup.

References