AI benchmarks vs your own evals: how to read a model score
A benchmark score answers one question: how did this model do on that public test, under the settings chosen by whoever ran it. A buyer needs the answer to a different question: how will the model do on my work. This guide explains the benchmarks vendors quote, including SWE-bench, MMLU, GPQA, ARC-AGI and the agent tests, with each one's size and method from its own paper. It covers three ways a score misleads: the questions were in the training data, every model scores near the top, and the vendor chose the settings. It ends with a seven-step method for choosing a model on your own cases.
Published September 30, 2026. Editorial.
Key takeaways
- A benchmark is a public, fixed set of items with a marking rule and a score table. A review of 445 LLM benchmarks found that less than 10% used complete real-world tasks.
- OpenAI reported its GPT-4o model at 16% on the original SWE-bench and 33.2% on the human-checked Verified set, so cleaning the test changed the score by a factor of 2.075.
- OpenAI stopped reporting SWE-bench Verified scores on 23 February 2026, 18 months after publishing the set, and on 8 July 2026 it withdrew its recommendation of the replacement benchmark.
- Researchers estimated that 6.49% of MMLU questions contain errors, which limits what a score above 90% on that test can show.
- In the tau-bench paper, agents built on gpt-4o passed fewer than 50% of tasks, and fewer than 25% of retail tasks when the same task had to succeed 8 times out of 8.
- OpenAI reported GPT-4 at 2.7% and at 28.3% on SWE-bench Lite depending on the program that ran the model, so ask for the settings behind every published score.
- On OpenAI's price page on 30 September 2026, one example task cost $55.00 or $5,500.00 per 100,000 requests depending on the model, and only your own cases show which model is good enough.
On 13 August 2024 OpenAI published SWE-bench Verified, a set of 500 coding tasks chosen after 93 software developers screened 1,699 samples by hand [1]. On 23 February 2026, 18 months later, OpenAI wrote that it had "stopped reporting SWE-bench Verified scores" and recommended "that other model developers do so too" [2]. Its stated reason was that higher scores on the test had come to reflect how much a model was exposed to the benchmark during training [2]. In the same post it recommended a replacement, SWE-bench Pro. On 8 July 2026 it withdrew that recommendation, after its own audit counted broken tasks in the replacement at 27.4% by one method and 34.1% by another [3]. Each of these statements is OpenAI's account of its own work, read from the Internet Archive copies of its posts.
A benchmark is a public, fixed set of questions or tasks that many models are scored on. The events above show what that means for a buyer: a benchmark score has an author, a date and a set of conditions, and all three can change. The score answers one question, "how did this model do on that test?". A buyer needs the answer to a second question, "how will it do on my work?". This guide explains the benchmarks you will see quoted, the three ways a score misleads, and how to run a small eval of your own to choose a model. An eval of your own is a test built from your product's real cases. Each section below gives the main point of one page and links to it.
The three ways a benchmark score misleads
| Problem | What happens | One piece of evidence |
|---|---|---|
| The questions were in the model's training text | The model scores well on questions it has already seen | On 1,205 newly written arithmetic problems, some model families lost up to 8% accuracy against the public set they resemble [4] |
| Every strong model scores near the top | The gaps that remain are too small to rank models | As of early 2026 the 15 leading models on MMLU-Pro all scored above 87% [5] |
| The vendor chose the settings | One model gets different scores under different conditions | OpenAI reported GPT-4 at 2.7% and at 28.3% on one coding benchmark, depending on the program that ran the model [1] |
The first four pages of the guide describe the tests. The next three take these problems one at a time, and the last page gives the method for deciding with your own data.
What an AI benchmark measures and what it leaves out
A benchmark has three parts: a fixed set of items chosen by its authors, a rule for marking each answer, and a public table of model scores. MMLU, a knowledge test from 2020, is a clear example. It has 15,908 multiple-choice questions across 57 subjects, each with four options, so guessing scores 25% [6]. A score of 90% on it is a true statement about those questions.
The difficulty starts when a reader treats the test score as a measure of the skill the test is named after. A review of 445 benchmarks for large language models (LLMs), carried out by 29 expert reviewers and presented at the NeurIPS 2025 conference, found that less than 10% used complete real-world tasks [7]. It also found that 16.0% used uncertainty estimates or statistical tests to compare results [7]. In plain words, most benchmarks test short questions in place of whole jobs, and most publish no margin of error, so a reader cannot tell whether a gap of one point between two models means anything.
A benchmark also leaves out everything particular to you: your documents, your users' wording, your rules for a correct answer, your limits on cost and response time. Your own eval contains exactly those things, and its cases stay private. What an AI benchmark measures and what it leaves out explains who makes benchmarks, how a score is produced, and how a public test differs from an eval you write yourself. If the word eval is new to you, start with what is an AI eval.
SWE-bench: what a coding benchmark score means
SWE-bench is the coding test quoted most often in model announcements. The original set has 2,294 tasks taken from real problem reports in 12 software projects written in the Python programming language, and the project site defines the score as "the percentage of task instances solved" [8]. By OpenAI's description, a task counts as solved when the model's change to the code makes the project's own tests pass [1].
The history of the Verified version shows how much the test itself can decide the score. When OpenAI and the SWE-bench team had people screen the tasks, 68.3% of the screened samples were removed, and 61.1% were flagged because their tests could mark a valid solution as wrong [1]. On the cleaned set of 500 tasks, OpenAI reported that its GPT-4o model scored 33.2%, against 16% on the original set [1]. That is 2.075 times the earlier score. Both figures describe the same model, so the difference comes from which tasks were in the test.
Sources also disagree on the current top score. OpenAI's February 2026 post gives 80.9% as the best result on Verified [2]. Stanford's AI Index 2026 gives 76.8% for the leader as of February 2026, a figure the report words as approximate [5]. Quote either one with its source and date.
A SWE-bench score counts tasks from other people's public Python projects that already have tests. It says little about private code, other programming languages, or a request that arrives without a clear description. SWE-bench explained covers every version, the grading method and the published critiques. For testing code written for you, see our guide to eval-driven development.
MMLU, GPQA and ARC-AGI: knowledge and reasoning benchmarks
Knowledge benchmarks ask exam questions. On MMLU, people recruited through Amazon Mechanical Turk, a website where people are paid to do small online tasks, with no special training, scored 34.5%, and the largest model tested in 2020 scored under 20 percentage points above guessing [6]. Once models score close to the maximum, the test's own mistakes matter. Researchers at the University of Edinburgh and other institutions re-checked 5,700 MMLU questions by hand and estimated that 6.49% of MMLU questions contain errors [9]. In the virology subject, 57% of the 100 questions they analysed contained errors [9]. The 6.49% figure is an estimate from a sample, and it limits what a score above 90% can tell you: part of the remaining gap comes from faulty questions.
Newer tests were built to stay hard for longer. GPQA uses science questions written by subject experts, ARC-AGI uses grid puzzles, and Humanity's Last Exam collects expert-written questions across dozens of subjects. Each has its own size and format. Scores on the newest tests rise fast: the AI Index reports that top accuracy on Humanity's Last Exam went "from under 10% to 38.3%" in a single year [5].
The point for a buyer is about fit. A multiple-choice knowledge score shows what a model can recall and reason about in an exam format. A product that answers customers from your documents depends on a different ability: finding the right passage in your text and staying inside it. Only a test on your documents measures that. MMLU, GPQA and ARC-AGI explained gives each benchmark's contents, scoring and known limits.
Agent benchmarks: GAIA, WebArena, tau-bench and METR time horizons
An agent is an AI system that takes actions through tools, such as searching a website or changing a record, over many steps. Agent benchmarks mark the end result. In tau-bench, a 2024 benchmark from the company Sierra, the agent talks to a simulated customer and uses tools under a written policy. The test then "compares the database state at the end of a conversation with the annotated goal state" [10]. In plain words, it checks whether the records ended up as they should. It has 165 tasks: 115 in a retail setting and 50 in an airline setting [10].
The tau-bench paper gives the fact that matters most for a buyer. Its authors report that agents built on the model gpt-4o succeeded on fewer than 50% of the tasks. When the same retail task had to succeed in 8 tries out of 8, the rate fell below 25% [10]. A single-run score does not show how often an agent fails when the same request comes in again, and a product receives the same request many times.
The other benchmarks in this group test different work. GAIA asks questions that need several tools to answer, WebArena sets tasks on working copies of websites, and METR's time horizon expresses a model's ability as the length of task, measured in the time a skilled person needs, that the model completes half the time. All of them run in a prepared test setting. Agent benchmarks explained gives the task counts, the human scores and the best model scores at publication. To test an agent of your own, use our guide to AI agent evals.
Benchmark contamination: the test questions were in the training data
A benchmark's questions and answers are public, so they can end up in the text a model is trained on. The model can then score well by repeating what it has seen. Researchers call this contamination.
The most direct way to measure it is to write new questions of the same kind and difficulty and compare. A team at Scale AI did that for a public set of school arithmetic problems. They wrote 1,205 new problems using human writers only and found "accuracy drops of up to 8%", with several model families showing the pattern across almost all model sizes [4]. The same paper reports that many models, including the leading ones, showed minimal signs of the problem [4]. Scale AI sells evaluation services, and its result is mixed.
Coding benchmarks have the same exposure, because their tasks come from public code. In its February 2026 post OpenAI wrote that every leading model it tested could reproduce the original human-written fix, or exact details of the problem statement, for certain SWE-bench Verified tasks [2]. The post names three models from three vendors and publishes no rate for any of them.
A buyer has one reliable protection. Cases taken from your own product have never been published, so no model has seen them. Benchmark contamination sets out each study, what was done and what was found, and describes the benchmarks that replace their questions on a schedule.
Benchmark saturation and Goodhart's law
A benchmark stops being useful when every strong model scores near the maximum. The AI Index calls this saturation, "where models reach scores so high that a test can no longer distinguish between them" [5]. The same report says that "Evaluations intended to be challenging for years are saturated in months" [5]. Its example is MMLU-Pro, a harder version of MMLU released in 2024. As of early 2026 the 15 leading models all scored above 87%, and the gap from first place to fifteenth was "just over 4 percentage points" [5]. A test can rank fifteen models inside a gap of that size only if its margin of error is smaller than the gap, and most benchmarks publish no margin of error [7].
The second half of the problem has a name from economics. Charles Goodhart was Chief Adviser to the Bank of England, and the anthropologist Marilyn Strathern restated Goodhart's idea in 1997 in the form most people quote: "When a measure becomes a target, it ceases to be a good measure" [11]. That wording is taken from a University of Cambridge page that cites both authors. A public benchmark is a target for every model vendor. Once a vendor's release is judged by a score, the vendor has a reason to tune the model toward that test, and the score rises faster than the skill.
The same law applies inside your own company. A team that adjusts its instructions against one fixed set of cases will end up fitting that set. Keep part of your cases unseen by the people who tune the product, and add new cases from real use. Benchmark saturation and Goodhart's law explains the signs that a benchmark has stopped separating models and how to protect your own set.
How to read a model vendor's benchmark claims
A benchmark number in an announcement depends on choices the vendor made: how many attempts the model had, which tools it could use, how long it could reason, which subset of tasks was run, and which program ran the model. That program has a name, scaffold: the software wrapped around a model that hands it tools and runs it step by step. OpenAI's 2024 post gives the size of the effect. GPT-4 scored between 2.7% and 28.3% on SWE-bench Lite depending on the scaffold [1]. The model was the same in both runs.
Public ranking sites have their own version of this problem. Arena is a site where people compare two unnamed models' answers and vote. A 2025 paper, The Leaderboard Illusion, reported that some providers test many private versions of a model before release, and it identified "27 private LLM variants tested by Meta in the lead-up to the Llama-4 release" [12]. Several of the paper's authors work at Cohere, which competes on the same ranking. Arena replied that the paper "contains several incorrect claims" and that its policy lets "any model provider" submit as many public and private variants as they would like [13]. Both sides agree on the practice that matters to a buyer: a vendor can test several versions privately and release the one that scored best.
How to read a model vendor's benchmark claims is a printable list of questions to ask about any published score. For claims about a supplier's private tests, see how to evaluate a vendor's eval suite.
How to choose an AI model with your own evals
Public rankings place the leading models close together. As of March 2026 the AI Index listed the top models of six companies between 1,424 and 1,503 rating points on Arena, the voting site described above, a spread of 79 points, and said the narrow gaps are "shifting competitive pressure toward cost, reliability, and domain-specific performance" [5]. Google Cloud's documentation says that results on your specific tasks give "valuable insights which cannot be derived from public leaderboards and general benchmarks" [14]. A leaderboard is a public table that ranks models by score, and that Google page describes a paid service.
Prices do separate the models. Take an invented task of 3,000 input tokens and 500 output tokens, where a token is a piece of a word and the unit vendors bill in. At the prices on OpenAI's own pricing page on 30 September 2026, that task costs $55.00 per 100,000 requests on the model gpt-6-luna and $5,500.00 on gpt-6-astra, a ratio of 100 [15]. Prices change, so date any price you quote. Whether the lower-priced model is good enough for your task is a question only your cases can answer.
The method has seven steps: write the task and what correct means, collect real cases, run every candidate the same way, score quality with cost per task and response time, read the failures, decide, and keep the set to run again when a model is replaced. Running 200 such cases on the higher-priced model costs $11.00 at the same published price [15]. How to choose an AI model with your own evals gives each step with the full price table. Building the case set is covered in our guide to LLM evals.
How Reveneau uses benchmarks and evals
At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. We treat model choice the same way. A public benchmark score helps us decide which models to try on a client's task. The choice itself comes from evals written for that client's product, with the client's definition of a correct answer, and the evals stay in the project so they can run again when a vendor replaces a model.
We grade the judged checks in our suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. If you have a model decision to make and want it made on your own cases, our AI development service is where that work starts, or you can contact us.
Explore the guide
The benchmarks vendors quote
SWE-bench explained: what a coding benchmark score means
SWE-bench is a public test of AI coding models. It gives a model a real bug report from an open software project and counts the task as solved when the project's tests pass after the model's change. The original set has 2,294 tasks from 12 Python projects. The score is the percentage of tasks solved, and it changes with the version of the test, the program wrapped around the model and the date. OpenAI stopped reporting the Verified version on 23 February 2026 after its own audit. A high score shows a model fixes public Python bugs that have tests. Your own requirements were outside the test.
MMLU, GPQA and ARC-AGI explained: knowledge and reasoning benchmarks
MMLU, GPQA, ARC-AGI and Humanity's Last Exam are public tests that model makers quote to show knowledge and reasoning. MMLU has 15,908 multiple-choice questions in 57 subjects, and the authors of a newer test write that models now score over 90% on it. GPQA has 448 graduate-level science questions on which experts scored 65%. ARC-AGI uses puzzles that ask the solver to work out a new rule. Humanity's Last Exam has 2,500 questions written by experts. Each score describes performance on that test's own questions. For a product that answers customers from your own documents, these scores say little, because none of them tests your documents, your search step or your customers' questions.
Agent benchmarks explained: GAIA, WebArena, tau-bench and METR time horizons
Agent benchmarks score whether an AI model can finish a task that takes several steps and tools. GAIA has 466 questions, WebArena has 812 website tasks, tau-bench has 165 customer service tasks and OSWorld has 369 desktop tasks. The METR time horizon reports the length of task, measured by how long it takes a skilled person, that a model completes half the time. At publication the best models scored between 12.24% and 61.2% on these tests, and the AI Index 2026 reports later best scores between 66.3% and 74.5%. Each score describes that benchmark's tasks on one date and one version. It leaves out repeat reliability, cost, and your own tools and rules, so use it to pick candidates and then test your own agent.
Why a score can mislead
Benchmark contamination: when the test questions were in the training data
Benchmark contamination happens when a benchmark's public questions and answers end up in the text a model was trained on, so the model can score well by repeating what it has read. The measured effect is real and uneven. Scale AI's GSM1k study found accuracy drops of up to 8% on 1,205 newly written maths problems for some model families, and little to no sign of the problem in Gemini, GPT and Claude models. On SWE-bench, one study saw a score fall from 12.47% to 3.97% after leaked and weakly tested tasks were removed. A buyer protects against contamination by testing on private cases from their own business, which no model has read.
Benchmark saturation and Goodhart's law: when a score stops separating models
A benchmark is saturated when every strong model scores so close to the top that the gaps between them are smaller than the benchmark's own errors or its margin of error. Stanford's AI Index 2026 reports the top 15 models on the MMLU-Pro knowledge test all above 87%, with just over 4 percentage points between first and fifteenth. Goodhart's law, in Marilyn Strathern's wording, explains the pattern: "When a measure becomes a target, it ceases to be a good measure." The same law applies to the eval set you build for your own product, so keep part of it unseen by the people who tune against it and add new cases on a schedule.
How to read a model vendor's benchmark claims
A benchmark number in a model announcement depends on settings the vendor chose: the version of the benchmark, the program that ran the model, the tools allowed, the thinking effort, the number of attempts and the scoring rule. By OpenAI's own account, GPT-4 scored 2.7% and 28.3% on the same coding benchmark under two different surrounding programs. In Anthropic's system card for Claude Opus 5.5, one benchmark is reported at 81.8 under partial scoring and 48.7 under strict scoring for the same runs. This page gives twelve questions to ask before you compare two published scores, and explains what a rating built from votes, such as Arena's, measures.
Common questions
What does a benchmark score in a model announcement tell me?
A benchmark score in a model announcement tells you how the model did on one public set of questions or tasks, on one date, under settings chosen by whoever ran the test. A 2025 review of 445 LLM benchmarks found that less than 10% used complete real-world tasks, so the score seldom describes a whole job. Use it to decide which models to try.
Which AI benchmarks am I most likely to see quoted?
The AI benchmarks quoted most often fall into three groups. SWE-bench covers coding, with 2,294 tasks in its original set according to its project site. MMLU, GPQA, ARC-AGI and Humanity's Last Exam cover knowledge and reasoning. GAIA, WebArena, tau-bench and METR's time horizon cover agents, which are AI systems that act through tools over many steps.
What are the three ways a benchmark score can mislead a buyer?
A benchmark score can mislead a buyer in three ways: the questions were in the model's training text, every strong model scores near the top, or the vendor chose favourable settings. For the second case, Stanford's AI Index 2026 reports that the 15 leading models on MMLU-Pro all scored above 87% as of early 2026, too close together to rank.
What does the history of SWE-bench Verified show about benchmark scores?
The history of SWE-bench Verified shows that a benchmark score has a limited useful life. OpenAI helped publish the 500-task set on 13 August 2024 and stopped reporting scores on it on 23 February 2026, 18 months later. By OpenAI's own account, higher scores had come to reflect how much a model was exposed to the benchmark during training. The post is the company's own audit.
Is a newer benchmark more reliable than the one it replaces?
A newer benchmark is harder for a time, and it can have faults of its own. OpenAI recommended SWE-bench Pro as a replacement for SWE-bench Verified in February 2026, then withdrew the recommendation on 8 July 2026 after its own audit counted 27.4% of tasks as broken by one method and 34.1% by another. Check the publication date and any audits before relying on a benchmark.
Can the same model get two different scores on the same benchmark?
Yes. The same model can score differently on one benchmark when the conditions change. OpenAI reported in August 2024 that GPT-4 scored between 2.7% and 28.3% on SWE-bench Lite depending on the scaffold, which is the program that hands the model tools and runs it step by step. Ask which attempts, tools and scaffold were used before comparing two published scores.
Do different sources agree on the same benchmark figure?
Different sources sometimes give different figures for the same benchmark. For SWE-bench Verified, OpenAI's February 2026 post gives 80.9% as the top score, and Stanford's AI Index 2026 gives 76.8% for the leader as of February 2026, worded as approximate. Quote a benchmark figure with its source and its date, and go to the benchmark's own paper for its size and method.
How long does a public benchmark stay useful for comparing models?
A public benchmark stays useful until the leading models all score near its maximum or its questions reach training data, and that can take under two years. SWE-bench Verified was published on 13 August 2024, and OpenAI stopped reporting it on 23 February 2026. Stanford's AI Index 2026 says tests meant to be challenging for years now reach top scores within months.
What should I do instead of comparing benchmark scores?
Run the candidate models on your own cases. Collect real inputs from your product, write down what a correct answer is, run each model with the same instructions, and compare pass rate, cost per task and response time. Google Cloud's documentation says results on your specific tasks give insights which cannot be derived from public leaderboards, meaning public ranking tables, and general benchmarks.
How much does it cost to test models on my own cases?
Testing models on your own cases costs little in model fees. For an example task of 3,000 input tokens and 500 output tokens, 200 cases on gpt-6-astra cost $11.00 at the price on OpenAI's pricing page on 30 September 2026. The larger cost is staff time: collecting real cases, agreeing on the expected answers and reading the failures.
Do I need a technical background to use this guide?
Readers need no technical background for this guide. Each page explains a benchmark in plain words, with its size and method taken from the benchmark's own paper, and defines each term where it first appears. One page is a printable list of questions to ask a vendor about any published score. The page on choosing a model gives the price arithmetic step by step.
References
- [1] OpenAI, Introducing SWE-bench Verified (13 August 2024), read from the Internet Archive copy of 15 September 2026: 500 samples, 93 developers screening 1,699 samples, 68.3% filtered out, GPT-4o at 33.2% on Verified against 16% on the original set, and GPT-4 at 2.7% to 28.3% on SWE-bench Lite by scaffold. OpenAI's own account.
- [2] OpenAI, Why SWE-bench Verified no longer measures frontier coding capabilities (23 February 2026), read from the Internet Archive copy of 15 September 2026: OpenAI stopped reporting the scores, top score 80.9%, and all tested leading models could reproduce reference fixes for certain tasks. OpenAI's own audit.
- [3] OpenAI, Separating signal from noise in coding evaluations (8 July 2026), read from the Internet Archive copy of 22 September 2026: recommendation of SWE-Bench Pro withdrawn; 200 tasks (27.4%) flagged broken by one method and 249 (34.1%) by human annotation. OpenAI's own audit.
- [4] Zhang and co-authors (Scale AI), A Careful Examination of Large Language Model Performance on Grade School Arithmetic (arXiv, 1 May 2024, NeurIPS 2024): 1,205 new problems written by people, accuracy drops of up to 8% for some model families, minimal signs of overfitting for many leading models.
- [5] Stanford HAI, AI Index Report 2026, Chapter 2: Technical Performance: the definition of saturation, the statement that evaluations are saturated in months, MMLU-Pro top 15 above 87% with just over 4 points from 1st to 15th, SWE-bench Verified leader at 76.8% as of February 2026, and Arena ratings from 1,424 to 1,503 as of March 2026.
- [6] Hendrycks and co-authors, Measuring Massive Multitask Language Understanding (arXiv, 7 September 2020, ICLR 2021): 57 subjects, 15,908 questions, 34.5% for untrained Amazon Mechanical Turk workers, and the largest GPT-3 model under 20 points above chance.
- [7] Measuring what Matters: Construct Validity in Large Language Model Benchmarks (arXiv, 3 November 2025, NeurIPS 2025): 445 benchmarks reviewed by 29 experts, less than 10% used complete real-world tasks, 16.0% used uncertainty estimates or statistical tests.
- [8] SWE-bench team, SWE-bench Leaderboards project site (read 30 September 2026): 2,294 instances from 12 Python repositories and the definition of % Resolved.
- [9] Gema and co-authors, Are We Done with MMLU? (arXiv, 6 June 2024): an estimated 6.49% of MMLU questions contain errors, from 5,700 re-annotated questions; 57% of the analysed Virology questions contain errors.
- [10] Yao, Shinn, Razavi and Narasimhan (Sierra), tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, 17 June 2024): 115 retail and 50 airline tasks, scoring by end database state, gpt-4o agents under 50% of tasks and under 25% on eight repeated retail trials.
- [11] University of Cambridge, Department of Applied Mathematics and Theoretical Physics, Goodhart's Law (undated page, read 30 September 2026): Goodhart as Chief Adviser to the Bank of England and Marilyn Strathern's restatement, which the page places in Strathern's 1997 article in the journal European Review.
- [12] Singh and co-authors, The Leaderboard Illusion (arXiv, 29 April 2025; several authors at Cohere): private testing of model variants before release, including 27 private variants tested by Meta before the Llama-4 release.
- [13] Arena, Our Response to 'The Leaderboard Illusion' Writeup (9 May 2025): the platform's reply that the paper contains several incorrect claims and that any provider can submit public and private variants. The platform's own account.
- [14] Google Cloud documentation, Gen AI evaluation service overview (read 30 September 2026): results on your specific tasks give insights which cannot be derived from public leaderboards and general benchmarks. A product page for a paid service.
- [15] OpenAI API documentation, Pricing (read 30 September 2026, standard tier, short context, US dollars per 1M tokens): gpt-6-astra $10.00 input and $50.00 output, gpt-6-luna $0.10 input and $0.50 output. Prices change.
Related reading
What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified
On 23 February 2026 OpenAI said it had stopped reporting SWE-bench Verified, and on 8 July 2026 it withdrew its recommendation of SWE-Bench Pro. This post gives OpenAI's figures with their limits, then shows how to choose a coding model on 100 of your own tasks.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.