What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified

On 23 February 2026 OpenAI published a post that contains this sentence: "This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too." SWE-bench Verified is a coding test that OpenAI itself published on 13 August 2024, 18 months earlier. In the same February post OpenAI recommended a replacement test, SWE-Bench Pro. On 8 July 2026, which is 135 days later, it wrote: "we retract our earlier recommendation to adopt SWE-Bench Pro."
If you compare AI coding models by their published scores, these two posts change what that comparison can tell you. This post gives what OpenAI said, with its exact figures and their limits, and then argues one position: choose a coding model on a small set of your own tasks, and use published scores only to decide which models to test. OpenAI's website refuses automated requests, so we read the three OpenAI posts cited here in the copies kept by the Internet Archive, and every link to them goes to that copy.
What SWE-bench is, in everyday words
A benchmark is a fixed, public set of tasks that many AI models are scored on, so that their results can be put in one table. SWE-bench is a benchmark for coding work. Its 2023 paper describes 2,294 software problems taken from real issues and their matching pull requests in 12 popular Python repositories on GitHub. GitHub is a website where software teams store their code. An issue is a written bug report or request, a pull request is a proposed change to the code, and a repository is one project's folder of code with its history. Python is a programming language.
For each task the model receives the project's code and the text of the issue. Its job, in the paper's words, is "editing the codebase to address the issue". The SWE-bench project site publishes the result as "% Resolved", which it defines as "the percentage of task instances solved".
SWE-bench Verified is a 500-task part of that set. By OpenAI's account in its August 2024 post, 93 developers experienced in Python checked 1,699 samples by hand, and 68.3% of the checked samples were removed. The same post says 61.1% of the samples were marked because their unit tests, the small programs that check whether code behaves as expected, could mark a valid solution as incorrect.
What OpenAI said on 23 February 2026
Everything in this section is OpenAI's own statement, in its post of 23 February 2026, about a benchmark it helped to build.
OpenAI audited 138 Verified problems that its o3 model did not solve consistently over 64 runs. At least six experienced software engineers reviewed each problem. OpenAI found that 59.4% of the 138 problems "contained material issues in test design and/or problem description". It gives three parts: 35.5% had tests that demand one particular way of writing the code, 18.8% had tests that check for behaviour the description never asked for, and 5.1% had other faults. The three parts add up to 59.4.
The 138 problems are 27.6% of the 500, and OpenAI chose them because a model often failed them. The 59.4% is therefore the fault rate of a difficult subset. OpenAI gives no fault rate for the full 500, and a reader should avoid calculating one.
OpenAI also reported contamination, which means that the test material was in the text a model learned from. It says all the leading models it tested could reproduce the original human-written fix, or exact details of the task text, for certain tasks. It names the three models it tested: GPT-5.2-Chat, which is its own, Claude Opus 4.5 and Gemini 3 Flash Preview. The post publishes no rate per model, so it supports no ranking of the three.
Its conclusion is that improvements on the test "increasingly reflect how much the model was exposed to the benchmark at training time". It adds that the best published score had risen from 74.9% to 80.9% in six months. That is 6.0 percentage points. One task on a 500-task set is worth 0.2 points, so the gain equals 30 tasks (6.0 divided by 0.2).
Two sources disagree on that top score. Stanford's AI Index Report 2026 gives 76.8% for the leading model as of February 2026, a figure the report words as approximate. OpenAI gives 80.9%. Quote either one with its source and date.
Researchers outside OpenAI had published related findings earlier. The SWE-Bench+ paper of October 2024 examined the tasks one AI system had passed and reported that in 32.67% of them the solution was written in the issue text or its comments. The SWE-Bench Illusion paper of June 2025 found that models named the file holding the bug from the issue text alone with up to 76% accuracy on SWE-bench tasks, against up to 53% on projects outside the benchmark. Its authors read the gap as "possible data contamination or memorization". We read both papers as abstracts, and our guide page on benchmark contamination covers the wider evidence.
What OpenAI said on 8 July 2026
SWE-Bench Pro is a separate benchmark built by Scale AI. Its paper of September 2025 describes 1,865 problems from 41 repositories, and its authors call it "a contamination-resistant testbed".
OpenAI's post of 8 July 2026 reports an audit of that benchmark. OpenAI's own estimate is that 30% of the tasks are "broken", a figure it words as approximate. Two review methods produced two counts on the 731 tasks that are public. An automated analysis counted 200 broken tasks, which is 27.4%. A human review, with five engineers for each task it examined, counted 249, which is 34.1%. OpenAI's estimate sits between the two counts. The post also says leading models went from a 23.3% pass rate to 80.3% on those 731 tasks in eight months.
This is one company's audit of another company's benchmark, and the sources we read include no reply from Scale AI.
Another vendor made a different decision about the same test. Anthropic's system card for Claude Opus 5.5, dated 22 September 2026, reports a SWE-bench Pro result for Anthropic's own model. That date is 76 days after OpenAI's withdrawal. The card's results table has no row for SWE-bench Verified. Each vendor decides which tests to publish. The effect for you is practical: two vendors' tables may hold different tests, so you cannot read them side by side.
Why this matters to a buyer of AI coding models
A purchase lasts longer than these scores did. OpenAI reported Verified for 18 months and recommended its replacement for 135 days. A buyer who chose a model in January 2026 partly on a Verified score was told in February, by the company that published the test, that gains on it increasingly reflect exposure to the test during training.
A published score also covers other people's code. Every SWE-bench task comes from a public repository, and the original set and Verified contain Python only. Your codebase, your conventions and your requirements were outside the test. Our guide to AI benchmarks vs your own evals explains this for each benchmark vendors quote, and SWE-bench explained lists every version of this one.
The vendors give the same advice in their own documentation. An eval is a test you write for your own work. OpenAI's documentation says: "Design task-specific evals: Make tests reflect model capability in real-world distributions." Anthropic's documentation says: "Be task-specific: Design evals that mirror your real-world task distribution." In plain words, both tell the customer to build the test from the work the customer's product does.
How to choose on 100 of your own cases
The method has five parts, and choosing a model with your own evals gives it in full.
- Write down what a correct result is before you look at any model's output.
- Collect 100 to 200 real tasks from your own project, each with its expected result.
- Run every candidate model on all of them, with the same instructions.
- Record three numbers for each model: the share of cases that pass, the cost per task and the time per task.
- Read the failures before you decide.
Here is an invented example, with invented numbers. A team takes 100 past tasks from its own project. Model A passes 80 and model B passes 84. A pass rate from a sample has a margin of error. For 80% on 100 cases, 0.8 times 0.2 divided by 100 is 0.0016, the square root is 0.04, or 4.00 points, and 1.96 times that is 7.84 points. The gap between the models is 4 points, which is inside the margin, so the two totals alone cannot rank them.
Evan Miller's 2024 paper on statistics for evals, written at Anthropic, recommends comparing two models on the differences question by question. In the invented example both models pass 75 cases and both fail 11. Model A alone passes 5, and model B alone passes 9. Those four numbers add up to 100, and the 14 cases where the models differ are the ones to read. Suppose the 5 that only model A passed are the billing rules, and the 9 that only model B passed are changes to wording on a screen. The team would choose model A, the one with the lower total.
Your cases have one more property. They have never been published, so no model can have seen them in its training text. Keep them private, and run them again each time a vendor announces a new model or a new score.
What to do with a published score from now on
Use a published score to pick two to four candidates, and let your own run decide among them.
When a vendor quotes a coding score, ask which version of the test it used, which program ran the model, who ran it and on what date. The program matters: OpenAI's 2024 post reports one model, GPT-4, at 2.7% and at 28.3% on SWE-bench Lite, a 300-task version, depending on the program that ran it. How to read a model vendor's benchmark claims has the full list of questions.
Then ask the supplier to run its model on your tasks. Our earlier posts on how to evaluate a vendor's eval suite and on asking an AI vendor for time to production cover that conversation.
Credit is due to OpenAI for publishing an audit of a test it helped to build, with the sample size and the selection rule stated.
Where Reveneau fits
Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass a large eval suite before release. We write that suite from the client's specification before any code exists, so the test describes the client's work. When a vendor announces a model with a higher coding score, the question we ask is whether that model passes the suite for the client's project.
We grade the judged checks in the suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. Read about our AI development service.
Sources
- Why SWE-bench Verified no longer measures frontier coding capabilities: OpenAI, 23 February 2026. The archived copy of 15 September 2026 was read, because the original page refuses automated requests. OpenAI's own audit of 138 problems, the 59.4% figure and its parts, the 27.6% subset, the three models tested, and 74.9% to 80.9%.
- Separating signal from noise in coding evaluations: OpenAI, 8 July 2026. The archived copy of 22 September 2026 was read. OpenAI's own estimate that 30% of SWE-Bench Pro tasks are broken, the counts of 200 (27.4%) and 249 (34.1%) on 731 public tasks, and the withdrawn recommendation.
- Introducing SWE-bench Verified: OpenAI, 13 August 2024. The archived copy of 15 September 2026 was read. 500 tasks, 93 developers, 1,699 samples, 68.3% removed, 61.1% marked for their tests, and GPT-4 at 2.7% and 28.3% on SWE-bench Lite. OpenAI's own account.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?: Jimenez and co-authors, arXiv, 10 October 2023. 2,294 problems from 12 Python repositories and the task definition.
- SWE-bench Leaderboards: the SWE-bench team's project site, read 30 September 2026. The definition of % Resolved and the size of each set.
- AI Index Report 2026, Chapter 2: Technical Performance: Stanford Institute for Human-Centered AI, 2026 edition. The SWE-bench Verified leader at 76.8% as of February 2026.
- SWE-Bench+: Enhanced Coding Benchmark for LLMs: Aleithan and co-authors, arXiv, 9 October 2024, abstract. 32.67% of one system's successful patches had the solution in the issue.
- The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason: Liang, Garg and Moghaddam, arXiv, 14 June 2025, abstract. Up to 76% against up to 53% for naming the file with the bug.
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?: Deng, Da and co-authors at Scale AI, arXiv, 21 September 2025, abstract. 1,865 problems from 41 repositories, and the authors' description of the benchmark.
- System Card: Claude Opus 5.5: Anthropic, 22 September 2026. Anthropic's own report, which includes a SWE-bench Pro result and has no SWE-bench Verified row in its results table.
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations: Evan Miller at Anthropic, arXiv, 1 November 2024. The recommendation to compare two models question by question.
- Evaluation best practices: OpenAI's API documentation, read 30 September 2026.
- Define success criteria and build evaluations: Anthropic's Claude Platform documentation, read 30 September 2026.


