Strategy

What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified

Editorial · Reveneau · October 4, 2026

What a buyer should do now that OpenAI has stopped reporting SWE-bench Verified

On 23 February 2026 OpenAI published a post that contains this sentence: "This is why we have stopped reporting SWE-bench Verified scores, and we recommend that other model developers do so too." SWE-bench Verified is a coding test that OpenAI itself published on 13 August 2024, 18 months earlier. In the same February post OpenAI recommended a replacement test, SWE-Bench Pro. On 8 July 2026, which is 135 days later, it wrote: "we retract our earlier recommendation to adopt SWE-Bench Pro."

If you compare AI coding models by their published scores, these two posts change what that comparison can tell you. This post gives what OpenAI said, with its exact figures and their limits, and then argues one position: choose a coding model on a small set of your own tasks, and use published scores only to decide which models to test. OpenAI's website refuses automated requests, so we read the three OpenAI posts cited here in the copies kept by the Internet Archive, and every link to them goes to that copy.

What SWE-bench is, in everyday words

A benchmark is a fixed, public set of tasks that many AI models are scored on, so that their results can be put in one table. SWE-bench is a benchmark for coding work. Its 2023 paper describes 2,294 software problems taken from real issues and their matching pull requests in 12 popular Python repositories on GitHub. GitHub is a website where software teams store their code. An issue is a written bug report or request, a pull request is a proposed change to the code, and a repository is one project's folder of code with its history. Python is a programming language.

For each task the model receives the project's code and the text of the issue. Its job, in the paper's words, is "editing the codebase to address the issue". The SWE-bench project site publishes the result as "% Resolved", which it defines as "the percentage of task instances solved".

SWE-bench Verified is a 500-task part of that set. By OpenAI's account in its August 2024 post, 93 developers experienced in Python checked 1,699 samples by hand, and 68.3% of the checked samples were removed. The same post says 61.1% of the samples were marked because their unit tests, the small programs that check whether code behaves as expected, could mark a valid solution as incorrect.

What OpenAI said on 23 February 2026

Everything in this section is OpenAI's own statement, in its post of 23 February 2026, about a benchmark it helped to build.

OpenAI audited 138 Verified problems that its o3 model did not solve consistently over 64 runs. At least six experienced software engineers reviewed each problem. OpenAI found that 59.4% of the 138 problems "contained material issues in test design and/or problem description". It gives three parts: 35.5% had tests that demand one particular way of writing the code, 18.8% had tests that check for behaviour the description never asked for, and 5.1% had other faults. The three parts add up to 59.4.

The 138 problems are 27.6% of the 500, and OpenAI chose them because a model often failed them. The 59.4% is therefore the fault rate of a difficult subset. OpenAI gives no fault rate for the full 500, and a reader should avoid calculating one.

OpenAI also reported contamination, which means that the test material was in the text a model learned from. It says all the leading models it tested could reproduce the original human-written fix, or exact details of the task text, for certain tasks. It names the three models it tested: GPT-5.2-Chat, which is its own, Claude Opus 4.5 and Gemini 3 Flash Preview. The post publishes no rate per model, so it supports no ranking of the three.

Its conclusion is that improvements on the test "increasingly reflect how much the model was exposed to the benchmark at training time". It adds that the best published score had risen from 74.9% to 80.9% in six months. That is 6.0 percentage points. One task on a 500-task set is worth 0.2 points, so the gain equals 30 tasks (6.0 divided by 0.2).

Two sources disagree on that top score. Stanford's AI Index Report 2026 gives 76.8% for the leading model as of February 2026, a figure the report words as approximate. OpenAI gives 80.9%. Quote either one with its source and date.

Researchers outside OpenAI had published related findings earlier. The SWE-Bench+ paper of October 2024 examined the tasks one AI system had passed and reported that in 32.67% of them the solution was written in the issue text or its comments. The SWE-Bench Illusion paper of June 2025 found that models named the file holding the bug from the issue text alone with up to 76% accuracy on SWE-bench tasks, against up to 53% on projects outside the benchmark. Its authors read the gap as "possible data contamination or memorization". We read both papers as abstracts, and our guide page on benchmark contamination covers the wider evidence.

What OpenAI said on 8 July 2026

SWE-Bench Pro is a separate benchmark built by Scale AI. Its paper of September 2025 describes 1,865 problems from 41 repositories, and its authors call it "a contamination-resistant testbed".

OpenAI's post of 8 July 2026 reports an audit of that benchmark. OpenAI's own estimate is that 30% of the tasks are "broken", a figure it words as approximate. Two review methods produced two counts on the 731 tasks that are public. An automated analysis counted 200 broken tasks, which is 27.4%. A human review, with five engineers for each task it examined, counted 249, which is 34.1%. OpenAI's estimate sits between the two counts. The post also says leading models went from a 23.3% pass rate to 80.3% on those 731 tasks in eight months.

This is one company's audit of another company's benchmark, and the sources we read include no reply from Scale AI.

Another vendor made a different decision about the same test. Anthropic's system card for Claude Opus 5.5, dated 22 September 2026, reports a SWE-bench Pro result for Anthropic's own model. That date is 76 days after OpenAI's withdrawal. The card's results table has no row for SWE-bench Verified. Each vendor decides which tests to publish. The effect for you is practical: two vendors' tables may hold different tests, so you cannot read them side by side.

Why this matters to a buyer of AI coding models

A purchase lasts longer than these scores did. OpenAI reported Verified for 18 months and recommended its replacement for 135 days. A buyer who chose a model in January 2026 partly on a Verified score was told in February, by the company that published the test, that gains on it increasingly reflect exposure to the test during training.

A published score also covers other people's code. Every SWE-bench task comes from a public repository, and the original set and Verified contain Python only. Your codebase, your conventions and your requirements were outside the test. Our guide to AI benchmarks vs your own evals explains this for each benchmark vendors quote, and SWE-bench explained lists every version of this one.

The vendors give the same advice in their own documentation. An eval is a test you write for your own work. OpenAI's documentation says: "Design task-specific evals: Make tests reflect model capability in real-world distributions." Anthropic's documentation says: "Be task-specific: Design evals that mirror your real-world task distribution." In plain words, both tell the customer to build the test from the work the customer's product does.

How to choose on 100 of your own cases

The method has five parts, and choosing a model with your own evals gives it in full.

  1. Write down what a correct result is before you look at any model's output.
  2. Collect 100 to 200 real tasks from your own project, each with its expected result.
  3. Run every candidate model on all of them, with the same instructions.
  4. Record three numbers for each model: the share of cases that pass, the cost per task and the time per task.
  5. Read the failures before you decide.

Here is an invented example, with invented numbers. A team takes 100 past tasks from its own project. Model A passes 80 and model B passes 84. A pass rate from a sample has a margin of error. For 80% on 100 cases, 0.8 times 0.2 divided by 100 is 0.0016, the square root is 0.04, or 4.00 points, and 1.96 times that is 7.84 points. The gap between the models is 4 points, which is inside the margin, so the two totals alone cannot rank them.

Evan Miller's 2024 paper on statistics for evals, written at Anthropic, recommends comparing two models on the differences question by question. In the invented example both models pass 75 cases and both fail 11. Model A alone passes 5, and model B alone passes 9. Those four numbers add up to 100, and the 14 cases where the models differ are the ones to read. Suppose the 5 that only model A passed are the billing rules, and the 9 that only model B passed are changes to wording on a screen. The team would choose model A, the one with the lower total.

Your cases have one more property. They have never been published, so no model can have seen them in its training text. Keep them private, and run them again each time a vendor announces a new model or a new score.

What to do with a published score from now on

Use a published score to pick two to four candidates, and let your own run decide among them.

When a vendor quotes a coding score, ask which version of the test it used, which program ran the model, who ran it and on what date. The program matters: OpenAI's 2024 post reports one model, GPT-4, at 2.7% and at 28.3% on SWE-bench Lite, a 300-task version, depending on the program that ran it. How to read a model vendor's benchmark claims has the full list of questions.

Then ask the supplier to run its model on your tasks. Our earlier posts on how to evaluate a vendor's eval suite and on asking an AI vendor for time to production cover that conversation.

Credit is due to OpenAI for publishing an audit of a test it helped to build, with the sample size and the selection rule stated.

Where Reveneau fits

Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass a large eval suite before release. We write that suite from the client's specification before any code exists, so the test describes the client's work. When a vendor announces a model with a higher coding score, the question we ask is whether that model passes the suite for the client's project.

We grade the judged checks in the suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. Read about our AI development service.

Sources

Common questions

Why did OpenAI stop reporting SWE-bench Verified scores?

OpenAI said on 23 February 2026 that higher scores on SWE-bench Verified had stopped showing real gains in coding ability. Its audit of 138 problems found faults in the tests or the description in 59.4% of them, and it reported that the leading models it tested could reproduce the original human-written fix for certain tasks. These are OpenAI's own findings about a test it published in August 2024.

Does OpenAI's 59.4% figure mean that most SWE-bench Verified tasks are faulty?

The 59.4% figure applies only to the 138 problems OpenAI audited, which are 27.6% of the 500 tasks in SWE-bench Verified. OpenAI chose those problems because its o3 model failed to solve them consistently over 64 runs, so they are a difficult subset. OpenAI's post of 23 February 2026 gives no fault rate for the full set of 500, and a reader should avoid calculating one from it.

Why did OpenAI withdraw its recommendation of SWE-Bench Pro?

OpenAI withdrew the recommendation on 8 July 2026 after auditing SWE-Bench Pro, a coding benchmark built by Scale AI. OpenAI's own estimate is that 30% of the tasks are broken. Two review methods gave two counts on the 731 public tasks: an automated analysis counted 200 (27.4%) and a human review counted 249 (34.1%). The sources read for this post include no reply from Scale AI.

Do other AI vendors still report SWE-Bench Pro scores?

At least one does. Anthropic's system card for Claude Opus 5.5, dated 22 September 2026, reports a SWE-bench Pro result for Anthropic's own model, 76 days after OpenAI withdrew its recommendation of that benchmark. The card's results table has no row for SWE-bench Verified. Each vendor decides which tests to publish, so two vendors' tables may hold different tests, and a buyer cannot read them side by side.

What was the highest SWE-bench Verified score in February 2026?

Two sources give two figures for February 2026. OpenAI's post of 23 February 2026 says the best published score on SWE-bench Verified rose from 74.9% to 80.9% in six months. Stanford's AI Index Report 2026 gives 76.8% for the leading model, a figure the report words as approximate, with several other models between 70% and 76%. Quote either figure with its source and its date.

Should a buyer ignore published coding benchmark scores?

A buyer should use published coding scores for one job: choosing which two to four models to test. OpenAI's 2024 post reports one model, GPT-4, at 2.7% and at 28.3% on SWE-bench Lite depending on the program that ran it, so a score describes a model plus its settings. The decision between the candidates should come from a run on the buyer's own tasks.

How many of my own cases do I need to compare two coding models?

Start with 100 to 200 real cases from your own project. With 100 cases and a pass rate of 80%, the margin of error is 7.84 percentage points, so two models at 80% and 84% cannot be ranked by their totals. Evan Miller's 2024 paper on eval statistics recommends comparing two models question by question, so list each case where one model passed and the other failed.

What should I ask a vendor that quotes a SWE-bench score?

Ask the vendor which version of SWE-bench it used, which program ran the model, who ran the test and on what date. The SWE-bench project site lists sets of 2,294, 500 and 300 tasks, and each is a different list. Then ask for a run on your own tasks, with the pass rule written down by you before the vendor's model has produced any output.

How long does a published benchmark recommendation stay valid?

The public record gives two intervals, both from OpenAI. OpenAI published SWE-bench Verified on 13 August 2024 and said it had stopped reporting it on 23 February 2026, 18 months later. It recommended SWE-Bench Pro in that February post and withdrew the recommendation on 8 July 2026, after 135 days. If your contract for a model runs longer than either interval, keep a test set of your own cases.