The benchmarks vendors quote

SWE-bench explained: what a coding benchmark score means

SWE-bench is a public test of AI coding models. It gives a model a real bug report from an open software project and counts the task as solved when the project's tests pass after the model's change. The original set has 2,294 tasks from 12 Python projects. The score is the percentage of tasks solved, and it changes with the version of the test, the program wrapped around the model and the date. OpenAI stopped reporting the Verified version on 23 February 2026 after its own audit. A high score shows a model fixes public Python bugs that have tests. Your own requirements were outside the test.

Published September 30, 2026. Editorial.

Key takeaways

  • SWE-bench has 2,294 tasks taken from real issues in 12 public Python repositories, by its 2023 paper. A task counts as solved when the project's tests pass after the model's change.
  • SWE-bench Verified is a 500-task subset that OpenAI published in August 2024 after 93 developers screened 1,699 samples and 68.3% of them were removed, by OpenAI's own account.
  • The same model scores differently under different conditions: OpenAI reported GPT-4o at 16% on the original set and 33.2% on Verified, and GPT-4 between 2.7% and 28.3% on Lite depending on the program that ran it.
  • OpenAI stopped reporting Verified scores on 23 February 2026, saying 59.4% of 138 audited hard problems had faulty tests or descriptions, and withdrew its recommendation of SWE-Bench Pro on 8 July 2026.
  • A SWE-bench score covers bug fixes in public code that has tests. To learn whether a model can do your work, run it on tasks written from your own specification.

In October 2023 the best model tested on SWE-bench fixed 1.96% of the tasks [1]. On 23 February 2026 OpenAI wrote that the top score on the most quoted version of the benchmark had reached 80.9%, and in the same post it said it had stopped reporting that score and recommended that other model developers stop too [6]. Both facts describe the same test. This page explains what the tasks are and what the percentage counts, from the benchmark's own paper and project site, then gives the published criticism with its exact figures. It belongs to our guide to AI benchmarks vs your own evals.

What SWE-bench contains

A benchmark is a fixed, public set of tasks that many AI models are scored on, so that their results can be listed in one table. SWE-bench is a benchmark for coding work. Its tasks come from GitHub, a website where software teams store their code. On GitHub a written bug report or request is called an issue, and a proposed change to the code is called a pull request.

The SWE-bench paper, first published in October 2023, describes 2,294 software engineering problems drawn from real issues and their matching pull requests across 12 popular repositories written in Python, a programming language [1]. A repository is one project's complete folder of code together with its history. The authors started from 90,000 pull requests and filtered them down to the 2,294 tasks [1].

For each task the model receives the project's code and the text of the issue. Its job, in the paper's words, is "editing the codebase to address the issue" [1].

How a SWE-bench task is graded

Software projects include tests: small programs that run the code and check that it behaves as expected. SWE-bench uses those tests to grade the model. The model's change is applied to the project, the tests for that task are run, and the task counts as solved when they pass. OpenAI's description of the benchmark calls these unit tests, says they are not shown to the model, and says they mark a solution as correct or incorrect [2].

The published number is called "% Resolved", which the project site defines as "the percentage of task instances solved" [3]. A score of 50% on a 500-task set means that 250 tasks had passing tests after the model's change. One task on a 500-task set is worth 0.2 percentage points (1 divided by 500, written as a percentage).

Grading by tests has two consequences, and both appear in the audits further down this page. A change that passes weak tests counts as correct even when a reviewer would reject it. A correct change fails when the tests demand one particular way of writing the fix.

The versions of SWE-bench and what each one changed

"SWE-bench" in an announcement can mean any of seven sets. Each size below is the one its source states, read on 30 September 2026.

Version Size Who made it What changed
SWE-bench 2,294 tasks Jimenez and co-authors, 2023 [1] The original set, from 12 Python repositories
SWE-bench Lite 300 tasks The SWE-bench team [3] A smaller set, "curated for less costly evaluation"
SWE-bench Verified 500 tasks OpenAI with the SWE-bench team, August 2024 [2][3] Every task screened by human developers
Bash Only 500 tasks The SWE-bench team [3] The Verified tasks, with every model run in one fixed setup
SWE-bench Multilingual 300 tasks The SWE-bench team [3] 42 repositories across 9 programming languages
SWE-bench Multimodal 617 tasks (paper), 480 (site) Yang, Jimenez and co-authors, 2024 [10] Tasks from 17 JavaScript libraries, each with an image
SWE-Bench Pro 1,865 tasks Scale AI, September 2025 [8] A separate benchmark built from 41 repositories

The Multimodal paper counts 617 task instances [10] and the project site lists 480 [3], so quote each figure with its source. SWE-Bench Pro is absent from the SWE-bench project site. Its authors split it into a public set from 11 repositories, a held-out set (tasks kept back from the public) from 12 repositories, and a commercial set from 18 privately owned repositories [8].

Why SWE-bench Verified was made

OpenAI published SWE-bench Verified on 13 August 2024. Everything in this section is OpenAI's own account of its own work [2].

OpenAI says it worked with 93 software developers experienced in Python, who screened 1,699 random samples from the SWE-bench test set by hand. Each sample was labelled three times by separate people, and the most severe of the three labels was kept. The screening flagged 38.3% of samples for a problem statement that left out information a developer would need, and 61.1% for unit tests that could mark a valid solution as incorrect. In total 68.3% of the screened samples were removed. The Verified set is 500 of the samples that passed [2].

OpenAI reports that its GPT-4o model reached 33.2% on Verified against 16% on the original SWE-bench on the best-performing scaffold [2]. A scaffold is the program wrapped around the model that gives it tools and runs its steps one after another. Dividing 33.2 by 16 gives 2.075, so the same model scored more than twice as high after the tasks were cleaned.

What a SWE-bench percentage depends on

Three things change the number besides the model's ability.

The version. A score on Lite, Verified or the original set is a score on a different list of tasks.

The scaffold. OpenAI's 2024 post says GPT-4 scored 2.7% on SWE-bench Lite with an early scaffold and 28.3% with one called CodeR [2]. That is the same model, with a result 10.48 times higher (28.3 divided by 2.7). The Bash Only view on the project site exists for this reason: it puts "every model in the same mini-SWE-agent environment" [3]. A leaderboard, the public ranking table of scores, ranks a model together with its scaffold.

The date. On 5 August 2024 the top agents, meaning a model together with its scaffold, scored 20% on SWE-bench and 43% on Lite, according to the leaderboard as OpenAI quoted it [2]. For February 2026 two sources give two figures for Verified. Stanford's AI Index 2026 gives 76.8% for the leader as an approximate figure, with several other models between 70% and 76% [11]. OpenAI's post of 23 February 2026 says the best published score moved from 74.9% to 80.9% in the six months before it [6].

The settings a vendor chooses for a published score are the subject of how to read a model vendor's benchmark claims.

What the published audits found

Three pieces of criticism give exact figures, and each has a limited scope.

SWE-Bench+, October 2024. The authors read by hand the patches that one agent, SWE-Agent with GPT-4, had passed. A patch is the set of code changes a model proposes. They report that 32.67% of the successful patches had the solution given in the issue report or its comments, and that 31.08% passed because the tests were too weak to check the patch. With those cases removed, the agent's score fell from 12.47% to 3.97%. They add that over 94% of the issues were created before the models' knowledge cut-off dates, the dates after which a model has no training text [4]. These are shares of one agent's successful patches, and this page relies on the paper's abstract.

The SWE-Bench Illusion, June 2025. Given only the issue text, models named the file that held the bug with up to 76% accuracy on SWE-bench tasks, against up to 53% on repositories outside the benchmark. Word-for-word overlap with the reference fix (the human-written change stored as the correct answer), measured in runs of five consecutive pieces of text, reached up to 35% on SWE-bench Verified and the full set, against up to 18% on other benchmarks [5]. The authors read the gap as "possible data contamination or memorization" [5]. Contamination means the test material was in the text a model learned from. Our page on benchmark contamination covers the wider evidence.

OpenAI's audit, 23 February 2026. OpenAI audited 138 Verified problems that its o3 model failed to solve consistently over 64 runs, with at least six experienced engineers reviewing each one. It found that 59.4% of the 138 had material faults in the tests or the description: 35.5% had tests that demand one particular way of writing the code, 18.8% had tests that check for behaviour the description never asked for, and 5.1% had other faults [6]. The 138 problems are 27.6% of the 500 and were chosen because a model failed them, so 59.4% is the fault rate of a hard subset.

OpenAI also says every leading model it tested could reproduce the original human-written fix or exact details of the task text for certain tasks, after probing GPT-5.2-Chat, Claude Opus 4.5 and Gemini 3 Flash Preview. Its conclusion is that improvements on Verified "increasingly reflect how much the model was exposed to the benchmark at training time" [6]. This is OpenAI's audit of a benchmark it helped build.

What happened to SWE-Bench Pro

In the February 2026 post OpenAI recommended that developers report SWE-Bench Pro until better benchmarks existed [6]. On 8 July 2026 it withdrew that advice [7].

OpenAI's July audit gives 30% as its approximate estimate of the share of SWE-Bench Pro tasks that are broken. Two review routes produced two counts on the 731-task public set: an automated analysis flagged 200 tasks (27.4%) and a human review, with five engineers per flagged task, identified 249 (34.1%). OpenAI also reports that leading models went from a 23.3% pass rate to 80.3% on that public set in eight months [7].

The authors of SWE-Bench Pro describe it as "a contamination-resistant testbed" [8]. A preprint (a paper posted before review by other researchers) dated 8 September 2026, read as an abstract, reports leaked reference solutions and task quality faults in SWE-Bench Pro and says some models score lower on its cleaned version [9]. The sources for this page include no reply from Scale AI to OpenAI's audit, so the July figures are one company's estimate of another company's benchmark.

Vendors still report the benchmark. Anthropic's system card for Claude Opus 5.5 is dated 22 September 2026, which is 76 days after OpenAI's withdrawal, and it reports a SWE-bench Pro result for Anthropic's own model [12].

What SWE-bench leaves out

Other languages. The original set and Verified are Python only [1]. When the Multimodal set, built from projects in the JavaScript programming language, was published in October 2024, the best system resolved 12% of them and the next best 6% [10].

Private code. Every task comes from a public repository. No task contains your codebase or its conventions.

Vague requests. Verified removed tasks whose description was incomplete. A real request often arrives incomplete, and working out what was meant is part of the job.

Work that has no test. A task can only be graded when tests exist for it. Choices about design, wording on a screen, or whether a feature should exist at all fall outside the score.

Large pieces of work. OpenAI labels 196 of the 500 Verified tasks, which is 39.2%, as fixes that take a person under 15 minutes, and 45 as taking over an hour [2]. A product is many such changes that have to agree with each other.

Does a high score mean a model can build your software

A SWE-bench score is evidence about one model, inside one scaffold, on bug fixes in public Python projects that have tests, on one date. It says little about your product, because none of your requirements were in the test.

The measurement that answers your question is an eval: a test you write for your own work. How to choose an AI model with your own evals gives the steps, and how to write your first eval suite covers the coding case. For a supplier's own suite, how to evaluate a vendor's eval suite lists what to ask.

How Reveneau applies this

All of Reveneau's code is written by AI, and every change must pass a large eval suite before release. We write that suite from the client's specification before any code exists. That order is the method described in our guide to eval-driven development.

SWE-bench tests a model on other people's repositories, against tests those projects happened to have. A Reveneau eval suite tests the change the client asked for, against checks written for that request. When a new model is announced with a higher SWE-bench score, the question we ask is whether it passes the suite for the client's project.

We grade the judged checks in that suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. Because AI does work that would otherwise need more engineers, a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release. To see how this applies to your build, read about our AI development service.

Common questions

What is SWE-bench?

SWE-bench is a public benchmark that tests whether an AI model can fix real software problems. Its 2023 paper describes 2,294 tasks taken from real issues and their matching pull requests in 12 popular Python repositories on GitHub. The model receives the project's code and the text of the issue, and has to edit the code so that the problem is resolved. The published score is the percentage of tasks solved.

What is SWE-bench Verified?

SWE-bench Verified is a 500-task subset of SWE-bench that OpenAI published on 13 August 2024. By OpenAI's account, 93 developers experienced in Python screened 1,699 samples by hand, each sample was labelled three times, and 68.3% of the screened samples were removed for unclear descriptions or unfair tests. OpenAI stopped reporting scores on Verified on 23 February 2026.

What does a SWE-bench score mean?

A SWE-bench score is the percentage of tasks in the set that were solved, which the project site calls % Resolved. On the 500-task Verified set, 50% means 250 tasks had passing tests after the model's change, and one task is worth 0.2 percentage points. The number describes a model together with the program wrapped around it, on one version of the test, on one date.

How is SWE-bench graded?

SWE-bench is graded by running tests. The model's change is applied to the project, the tests for that task are run, and the task counts as solved when they pass. OpenAI's 2024 post describes them as unit tests that mark a solution correct or incorrect. Its screening flagged 61.1% of 1,699 samples for tests that could mark a valid solution as incorrect, which is why grading quality matters.

Does a high SWE-bench score mean a model can build my software?

A high SWE-bench score shows that a model, with the program that runs it, fixes bugs in public Python projects that have tests. Building your software also needs work the benchmark leaves out: private code, incomplete requests, design choices that no test checks, and many changes that must agree with each other. By OpenAI's labels, 39.2% of Verified tasks take a person under 15 minutes. Test the model on your own tasks.

What are the problems with SWE-bench?

Published audits of SWE-bench report two kinds of problem: faulty tasks and leaked answers. The SWE-Bench+ paper found that 32.67% of one agent's successful patches had the solution written in the issue, and 31.08% passed on weak tests. OpenAI's February 2026 audit found material faults in 59.4% of 138 hard Verified problems, and said every leading model it tested could reproduce reference fixes for certain tasks.

Why did OpenAI stop reporting SWE-bench Verified scores?

OpenAI said on 23 February 2026 that SWE-bench Verified no longer measured real progress in coding ability. Its audit of 138 problems that its o3 model often failed found faulty tests or descriptions in 59.4% of them, and it reported that the models it probed could reproduce original fixes for certain tasks. OpenAI concluded that gains increasingly reflect exposure to the benchmark during training. These are OpenAI's own findings.

What is SWE-Bench Pro, and is it more reliable than SWE-bench Verified?

SWE-Bench Pro is a separate coding benchmark from Scale AI with 1,865 problems from 41 repositories, part of them kept private. OpenAI recommended it in February 2026 and withdrew that recommendation on 8 July 2026, estimating that 30% of its tasks are broken: an automated analysis flagged 27.4% of the 731 public tasks and a human review flagged 34.1%. The sources read for this page include no reply from Scale AI.

Does SWE-bench cover programming languages other than Python?

The original SWE-bench and the Verified subset contain Python tasks only. Two later sets widen the range. SWE-bench Multilingual has 300 tasks from 42 repositories across 9 programming languages, by the project site. SWE-bench Multimodal uses 17 JavaScript libraries and includes an image in every task; its paper counts 617 tasks and the project site lists 480. Check which set a quoted score refers to.

Why does the same model get different SWE-bench scores from different testers?

The same model gets different SWE-bench scores because the score includes the scaffold, the program that gives the model tools and runs its steps. OpenAI's 2024 post reports GPT-4 at 2.7% on SWE-bench Lite with an early scaffold and 28.3% with another, a result 10.48 times higher. The Bash Only view on the SWE-bench site runs every model in one fixed setup to remove that difference.

Is there a cheaper version of SWE-bench to run?

SWE-bench Lite is the smaller set, with 300 tasks, and the project site describes it as curated for less costly evaluation. Running fewer tasks costs less and gives a less precise number: on 300 tasks, one task is a third of a percentage point. A score on Lite cannot be compared with a score on Verified or on the full 2,294-task set, because each is a different list of tasks.

What should I do when a vendor quotes a SWE-bench score to me?

When a vendor quotes a SWE-bench score, ask which version it used, which program ran the model, and on what date. Then ask for evidence on your own work. Pick real tasks from your project, write down what a correct result is, and have each candidate model attempt them under the same conditions. That small eval answers the question a SWE-bench score leaves open: whether the model can do your tasks.

References

More in The benchmarks vendors quote

MMLU, GPQA and ARC-AGI explained: knowledge and reasoning benchmarks

MMLU, GPQA, ARC-AGI and Humanity's Last Exam are public tests that model makers quote to show knowledge and reasoning. MMLU has 15,908 multiple-choice questions in 57 subjects, and the authors of a newer test write that models now score over 90% on it. GPQA has 448 graduate-level science questions on which experts scored 65%. ARC-AGI uses puzzles that ask the solver to work out a new rule. Humanity's Last Exam has 2,500 questions written by experts. Each score describes performance on that test's own questions. For a product that answers customers from your own documents, these scores say little, because none of them tests your documents, your search step or your customers' questions.

Agent benchmarks explained: GAIA, WebArena, tau-bench and METR time horizons

Agent benchmarks score whether an AI model can finish a task that takes several steps and tools. GAIA has 466 questions, WebArena has 812 website tasks, tau-bench has 165 customer service tasks and OSWorld has 369 desktop tasks. The METR time horizon reports the length of task, measured by how long it takes a skilled person, that a model completes half the time. At publication the best models scored between 12.24% and 61.2% on these tests, and the AI Index 2026 reports later best scores between 66.3% and 74.5%. Each score describes that benchmark's tasks on one date and one version. It leaves out repeat reliability, cost, and your own tools and rules, so use it to pick candidates and then test your own agent.