How autonomous are AI coding agents, really?

In January, the head of Claude Code at Anthropic said he had not hand-edited a line of code in two months. His exact words to Fortune were "I shipped 22 PRs yesterday and 27 the day before, each one 100% written by Claude." An OpenAI researcher said the same thing in fewer words: "100%, I don't write code anymore." Anthropic's spokesperson put the company-wide figure at somewhere between 70 and 90 percent, and Fortune's reporting notes that is well ahead of Microsoft, which has publicly given a figure around 30 percent (Fortune, January 2026). These are the companies' own accounts of their own work, so take them as claims rather than as measurements. They are still the most direct evidence available.
Now read the first quote again and notice what is still in it. Pull requests. Twenty-two of them, opened by a person who decided they were ready.
That detail describes the current situation. Models now write the code, and people are still accountable for it. We work inside this every day, since every line of code we release is written by a model, and the gap between those two things is what our entire process is designed for. Here is what the 2026 evidence actually supports, in three parts.
1. "The model wrote all of it" is a claim about typing, not about judgment
A pull request is a request. It is the moment the writing stops and someone decides whether the work is good enough to become part of the product. When an engineer reports 22 of them in a day, the interesting number is 22, not 100 percent. The model now does the typing, and people still do the deciding.
This is why the headline percentages are less informative than they sound. A team can go from 30 percent AI-written to 90 percent AI-written without changing a single thing about who is responsible when a change breaks. What changes is the ratio of hours spent producing to hours spent judging, and that ratio reverses fast. Producing a change is now close to free. Deciding whether it is correct costs exactly what it always did, and you now have to do it far more often.
So the useful question for your own team is not what share of your code a model wrote. It is what happens between the model finishing and the change reaching a customer, and whether that step grew at the same rate the writing did. In most organisations it did not, which is the slowest step people are describing when they say AI has not made them faster.
2. A five-hour time horizon means a 50 percent success rate
The best public measurement of agent autonomy is METR's time horizon work. Their updated methodology, published in January 2026, puts Claude Opus 4.5 at a 320-minute time horizon, with GPT-5 at 214 minutes and o3 at 121. The measured doubling time since 2023 is about 131 days (METR, Time Horizon 1.1). Just over five hours, doubling roughly twice a year. It is a genuinely striking trend and it is easy to read as "an agent can now work unattended for most of a day."
METR themselves say that reading is wrong. A week before publishing those figures they put out a note clarifying what the metric does and does not mean, and the first line of it is blunt: "Time horizon is not the length of time AIs can work independently." What it measures is the amount of serial human labour a model can replace at a 50 percent success rate (METR, January 2026).
Fifty percent. The task fails half the time. Almost nothing you would actually delegate is worth doing at those odds, and METR state that too: a 50 percent time horizon "does not mean we can delegate tasks under X hours to AIs," because many tasks need success rates above 98 percent before automating them makes any sense.
If you want the number you would plan against, it is the 80 percent horizon, and it is much shorter. In METR's original study, Claude 3.7 Sonnet had the longest 80 percent horizon of any model examined at around 15 minutes, against a 50 percent horizon of 59 minutes (METR, March 2025). Roughly four times shorter. The doubling rate is about the same at both thresholds, so the gap stays as models improve and grows in proportion with them.
Two more caveats from the same researchers, both of which matter more than the headline. The error bars span about a factor of two in each direction, so precise comparisons between models are not meaningful. And the tasks in the suite are cleaner than real work. METR scored their tasks for "messiness," meaning things like unclear success criteria, under-specified scope, and information the agent has to go and find rather than being handed. Models perform worse on messier tasks than task length alone predicts, at roughly eight percentage points of success rate per point on that scale. Your backlog is messy. That is what a backlog is.
The practical version of all this is simple. Break work into units small enough that failure is cheap to notice and cheap to redo. A four-hour task handed over whole is a four-hour task where you find out at the end that step three was wrong.
3. Both of the reassuring numbers are weaker than they look
There are two figures people use to argue that agents are already doing real engineering. Benchmark scores, and the share of agent-written pull requests that get merged. Both are weaker than their headline.
Start with benchmarks. SWE-bench Verified is a human-filtered subset of 500 real GitHub issues, built with OpenAI, and it is the closest thing the field has to a standard (SWE-bench). Its scores are also measurably too high. A peer-reviewed study of patches from three leading issue-solving tools found that weaknesses in the validation mechanism cause 7.8 percent of all patches to count as correct while failing the project's own developer-written test suite. Beyond that, 29.6 percent of plausible patches behave differently from the ground truth fix, and manual inspection found 28.6 percent of those to be certainly incorrect. Combined, the authors put the inflation of reported resolution rates at 6.2 absolute percentage points (Are "Solved Issues" in SWE-bench Really Solved Correctly?). And a benchmark issue arrives already scoped, already reproduced, already agreed to be worth fixing. That framing work is usually the hard part.
Now the merged pull requests, which is the more interesting number. Researchers found 567 pull requests explicitly marked as generated with Claude Code, across 157 open-source projects with at least ten stars, and tracked what happened to them. 83.8 percent were merged. That sounds like a settled argument until you read the second figure: only 54.9 percent of the merged ones went in without further modification (On the Use of Agentic Coding, arXiv). Nearly half needed a human to change something before it was ready to merge.
There is a larger problem with reading that study as evidence of autonomy, and it is not a criticism of the study. Every one of those 567 pull requests exists because a developer looked at what the agent produced and decided to open it. The runs that went badly never became pull requests at all. A sample selected by human approval cannot tell you the unsupervised failure rate, because human approval is the variable you were trying to remove.
The security results show the same thing. Veracode's spring 2026 testing across more than 150 models and 80 coding tasks found only 55 percent of generations produced secure code, a number that has stayed roughly flat for two years while syntax correctness climbed past 95 percent (Veracode). Models got much better at writing code that runs and barely better at writing code that is safe. Those two skills were never the same skill, and nothing in the current trend suggests one improves the other.
What we actually do about it
None of this is an argument for using agents less. DORA's 2025 research across nearly 5,000 technology professionals found 90 percent of respondents use AI at work, and that adoption has a positive relationship with delivery throughput alongside a negative relationship with delivery stability (DORA, 2025). More changes reaching production, a higher share of them going wrong. That is not a reason to write less code. It is a description of a system generating changes faster than its checks can handle them, and the fix is on the checks side.
So the rule we hold to is narrow. Delegate the writing, never delegate the definition of correct. Every change we release starts as a written specification, the checks come from that specification before the implementation exists, and grading happens in a fresh context so the thing that wrote the code is not the thing that certifies it. That ordering matters more than any of it, because a check written after the code will slowly change to match whatever the code already does.
Everything in the research above is a version of the same finding. The models have become extraordinary at producing plausible work and have not become able to tell you whether the work is right, and nothing in the current trend suggests those two abilities arrive together. The people running this well are not the ones who found a way to stop checking. They are the ones who made checking cheap enough to do on every change.
We thank METR, Veracode, the DORA team, and the academic groups doing the unglamorous work of auditing benchmarks everyone else quotes without reading. Their caveats are the most valuable part of their research, and the part that people leave out first.
An agent can write the whole thing. Deciding what "done" means is still the job.
Sources
- Fortune, 29 January 2026: 100 percent of code at Anthropic and OpenAI is now AI-written: the Cherny and Roon quotes, the 70 to 90 percent company-wide figure, and the Microsoft comparison. These are self-reported company claims.
- METR, Time Horizon 1.1 (January 2026): Claude Opus 4.5 at 320 minutes, GPT-5 at 214, o3 at 121, and a 130.8-day doubling time since 2023 across a 228-task suite.
- METR, Clarifying limitations of time horizon (22 January 2026): "Time horizon is not the length of time AIs can work independently," the 50 percent framing, the ~2x error bars, and the 98 percent automation threshold.
- METR, Measuring AI Ability to Complete Long Tasks (March 2025): Claude 3.7 Sonnet at a 15-minute 80 percent horizon against a 59-minute 50 percent horizon, and the messiness penalty of about 8.1 percentage points per point.
- SWE-bench Verified: 500 human-validated instances, produced in collaboration with OpenAI.
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study: 7.8 percent of patches counted correct while failing the developer test suite, 29.6 percent behaviourally divergent, 6.2 absolute points of inflation.
- On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub: 567 pull requests across 157 projects, 83.8 percent merged, 54.9 percent merged without further modification.
- Veracode, Spring 2026 GenAI Code Security update: 55 percent of generations secure against syntax correctness above 95 percent.
- Google Cloud / DORA, 2025 State of AI-assisted Software Development: 90 percent AI use, throughput up and delivery stability down, across nearly 5,000 respondents.


