AI is an amplifier, not a fix

Every supplier you talk to this year will tell you that AI makes them faster. It is a true statement and a useless one, because it is now true of everybody. The 2025 DORA report, which surveyed nearly 5,000 technology professionals, found that 90% of respondents use AI at work. Speed is no longer a differentiator between teams. What separates them is which checks the work must pass before it reaches production.
The same report has two findings that people tend to quote separately, and they should be read together.
The first: "Unlike last year, we observe a positive relationship between AI adoption on both software delivery throughput and product performance." The second, in the next sentence: "However, AI adoption does continue to have a negative relationship with software delivery stability."
Read that as one sentence. Teams using AI release more, and more of what they release breaks. Both are happening at once, to the same teams, right now.
1. Amplifier is the right word
DORA's own framing is that AI acts as an amplifier, making the strengths of a strong organization and the weaknesses of a struggling one larger. That conclusion is the most practically useful thing in the report, because it tells you where to look.
If a team had a habit of merging changes nobody had verified, AI does not introduce that habit. It just runs it at a rate the old process was never designed for. A team that used to open six pull requests a week now opens forty, and the same one reviewer is still the only check before the code reaches the customer. Nothing about that person changed. Their workload went up by a factor of six.
The report says exactly why this ends the way it does: "AI accelerates software development, but that acceleration can expose weaknesses downstream. Without robust control systems, like strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability."
Note what that sentence does not say. It does not say the generated code is bad. It says the volume is higher and the control system did not grow with it. Those are different problems, and only one of them is fixable by choosing a better model.
2. The check before a merge is the whole answer
So the question worth asking a team, whether it is yours or a supplier's, is not which tools they use. Everyone uses the same three, and the answer tells you nothing.
The question is: what has to pass before a change reaches the main branch, and who wrote it?
There are only a few honest answers.
A person reads it. This is where most teams are, and it is the answer that sounds most responsible while working worst as volume grows. A reviewer looking at a diff that nothing has verified is checking the logic in their head, at volume, under time pressure. It works at six pull requests a week. At forty it stops working, and nobody notices, because what goes wrong is that the reviewer starts saying yes faster.
A test suite runs. Better, and it depends entirely on where those tests came from. A suite written after the implementation tends to encode what the code already does. It will pass on a feature that never matched the requirement, because no step in the process ever checked the requirement.
A suite written from the specification, before the implementation existed. This is the one that changes the problem. The check is derived from what you said you wanted, so the code has to meet a requirement set before the code existed. A change either passes it or is not released, and that decision takes no human time at merge time.
This is the practice we build on, and we have written up what it cost us to adopt in adding eval tests was the best decision we made. The short version is that it moved the hard thinking to the start of the work, where changes cost less, and made the merge itself routine. Routine merges are the goal. A merge that depends on one reviewer's extra effort will one day be done by a tired reviewer at the end of a long week.
If you want a longer explanation, the eval-driven development guide covers how to build a suite, what to check first, and how to stop it from growing into two thousand checks nobody trusts.
3. You cannot feel the difference, so stop trying
Here is the finding that should make everyone cautious about their own judgment.
METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real issues, in their own repositories, which averaged more than 22,000 stars and a million lines of code. Before starting, the developers expected AI tools to speed them up by 24%. After finishing, they estimated they had been sped up by 20%.
They were 19% slower.
Think about the size of that difference. These were experienced engineers, working in code they knew well, and their sense of their own productivity was wrong by roughly forty percentage points, in their own favour.
Two honest caveats, because people often quote the number wrongly. METR itself now labels that result historical and says it no longer reflects current tools. Their 2026 continuation estimates a speedup of 18% for developers from the original study, but with a confidence interval running from -38% to +9%, and they openly name a problem with which tasks were included: between 30% and 50% of developers said they were choosing not to submit tasks they did not want to do without AI.
So do not use this to argue that AI slows people down. Use it for the thing it actually establishes, which no update changes: nobody can feel their own throughput accurately when their tools change. Not the developers, not their manager, and not the supplier telling you they are three times faster now.
Which means the only defensible claim about speed is a measured one. If a team cannot tell you how they know, they do not know. They feel it, and the evidence on feelings here is not good.
What this looks like in practice
If you are running the team, the order matters. Fix the check before a merge, and only then raise the volume. Raising output through a weak control system is how you get the DORA outcome: more released, more broken, and a growing sense that the tools made everything worse when what they did was show what was already weak.
Start small. Take the few behaviours that would genuinely hurt you if they broke, the payment process, the permission check, the data your customers would notice going missing. Write a check for each one from the requirement, not from the code. Make it block a merge. That is a small suite, and a small suite that blocks a merge on the right things is better than a large one that blocks nothing, which is the argument we make at length in how many tests does AI-generated code need.
If you are buying the work, the question for checking a supplier is short enough to ask on a first call. What has to pass before a change reaches our branch? Who wrote that check, and did they write it before or after the implementation? What happens when it fails?
A team with a real practice answers in specifics and does not need to think about it. A team without one will talk about their seniority, their culture, and how careful they are. Carefulness is a personality trait. A control system is a set of checks, and only the checks keep working in a bad quarter.
Our own answer, so you can check that we keep to it: all of our code is generated, every change has to pass an eval suite written from the specification before it reaches your branch, and the checks come from the requirement rather than the implementation. That is the entire reason we can move as fast as AI writes code and still take responsibility for what we release. If you want to test that claim, ask us.
The new tools raised the top speed for everyone at once. Each team still has to build its own minimum level of quality.
Sources
- Announcing the 2025 DORA Report, Google Cloud. Survey of nearly 5,000 technology professionals: 90% AI adoption, positive relationship with throughput and product performance, continued negative relationship with delivery stability.
- DORA State of AI-assisted Software Development 2025, DORA. The amplifier framing.
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, METR. 16 developers, 246 issues, 19% slowdown against a forecast 24% speedup and a self-reported 20% speedup.
- We are Changing our Developer Productivity Experiment Design, METR. The 2026 continuation, its confidence intervals, and the selection effects the team names.


