The verification gap: why AI made writing code cheap and checking it expensive
Generating code got roughly one hundred times cheaper in three years. Reading it did not get cheaper at all, because a person still reads at the speed a person reads. That mismatch is the verification gap, and it explains why teams adopting AI tools often feel much faster while producing more work that has to be redone. The research on this is now strong enough to settle the question.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- AI adoption raises delivery throughput and lowers delivery stability at the same time, according to DORA's 2025 research across nearly 5,000 professionals.
- Developers cannot reliably sense their own speed with AI tools, so their sense of quality should not decide whether a change passes either.
- Security is the clearest example of the gap: AI models write compiling code more than 95 percent of the time and secure code around 55 percent of the time, and that has not improved in two years.
- The most common complaint about AI output is that it is almost right, which is precisely the failure a fast human read misses and an automated check catches.
For most of software's history, the slow step was typing. Everything about how teams work, from sprint planning to code review, was designed around that assumption. The assumption stopped being true and the practices have not changed to match.
What the research actually says
Start with throughput and stability. Google's DORA program surveyed nearly 5,000 technology professionals for its 2025 report on AI-assisted software development. Ninety percent use AI at work. More than 80 percent believe it makes them more productive. And the report found a positive relationship between AI adoption and delivery throughput alongside a negative relationship with delivery stability [1]. Read that twice. More changes are reaching production, and a higher share of them are going wrong.
DORA's own description is that AI increases the effect of whatever practices the organisation already has. A team with strong checks gets faster. A team without them gets faster at producing rework.
Then take individual judgment. In July 2025, METR published a randomised controlled trial with 16 experienced open-source developers working 246 real tasks in repositories they knew well, averaging around five years of familiarity. The result was that AI-assisted work took 19 percent longer, and the participants estimated afterwards that it had been 20 percent faster [2]. METR labels this result historical, since the tools available in early 2025 are not the tools available now, and that caveat is fair. The lasting finding is that skilled professionals were wrong by 39 points about their own performance, and the error made the tool look better than it was. If your intuition cannot measure your own speed, it should not be the thing that certifies correctness.
Now take the code itself. Veracode's spring 2026 update tested more than 150 language models across 80 coding tasks in four languages. Only 55 percent of generations produced secure code. Java came out at 29 percent. Cross-site scripting was handled correctly in roughly 13 to 15 percent of relevant cases. Meanwhile syntax correctness exceeded 95 percent, and the security number has stayed the same for about two years while the syntax number rose [3]. Their conclusion is the sentence to remember: models have become excellent at writing code that compiles, and have failed at writing code that is safe.
Finally, take what developers themselves report. Stack Overflow's 2025 survey found 84 percent using or planning to use AI tools, 46 percent actively distrusting the accuracy of the output against 33 percent who trust it, and the top frustration at 66 percent being AI solutions that are almost right, but not quite [4].
Almost right is the hard case
Every one of those findings points to the same kind of failure, and it is a kind human review handles badly.
An obviously wrong change is easy. It does not compile, it fails on the first click, someone catches it in a minute. A subtly wrong change is the expensive one: correct structure, plausible variable names, a reasonable-looking guard clause that checks the wrong boundary. It reads as competent because it is competent everywhere except the one place that matters. A reviewer skimming 600 lines on a Friday afternoon will approve it, and a reviewer skimming 6,000 lines will approve it faster.
This is a statement about what humans are good at. People are excellent at asking whether a design will cause problems in a year. They are poor at mentally executing code against 40 unusual cases, and they get worse as the diff gets longer. A machine is the reverse. The gap grows when volume rises and the only checking is done by people, whose attention is limited.
Why the gap keeps getting worse
Three causes make it worse together.
Volume. Generation cost per line keeps falling, so the natural size of a change keeps rising. Review capacity is fixed by headcount and hours.
Uniformity. Human mistakes are idiosyncratic and scattered. Model mistakes cluster, because models reproduce the patterns in their training data. When one flawed pattern appears, it tends to appear in many places at once, which is why the security numbers stay the same as capability improves.
Confidence. Generated code shows no doubt. Hand-written code often signals doubt through a comment, an unfinished part, a variable named tmp. Generated code is uniformly tidy, and tidy code looks careful even when nobody took care.
What fixes the gap
Only one thing scales with volume, and that is automated verification. Checks read as fast as the machine can run them, do not get tired, do not skim, and do not care whether the diff is 60 lines or 6,000.
That is the argument for eval-driven development, and it is a practical argument rather than a moral one. If your team is generating more code, you need proportionally more machine-checkable assertions about what that code must do. More meetings, a stricter review policy, or a rule that all pull requests must be under 400 lines will not solve it. What works is more checks, running on every change, with the power to block a merge.
The human review does not go away. It gets better, because people stop trying to do the checking work that a computer does better. That division is covered in evals vs tests vs code review, and the wider trust question, including the human review discipline itself, is in the companion guide on taking AI-generated code to production. Evals prove a change is correct; the process around them, from sizing the work to releasing it and measuring what arrived, is the AI-native delivery guide.
The honest version for a buyer
If you are paying someone else to build software with AI, this is the thing to ask about. Asking whether they use AI tells you little, since everyone does, and a supplier who avoids it is slower with no gain in safety. Ask what the automated check is, whether it runs on every change, who wrote it, and what happens when it fails on a Friday. A supplier who can answer that specifically has dealt with the verification gap. A supplier who answers by describing their senior review culture has not, however senior the people are.
Best for
- Teams whose code volume has risen faster than their review capacity
- Codebases where a subtle behaviour change is expensive: payments, permissions, data
- Any team letting a coding agent work with limited supervision
Avoid if
- Do not rely on a stricter review policy alone to handle a large increase in generated code
- Do not treat a supplier's seniority as a substitute for an automated check that blocks bad changes
- Do not read a passing build as verification when the suite asserts almost nothing
Check before you decide
- Confirm what share of merges are blocked by an automated check versus a human opinion
- Confirm your revert and rollback rate before and after AI adoption, as well as your throughput
- Confirm the security-relevant paths have explicit checks, since generated code is weakest there
Common questions
What is the verification gap in AI-assisted development?
It is the mismatch between how fast code can now be produced and how fast it can be checked by a person. Generation cost has fallen sharply while human reading speed is unchanged, so the risk in a codebase grows with volume unless the checking is automated.
Does AI-assisted development make software less stable?
The 2025 DORA research found a positive relationship between AI adoption and delivery throughput and a negative relationship with delivery stability, across nearly 5,000 respondents. The report describes AI as something that increases the effect of an organisation's existing practices, which means the instability is a symptom of missing checks rather than an inevitable property of the tools.
Is AI-generated code less secure than hand-written code?
Veracode's spring 2026 testing of more than 150 models found only 55 percent of generations produced secure code, with Java at 29 percent, and that figure has stayed the same for about two years while syntax correctness rose above 95 percent. The practical reading is that models learned to write code that compiles faster than they learned to write code that is safe, so security-relevant paths need explicit automated checks.
Why is human code review not enough for AI-generated code?
Because the characteristic failure is code that is almost right, which is the hardest thing for a person to spot and the easiest thing for a test to catch. Review capacity is also fixed by hours and attention, so it cannot scale with a rising volume of changes the way an automated suite can.
Why did developers get worse at judging their own speed with AI tools?
METR's July 2025 randomised trial gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. The developers using AI tools took 19 percent longer to finish, yet estimated afterwards that AI had made them 20 percent faster, a difference of 39 points between the real and the perceived result on the same work.
Does the verification gap get better as AI models improve?
Not on its own. Veracode's spring 2026 testing found syntax correctness for generated code rose above 95 percent while the share of secure generations stayed near 55 percent for about two years. Models are improving at producing code that compiles and have stayed the same at producing code that is safe, so the gap between writing speed and checking speed keeps widening unless the checking is automated.
How much faster did code generation get compared to code review?
Generating code got roughly one hundred times cheaper over three years, while reading it did not get any cheaper, because a person still reads at the speed a person reads. That mismatch between a falling cost of writing and a fixed cost of checking is what the verification gap describes.
What is the practical fix for the verification gap?
Automated verification, because it is the only part of the process that scales with volume. Checks run as fast as the machine executes them regardless of whether a diff is 60 lines or 6,000, while human review capacity stays fixed by headcount and hours, so the fix is more machine-checkable assertions rather than a stricter review policy.
References
- [1] Google Cloud / DORA, 2025 State of AI-assisted Software Development report: 90% use AI at work, more than 80% report increased productivity, positive relationship with throughput and negative relationship with delivery stability, from nearly 5,000 respondents.
- [2] METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025): 19% longer completion times against a self-estimated 20% speed-up. METR labels the result historical.
- [3] Veracode, Spring 2026 GenAI Code Security update: more than 150 models, 80 tasks, 55% secure generations, Java at 29%, cross-site scripting handled correctly in 13 to 15% of cases, syntax correctness above 95%.
- [4] Stack Overflow 2025 Developer Survey, AI section: 84% using or planning to use AI, 46% distrust accuracy against 33% who trust it, 66% cite output that is “almost right, but not quite”.
Related reading
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.
How to move fast without breaking the product
Speed and quality are usually described as a trade-off. In practice, the teams that work fastest over a long time are the ones that made quality cheap to keep.
More in Start here
What is eval-driven development?
Eval-driven development is a way of building software where you write the automated check first, encode a requirement from the specification in it, and then let an AI coding agent write and rewrite the implementation until the check passes. The check is the deliverable your team owns and reviews. The code is what satisfies it. That reversed order matters more now than it did, because the code is no longer the expensive part.
Evals vs tests vs code review: which one catches what
These three get treated as interchangeable quality activities and they are not. A unit test proves a function behaves. An eval proves a requirement holds. A human review judges whether the change was a good idea. Only the third can tell you the feature was pointless, and only the first two will still be checking next year when everyone who wrote it has left.