Strategy

The real cost of shipping unverified code

Editorial · Reveneau · August 27, 2026

The real cost of shipping unverified code

The cost of unverified code is real and it almost never appears as a line item. That is why teams pay too little attention to it: the cost arrives in four separate places, and only one of them looks like a software defect.

This piece is written for whoever has to approve the time. We have deliberately kept our own numbers out of it, for reasons we explain at the end.

Cost one: the same defect, two prices

Take one bug. A permission check that stops working for users in a second organisation.

Caught by an automated check in a pull request, it costs about ten minutes. An engineer reads a failed build, fixes the boundary, and moves on. Nobody else in the company ever learns it happened.

Caught by a customer, the same bug costs a support conversation, an escalation, an engineer taken off whatever they were doing, a diagnosis, a fix, an out-of-band release, a retrospective, a note to whoever was affected, and some amount of trust that does not come back on a schedule.

Same defect. The multiplier between those two numbers is the entire argument for verification, and it is why we do not need to invent a statistic to make the case. Anyone who has run both weeks knows the ratio.

Cost two: senior attention spent on the wrong job

Your most expensive engineers are, right now, reading diffs to work out whether they function. Mentally executing code against unusual cases. Checking whether a boundary condition was handled.

That is the one part of their job a computer does better than they do. It is exact, it does not tire, and it does not get worse as the diff gets longer, which human attention certainly does.

The research on this is worth thinking about carefully. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in codebases they knew well. With AI tools they took 19 percent longer, and afterwards estimated AI had made them 20 percent faster. METR labels the finding historical, since the tools have changed, and that is fair. The part that is still true is that skilled professionals could not tell which direction their own performance had moved. If intuition cannot measure speed, it should not be the certification step for correctness.

Move the mechanical checking to a machine and the same people spend their week on design, on product judgment, and on whether the thing was worth building. That is a reallocation of your scarcest resource, and it rarely appears in any business case because nobody bills for it.

Cost three: the codebase nobody will change

This is the largest cost and the one almost nobody puts in a document, because it does not look like a cost. It looks like caution.

When a team cannot prove that a change is safe, it stops making changes. The dependency upgrade gets deferred. The refactor that would make the next three features easy never happens. A module becomes known as risky, and then the team gives it a nickname. New work avoids it, adding a layer instead of fixing the thing, and each layer makes the next change harder.

Most of what people call technical debt is exactly this: code nobody dares change. Not badly written code, unprovable code. And the interesting thing is that a check can convert it back. Capture what a module the team is afraid of currently does with fixtures, assert the outputs, and it becomes editable again, sometimes without changing a line of it. We wrote up that technique in adding evals to an existing codebase, and our wider view on debt is in what is technical debt and when to pay it.

This cost grows over time, which is why it eventually becomes larger than the other three. Every month the code stays unchanged makes the next change more expensive.

Cost four: a pipeline nobody believes

The last one is cultural and it removes the return on everything else.

Once engineers stop trusting a passing build, they compensate. They re-run builds. They review defensively. They add manual checking before release. They avoid letting an agent work unattended, which removes most of the speed advantage that justified the tooling in the first place.

At that point you are paying for the suite and getting no value from it. And the trust is broken by specific, identifiable things: checks that cannot fail, flaky checks, and checks that mirror the implementation. All three are fixable, and all three need to be found deliberately, because a board of passing checks does not report them. The warning signs are in when evals give false confidence.

The pattern to watch for in your own numbers

There is a documented pattern to this, and you can check your own data against it this week.

Google's DORA program surveyed nearly 5,000 technology professionals for its 2025 report on AI-assisted software development. Ninety percent use AI at work. More than 80 percent believe it makes them more productive. And AI adoption showed a positive relationship with delivery throughput alongside a negative relationship with delivery stability. More changes reaching production, and a higher share of them going wrong.

DORA's framing is that AI makes whatever the organisation already is stronger. Strong verification, and you get faster delivery. Weak verification, and you get faster rework, which feels like speed for about two quarters.

So look at your throughput next to your change failure rate and your rework rate. If the first is up and the other two are up with it, you are not faster. You are delaying the cost, and it arrives later as incident load and as parts of the codebase quietly becoming unmaintainable. The measures worth tracking are in metrics for AI code quality.

What the fix costs

Less than the framing usually implies, provided it is not run as a project.

One critical flow is a day or two: write down what must always be true as plain sentences, turn each into an automated check, add them to the pipeline as a required check that blocks a release. Three flows is two or three weeks. After that the suite grows inside normal work, because the rules are to cover the area you are about to change and to turn every defect that reached users into a permanent check.

The version that fails is the testing initiative with a coverage target running for a quarter. It produces a large number of low-value tests around the easiest code, gets cancelled when a deadline arrives, and then gets quoted for years as evidence that testing does not pay. Avoid it. Start with three flows and stop.

What we will not tell you

We will not give you a percentage for what this saved us. We have not run the controlled experiment that would justify a number, and a company whose pitch is accountability should not publish figures it cannot show the working for. That constraint costs us a more striking blog post and it is the right trade.

What we will say is that the character of our problems changed. What we find now are design disagreements, argued about in a pull request between two engineers. Not behaviour surprises reported by a customer.

If you want numbers for a business case, use the ones above. DORA on throughput and stability, Veracode on generated code being secure in only about 55 percent of generations while compiling more than 95 percent of the time, and METR on the unreliability of professional intuition about AI-assisted work. Those were measured by people with no interest in what you decide, which makes them worth more than anything we could claim about ourselves.

Thanks to the finance and engineering leaders who pressed us for a number and accepted an honest answer instead. That conversation is the reason this piece exists.

Unverified code only delays its cost, and whoever has to change it next pays that cost with extra added.

Related guide: Eval-driven development: how to prove AI-written code works.

Sources

Common questions

What does unverified code actually cost a business?

Four things: rework on defects that reach customers, senior engineering attention spent verifying diffs by hand, a codebase nobody will change because nobody can prove a change is safe, and the loss of trust in the pipeline that makes every automation gain exist only on paper. Only the first shows up in a bug tracker.

Why is a defect caught late so much more expensive?

Because the cost is not the fix. A defect found by a customer brings a support conversation, an out-of-band release, a retrospective, the interruption of whatever the team was doing, and some lost trust, while the same defect caught by a check in a pull request costs a few minutes of an engineer's time.

What is the connection between unverified code and technical debt?

Most of what teams call technical debt is code nobody dares change, and code becomes untouchable when there is no way to prove a change did not break it. Adding a characterisation check to a module the team is afraid of often makes it editable again without any refactor at all.

Does AI-assisted development increase or decrease delivery risk?

It does both, and the 2025 DORA research across nearly 5,000 respondents found AI adoption raising throughput while lowering delivery stability. The report describes AI as something that makes existing practice stronger, good or bad, which means the risk depends on whether automated verification scaled with the volume of code.

How much does it cost to start verifying properly?

Usually a day or two per critical flow, and two or three weeks for the three flows where a silent failure would cost the most. The expensive version is a quarter-long testing initiative with a coverage target, which tends to stop making progress and then gets cited as evidence that testing does not pay.

How do I make the case to someone who thinks the team is already fast?

Show the throughput number next to the change failure rate and the rework rate. A team releasing more while breaking more is not faster; it is delaying costs, and those costs are paid later in incident load and in the areas of the codebase that quietly become unmaintainable.

What should we measure to see this cost?

Track change failure rate, failed deployment recovery time, escaped defects, and the share of changes that revise something released in the last 30 days. Volume metrics stopped carrying information once producing code became cheap, so a team that only counts commits or pull requests merged will see a false picture of how well the pipeline is actually working.

Is unverified code ever an acceptable trade?

Yes, for genuinely disposable work: a prototype that will be deleted, an internal script, an experiment whose only purpose is to answer a question. The trade goes wrong when a prototype quietly becomes the product without anyone deciding that it had, because by then the codebase carries production risk with none of the checks that production work is supposed to have.

What is the first thing to fix on a codebase with no verification?

Fix the flows where failure is expensive and quiet: permissions, money movement, data deletion, and the numbers customers make decisions from. On an AI-built codebase, start with security and permission assertions, since generated code tends to work well in normal use and is weakest exactly there.

How does this affect a company being acquired or audited?

A trustworthy check suite is one of the few documents that describes what software actually guarantees rather than what a specification claims, so it makes diligence faster and cheaper. Its absence usually turns into a list of unquantified risks in a report you do not control.