The business case for evals
The case for evals is that AI moved your delivery risk from the writing step to the checking step, and the checking step is the one you are still doing with human attention that does not grow with the amount of code. This page is written for the person who has to approve the time for evals, and it avoids invented numbers on purpose.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- AI adoption raises throughput and lowers delivery stability. Evals are how you keep the first without paying for the second.
- The cost comes early and is small: usually days per critical flow, rather than three months.
- The saving is in rework, incidents, and the senior attention that review consumes, which is your scarcest resource.
- Evals are also an asset in a diligence or handover: they document what the software guarantees.
- Ask any development partner what the automated check is, who wrote it, and what happens when it fails.
Someone has to approve the time, and that person is usually being asked to slow down a team that currently feels fast. So it is worth making the argument in terms of risk and cost rather than engineering taste.
What you are actually buying
You are buying the ability to keep releasing quickly without the failure rate rising with the volume, which is the specific thing that goes wrong when a team adopts AI coding tools.
That failure pattern is documented. Google's DORA program surveyed nearly 5,000 technology professionals in 2025 and found AI adoption has a positive relationship with delivery throughput and a negative relationship with delivery stability [1]. Their view is that AI increases the effect of whatever practices the organisation already has. If the checks are strong, you get faster delivery. If they are weak, you get faster rework.
The second thing you are buying is scarce senior attention. Right now that attention is being spent reading diffs to work out whether they function, which a computer does better. Move that work to the machine and the same people spend their time on design and product judgment instead. That is a real reallocation of your most expensive resource, and it is the part that rarely appears in a business case.
What it costs
Less than people expect, if it is not run as a project. One critical flow is typically a day or two: write the invariants as sentences, turn them into checks, connect them to the pipeline as a blocking check. Three flows is a couple of weeks. After that the suite grows inside normal work, because the rule is that you cover the area you are about to change and you turn every incident into a permanent check.
The expensive version is the one to avoid: a testing initiative with a coverage target, running for three months, producing a large number of tests around the easiest code. That is the version that gets cancelled and then quoted as evidence that testing does not pay. Adding evals to an existing codebase is the staged alternative.
Where the return shows up
Four places, in order of size.
Rework avoided. A defect caught by a check in a pull request costs minutes. The same defect caught by a customer costs a support conversation, a hotfix, a release, a retrospective, and the trust of whoever reported it. The difference in cost between those two is the whole argument, and it is why we do not need an invented percentage to make the case.
Incidents avoided. Most of the failures that matter are ordinary: a permission boundary that stopped working, a migration that dropped a column, an integration whose payload changed structure without notice. All three are cheap to assert and expensive to discover in production.
Review time recovered. Senior engineers reading verified diffs go much faster than senior engineers verifying unverified ones, and they enjoy it more, which matters for retention.
Change confidence. The least visible one, and its value grows over time. Teams with trustworthy checks refactor, upgrade dependencies, and let agents restructure code, because they can prove nothing changed. Teams without them build up code nobody is willing to change, and that is what technical debt mostly is.
What we will not tell you
We will not give you a percentage for how much our own defect rate fell after we adopted this, because we have not run the controlled experiment that would justify the number. A consultancy that sells verification should not be careless with its own claims, and any figure we published would be exactly the kind of confident, unverifiable statement this practice exists to prevent.
What we will say is what changed in kind. The problems we find now are design disagreements found in a pull request, rather than unexpected behaviour reported by a customer. That is a qualitative claim and it is honest, which is the trade we would rather make.
If you want the outside numbers, use the ones in this guide: DORA on throughput and stability [1], Veracode on the unchanged security pass rate of generated code [2], and Stack Overflow on almost-right output being the top developer complaint [3]. Those are measured by people with nothing to gain from your decision.
Questions to ask a development partner
If you are buying software rather than building it, this is where the practice becomes a purchasing criterion. Everyone uses AI now, so asking whether they do tells you nothing. Ask instead:
What automated check must pass before a merge in our repository, and can we see it? Who writes those checks, and are they written before or after the implementation? What happens when a check fails on a Friday afternoon? How do you stop an agent from weakening a check to get a passing result? Which of our flows will have checks in the first two weeks?
Specific answers mean the discipline is real. A pitch about senior review culture, with no automated check behind it, means the verification gap is being handled only by human attention, and human attention does not scale with generated code volume however senior it is.
Our own answer, so you can hold us to the same standard: all of our code is generated, every change has to pass a large eval suite before it reaches your branch, and those checks are written from the specification rather than from the implementation. The companion question set for choosing an AI development partner is in how to evaluate an AI development partner, and if you want to talk about a specific codebase, get in touch.
Translating this into numbers a finance or operations leader will use
The four returns above are real but qualitative, and the person approving budget often needs at least a rough structure for comparing them against the cost of the work, even without a published percentage. Here is how to build that structure honestly, without inventing a figure this guide has already promised not to invent.
Take your own change failure rate and your own defect escape count from the last quarter, both of which your existing tools can usually produce. Multiply the number of defects that reached customers by whatever your team already knows a support ticket and hotfix cycle costs in engineer hours, since most teams already track this informally even if nobody has written it down as a metric. That gives a real cost of the current state, built from your own data rather than an industry average that may not describe your product. The cost of the first three weeks of eval work is comparatively easy to estimate, because it is bounded: a day or two per flow, a small number of flows. Comparing a bounded, known cost against an ongoing, currently unmeasured cost is usually enough to make the case, without needing a published return on investment figure that nobody could defend under scrutiny anyway.
Why the cost argument gets easier once volume is high
There is a detail worth stating plainly for a buyer weighing whether to fund this now or later: the case for eval-driven development gets stronger, not weaker, as the volume of AI-generated code in a codebase rises. A team generating a small fraction of its code with AI tools can plausibly rely on senior review catching most problems, because the volume is still within what careful human attention can cover. That same team, six months later, generating most of its code with an agent, faces the mismatch described in agent debt: review capacity did not grow, generation volume did, and the argument for automated checks that scale with volume rather than with headcount grows stronger with it. The practical implication for someone approving budget is that the cost of building the suite now is close to fixed, while the cost of delaying it rises with every quarter that AI-generated volume keeps growing. Waiting for a clearer business case is, in effect, waiting for the problem to get more expensive before addressing it.
What this looks like in a due diligence conversation specifically
The diligence angle deserves a concrete example, because "an asset in diligence" is easy to state and easy to underrate. A buyer evaluating a company's codebase typically has days, not months, to form a view of what is reliable and what is fragile. A codebase with a real eval suite gives that buyer something to inspect directly: a list of checks, what each one asserts, and a pass or fail history. A codebase with no suite, or with a suite of the kind described in when evals give false confidence, gives the buyer only the seller's word and whatever spot-checking a short engagement allows. The difference shows up in how a diligence report gets written: one version can state specifically which requirements are proven to hold, and the other has to hedge with language about apparent code quality and general impressions, which is precisely the kind of unverifiable claim serious buyers already discount.
The failure mode of approving the budget and then quietly deprioritising it
One more pattern worth naming for whoever holds the budget: teams that get eval work approved sometimes fail to protect the first three weeks from the ordinary pressure of a sprint, and the work slips a little every cycle until it never actually starts. The fix is not more budget. It is treating the first three flows as a fixed, short commitment with its own deadline, the same way a security patch or a compliance deadline would be treated, rather than as background work that competes with every feature request for the same engineer's attention. A commitment that can be deprioritised indefinitely is not a commitment, and the staged plan in adding evals to an existing codebase only works if the first three weeks actually happen on schedule.
Best for
- Approving the first two weeks of eval work on the flows that would cause the most harm
- Teams whose throughput rose after AI adoption while incidents also rose
- Buyers comparing development partners on something more useful than whether they use AI
Avoid if
- Do not fund a quarter-long testing initiative with a coverage target
- Do not expect a published percentage return: judge it on rework and incident history instead
- Do not accept seniority of reviewers as a substitute for an automated check
Check before you decide
- Confirm your own change failure rate and rework trend since AI adoption
- Confirm a partner can show the check that must pass before a merge, rather than only describe their culture
- Confirm who writes the checks and whether they exist before the implementation
Common questions
How do I justify the time to build an eval suite?
Frame it as keeping the throughput gain from AI without the stability loss that comes with it, which the 2025 DORA research found across nearly 5,000 respondents. The cost is usually a day or two per critical flow rather than a quarter, and the return shows up as avoided rework, fewer incidents, and senior review time recovered.
What is the return on investment for eval tests?
The mechanism is the difference in cost between a defect caught in a pull request and the same defect caught by a customer, which involves support, a hotfix, a release, and lost trust. We deliberately publish no percentage for our own results because we have not run a controlled experiment that would justify one.
What should I ask a software partner about their AI code quality?
Ask what automated check must pass before a merge in your repository, who writes those checks and whether they exist before the implementation, what happens when a check fails late on a Friday, and how they stop an agent from weakening a check to get a passing result. Specific answers indicate a real practice; a description of senior review culture does not.
Are evals useful outside engineering?
Yes, in two places. They are an asset in technical due diligence and in any handover, because they document what the software actually guarantees rather than what a document claims. They also reduce the class of risk that reaches customers, which is a commercial concern rather than an engineering one.
How much does building an eval suite typically cost in engineering time?
Usually a day or two per critical flow when it is run as motivated work rather than a project: write the invariants as sentences, turn them into checks, and connect them to the pipeline as a blocking check. Three flows is a couple of weeks, after which the suite grows inside normal work rather than needing a dedicated initiative.
What is the more expensive alternative to building an eval suite this way?
A testing initiative run as a project with a coverage target, spanning three months and producing a large number of tests around the easiest code to test rather than the code that matters. That version is the one that tends to get cancelled partway through and then cited afterwards as evidence that testing does not pay off.
Does an eval suite free up senior engineering time?
Yes, and that reallocation is a real part of the case even though it rarely appears in a written business justification. Senior engineers currently spend time reading diffs to work out whether they function, which an automated check does more reliably, freeing that attention for design and product judgment instead.
Is it safe to trust a vendor's claim about their AI code quality without evidence?
No. A pitch built around a senior review culture with no automated check behind it means the verification gap is being handled entirely by human attention, and human attention does not scale with a rising volume of generated code however experienced the reviewers are. Ask to see the specific check that must pass before a merge rather than accepting a description of process.
References
- [1] Google Cloud / DORA, 2025 State of AI-assisted Software Development report: positive relationship between AI adoption and throughput, negative relationship with delivery stability, from nearly 5,000 respondents.
- [2] Veracode, Spring 2026 GenAI Code Security update: security pass rates “stubbornly stuck at approximately 55%”, unchanged over roughly two years while syntax correctness exceeded 95%.
- [3] Stack Overflow 2025 Developer Survey, AI section: 66% cite “AI solutions that are almost right, but not quite” as their leading frustration.
Related reading
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.
How to prepare for technical due diligence before a raise or sale
Technical due diligence is where a deal can quietly fail. Here is what investors' technical reviewers actually look at, how to prepare before they do, and the warning signs that worry them.