Eval-driven development: how to prove AI-written code works / Run it continuously
Metrics for AI code quality: what to watch instead of volume
Once code became cheap to produce, every metric based on how much of it you produce stopped telling you anything useful. What still means something is what comes back: how often a change breaks something, how long recovery takes, how many defects reach a customer, and how much of last month's work is being redone. Those four are still useful after the change, and the first two have ten years of research supporting them.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Volume metrics measure the part that got cheap. They no longer distinguish a good week from a bad one.
- DORA's change failure rate and failed deployment recovery time are the two most useful numbers to start with.
- Escaped defects, meaning issues found by customers rather than by your suite, is the number that tells you whether the evals are working.
- Watch rework: the share of changes that revise something released in the last 30 days.
- Throughput without stability is the documented failure pattern of AI adoption, so never report one without the other.
A team we talk to often had a dashboard showing pull requests merged per week, and it had roughly tripled. Everyone knew things felt worse. Nothing on the dashboard could show it, because the dashboard measured the part of the job that a machine had taken over.
Here is what we watch instead.
Start with the two DORA numbers about failure
Google's DORA research has studied software delivery for over a decade, and its four key measures are the most validated set available: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time. The program's own view is that speed and stability are not a trade-off, and that the strongest teams do well on all of them at once [1].
Of the four, the two failure-side numbers are the ones that answer the question this guide is about.
Change failure rate. The share of changes to production that cause a degradation needing a fix, a rollback, or a patch. If your evals are doing their job, this falls even as your deployment frequency rises. If it climbs while volume climbs, you are releasing work that has to be redone and calling it progress.
Failed deployment recovery time. How long it takes to get back to working after a bad change. This is a measure of your ability to detect and reverse, which is exactly the ability an eval suite builds.
The reason to focus on these specifically is that the 2025 DORA report found AI adoption has a positive relationship with throughput and a negative relationship with delivery stability across nearly 5,000 respondents [2]. That is the pattern to watch for in your own numbers, and you cannot see it if you only report the throughput half.
Then add the two that are specific to this problem
Escaped defects. Count the issues found by a customer, or by production monitoring, rather than by your own checks. This is the closest thing to a direct measure of eval quality, and it is more honest than any coverage number because it counts what got through rather than what was attempted.
Track the ratio as well as the count: of the defects you found last month, what share did your suite catch first? A team improving its evals sees that share rise. A team with a large but weak suite sees no change and a lot of passing builds.
Rework rate. The share of changes that modify something released in the previous 30 days. Some rework is healthy iteration, so the number is not meant to be zero. A sharp rise usually means work is being merged before anyone knows whether it was right, which is the exact symptom of code being generated faster than it can be verified.
Two that sound useful and are not
Test coverage. It counts lines that executed, whether or not any behaviour was checked. It is possible to hold 90 percent coverage while proving nothing, and teams that aim for the number reliably reach it. Use it as a rough map of untouched areas, never as a target.
Eval pass rate. A suite passing at 100 percent tells you either that your code is correct or that your checks are weak, and the number cannot distinguish those. Watch what fails and how often instead, and be suspicious of a suite that has not failed in a month.
The measure nobody puts on a dashboard
Ask your engineers this question every quarter: when the pipeline passes, do you believe the change is safe?
The answer is the most predictive thing we know and it does not appear in any tool. A team that says yes will use the suite to block bad changes and let the agent work. A team that says no will re-run builds, review defensively, and rely on manual checking without saying so, which means the whole investment is producing formality instead of speed. If the answer is no, fix the trust problem before adding more checks. When evals give false confidence is about how that trust gets lost.
How to report it honestly
Two rules we hold ourselves to.
Never show throughput without stability next to it. A chart of changes released, with no change failure rate beside it, reports typing speed rather than engineering health, and typing is the part that got cheap.
Label a scoped target as a target. If a number on a page is a timeline you plan to hit rather than an average you measured, say so in the same sentence. We follow that rule on our own site because a company selling verification cannot be careless with its own claims, and it is worth following on an internal dashboard for the same reason.
For the business framing of these numbers, see the business case for evals. For measuring engineers rather than systems, our position is in measure engineers by outcomes, not output.
Common questions
What metrics should we use for AI-generated code quality?
Change failure rate and failed deployment recovery time from the DORA set, plus escaped defects and rework rate. Those four measure the problems that come back after release rather than the amount released, which is the only side that still tells you anything useful now that producing code is cheap.
Why is test coverage a bad target?
Because it counts which lines ran, whether or not any behaviour was proven, so a suite can reach a high percentage while asserting almost nothing. It is a reasonable map of which areas have no checks at all, and a poor goal to manage a team towards.
What is an escaped defect?
An issue found by a customer or by production monitoring rather than by your own checks. Tracking the share of defects your suite catches first is the closest thing to a direct measure of whether your evals are getting better.
Does higher throughput mean the team is doing better?
Not on its own, and the 2025 DORA research is the reason to be careful: AI adoption raised throughput while lowering delivery stability across nearly 5,000 respondents. Any report of volume should carry change failure rate beside it, or it describes typing speed rather than engineering health.
What are the four DORA metrics and which two matter most for eval quality?
The four are deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time, developed by Google's DORA research program over more than a decade of studying software delivery. Of the four, change failure rate and failed deployment recovery time answer whether an eval suite is doing its job, since both measure how often changes break something and how fast the team recovers.
What is rework rate and why does it matter for AI-generated code?
Rework rate is the share of changes that modify something released in the previous 30 days. Some rework is normal iteration, but a sharp rise usually means work is being merged before anyone knows whether it was right, which is the specific symptom of code being generated faster than it can be verified.
Is a 100 percent eval pass rate a good sign?
Not by itself. A suite passing every run either means the code is correct or means the checks are too weak to fail, and the number alone cannot tell you which. Watching what fails and how often is more informative than the pass rate, and a suite that has not failed in a month is worth investigating rather than trusting.
How should throughput be reported alongside stability?
Never show a count of changes released without the change failure rate beside it, because a chart of volume alone measures typing speed, which is the part of the job that became cheap, rather than engineering health. The 2025 DORA findings on AI raising throughput while lowering stability are the reason both numbers need to appear together.
References
- [1] Google DORA, Four Keys of software delivery performance: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time, with the finding that “speed and stability are not tradeoffs”.
- [2] Google Cloud / DORA, 2025 State of AI-assisted Software Development report: a positive relationship between AI adoption and delivery throughput, and a negative relationship with delivery stability, across nearly 5,000 respondents.
Related reading
You cannot measure engineers by how much they produce
Lines of code, tickets closed, and hours logged all measure activity rather than progress. Here is how we think about engineering output without numbers that only look good.
How to move fast without breaking the product
Speed and quality are usually described as a trade-off. In practice, the teams that work fastest over a long time are the ones that made quality cheap to keep.
More in Run it continuously
Running evals in CI when agents write the code
An eval that does not run on every change is only a note somebody wrote once. The value comes from blocking: no change reaches the main branch unless the relevant checks passed. That is simple to say and has a few real design decisions inside it, mostly about speed, about what an agent is allowed to do with a failed build, and about what happens on the day the check is wrong.
When evals give false confidence
A suite that catches nothing is worse than having no suite, because a team with no checks knows it is exposed and a team with passing checks believes it is covered. Six failure modes account for almost all of it, and each one has a specific sign you can look for this afternoon.
Agent debt: when generation outpaces review
Teams now generate most of their code with an agent and review almost none of it as carefully as before, because the agent is fast enough to make thorough review feel like the slowest step. The cost of that missing review appears later, in production, months after the pull request was merged, as an incident nobody can trace to a specific change.