Run it continuously

Metrics for AI code quality: what to watch instead of volume

Once code became cheap to produce, every metric based on how much of it you produce stopped telling you anything useful. What still means something is what comes back: how often a change breaks something, how long recovery takes, how many defects reach a customer, and how much of last month's work is being redone. Those four are still useful after the change, and the first two have ten years of research supporting them.

Published August 20, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Volume metrics measure the part that got cheap. They no longer distinguish a good week from a bad one.
  • DORA's change failure rate and failed deployment recovery time are the two most useful numbers to start with.
  • Escaped defects, meaning issues found by customers rather than by your suite, is the number that tells you whether the evals are working.
  • Watch rework: the share of changes that revise something released in the last 30 days.
  • Throughput without stability is the documented failure pattern of AI adoption, so never report one without the other.

A team we talk to often had a dashboard showing pull requests merged per week, and it had roughly tripled. Everyone knew things felt worse. Nothing on the dashboard could show it, because the dashboard measured the part of the job that a machine had taken over.

Here is what we watch instead.

Start with the two DORA numbers about failure

Google's DORA research has studied software delivery for over a decade, and its four key measures are the most validated set available: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time. The program's own view is that speed and stability are not a trade-off, and that the strongest teams do well on all of them at once [1].

Of the four, the two failure-side numbers are the ones that answer the question this guide is about.

Change failure rate. The share of changes to production that cause a degradation needing a fix, a rollback, or a patch. If your evals are doing their job, this falls even as your deployment frequency rises. If it climbs while volume climbs, you are releasing work that has to be redone and calling it progress.

Failed deployment recovery time. How long it takes to get back to working after a bad change. This is a measure of your ability to detect and reverse, which is exactly the ability an eval suite builds.

The reason to focus on these specifically is that the 2025 DORA report found AI adoption has a positive relationship with throughput and a negative relationship with delivery stability across nearly 5,000 respondents [2]. That is the pattern to watch for in your own numbers, and you cannot see it if you only report the throughput half.

Then add the two that are specific to this problem

Escaped defects. Count the issues found by a customer, or by production monitoring, rather than by your own checks. This is the closest thing to a direct measure of eval quality, and it is more honest than any coverage number because it counts what got through rather than what was attempted.

Track the ratio as well as the count: of the defects you found last month, what share did your suite catch first? A team improving its evals sees that share rise. A team with a large but weak suite sees no change and a lot of passing builds.

Rework rate. The share of changes that modify something released in the previous 30 days. Some rework is healthy iteration, so the number is not meant to be zero. A sharp rise usually means work is being merged before anyone knows whether it was right, which is the exact symptom of code being generated faster than it can be verified.

Two that sound useful and are not

Test coverage. It counts lines that executed, whether or not any behaviour was checked. It is possible to hold 90 percent coverage while proving nothing, and teams that aim for the number reliably reach it. Use it as a rough map of untouched areas, never as a target.

Eval pass rate. A suite passing at 100 percent tells you either that your code is correct or that your checks are weak, and the number cannot distinguish those. Watch what fails and how often instead, and be suspicious of a suite that has not failed in a month.

The measure nobody puts on a dashboard

Ask your engineers this question every quarter: when the pipeline passes, do you believe the change is safe?

The answer is the most predictive thing we know and it does not appear in any tool. A team that says yes will use the suite to block bad changes and let the agent work. A team that says no will re-run builds, review defensively, and rely on manual checking without saying so, which means the whole investment is producing formality instead of speed. If the answer is no, fix the trust problem before adding more checks. When evals give false confidence is about how that trust gets lost.

How to report it honestly

Two rules we hold ourselves to.

Never show throughput without stability next to it. A chart of changes released, with no change failure rate beside it, reports typing speed rather than engineering health, and typing is the part that got cheap.

Label a scoped target as a target. If a number on a page is a timeline you plan to hit rather than an average you measured, say so in the same sentence. We follow that rule on our own site because a company selling verification cannot be careless with its own claims, and it is worth following on an internal dashboard for the same reason.

For the business framing of these numbers, see the business case for evals. For measuring engineers rather than systems, our position is in measure engineers by outcomes, not output.

Common questions

What metrics should we use for AI-generated code quality?

Change failure rate and failed deployment recovery time from the DORA set, plus escaped defects and rework rate. Those four measure the problems that come back after release rather than the amount released, which is the only side that still tells you anything useful now that producing code is cheap.

Why is test coverage a bad target?

Because it counts which lines ran, whether or not any behaviour was proven, so a suite can reach a high percentage while asserting almost nothing. It is a reasonable map of which areas have no checks at all, and a poor goal to manage a team towards.

What is an escaped defect?

An issue found by a customer or by production monitoring rather than by your own checks. Tracking the share of defects your suite catches first is the closest thing to a direct measure of whether your evals are getting better.

Does higher throughput mean the team is doing better?

Not on its own, and the 2025 DORA research is the reason to be careful: AI adoption raised throughput while lowering delivery stability across nearly 5,000 respondents. Any report of volume should carry change failure rate beside it, or it describes typing speed rather than engineering health.

What are the four DORA metrics and which two matter most for eval quality?

The four are deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time, developed by Google's DORA research program over more than a decade of studying software delivery. Of the four, change failure rate and failed deployment recovery time answer whether an eval suite is doing its job, since both measure how often changes break something and how fast the team recovers.

What is rework rate and why does it matter for AI-generated code?

Rework rate is the share of changes that modify something released in the previous 30 days. Some rework is normal iteration, but a sharp rise usually means work is being merged before anyone knows whether it was right, which is the specific symptom of code being generated faster than it can be verified.

Is a 100 percent eval pass rate a good sign?

Not by itself. A suite passing every run either means the code is correct or means the checks are too weak to fail, and the number alone cannot tell you which. Watching what fails and how often is more informative than the pass rate, and a suite that has not failed in a month is worth investigating rather than trusting.

How should throughput be reported alongside stability?

Never show a count of changes released without the change failure rate beside it, because a chart of volume alone measures typing speed, which is the part of the job that became cheap, rather than engineering health. The 2025 DORA findings on AI raising throughput while lowering stability are the reason both numbers need to appear together.