Eval-driven development: how to prove AI-written code works / Run it continuously
When evals give false confidence
A suite that catches nothing is worse than having no suite, because a team with no checks knows it is exposed and a team with passing checks believes it is covered. Six failure modes account for almost all of it, and each one has a specific sign you can look for this afternoon.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- No suite makes a team careful. A weak suite makes a team confident. The second is more dangerous.
- The most common failure is an assertion that cannot fail: replaced by a mock, skipped, or comparing a value to itself.
- One flaky check devalues every other check, because it teaches the team to re-run instead of read.
- A suite written by the same pass that wrote the code tests the implementation's own interpretation rather than the requirement.
- If nothing in your suite has failed in a month, treat that as a finding rather than a reassurance.
Every team that has run a real incident review has had this moment: the pipeline passed, the change was released, the customer found the bug. The suite was answering a different question from the one everyone thought it was answering.
Here are the six ways that happens.
1. The assertion that cannot fail
The most common by a wide margin. The check runs, the check passes, and there is no situation in which it would fail.
The mock returns exactly the value the assertion expects, so the real code path is never exercised. The setup fails without any error and the test asserts on an empty result. The assertion compares a variable to itself after a round trip. The try block catches and hides the error and the test only checks that nothing was thrown.
The sign: make it fail on purpose. Break the implementation the check is supposedly protecting and confirm the build fails. Do this for every check when it is written and spot-check a handful whenever you take over a suite. In an existing codebase this exercise usually finds something within the first ten minutes.
2. Flakiness
A check that fails one run in twenty does more damage than it prevents. The main damage comes from what the team learns: a failure might not mean anything, so re-run it. Once that habit exists, it applies to every check, including the one that was correct.
The sign: ask how often people re-run a build without changing anything. If the answer is more than rarely, you have this problem, and the fix is to remove the unreliable checks from the blocking tier today rather than to schedule a cleanup.
3. The suite that describes the implementation
Checks that assert internal structure rather than external promises break on every refactor. The team responds rationally by updating the tests to match the new structure, which means the suite is being rewritten to agree with whatever the code now does. At that point it only copies what the code does and has stopped being a check.
The sign: what happens to your suite during a refactor that changes no behaviour? If a large share of checks need editing, they were describing the code rather than the requirement.
4. Written by the same pass that wrote the code
If one run produces both implementation and tests, the tests encode the implementation's interpretation of the requirement, including its misunderstandings. They will pass. They prove internal consistency, which is not the property you wanted.
Anthropic's own guidance points to the fix: a fresh context or a separate reviewer so the thing doing the work is not the thing grading it, since a fresh context is not biased towards code it just wrote [1]. Our stronger version is to derive the check from the specification before the implementation exists.
The sign: read the check and ask whether it would still be the right check if the feature had been built a completely different way. If not, it came from the code.
5. Everything tests only normal use
The suite covers the flows people demo and none of the ones that break. No empty collections, no expired sessions, no duplicate submissions, no unauthorised caller, no dependency timing out.
This matters more for generated code because its failures tend to be of this kind. The most cited developer frustration in Stack Overflow's 2025 survey, at 66 percent, was output that is almost right but not quite [2], and almost right nearly always means correct in normal use and wrong in an unusual case.
The sign: count the negative cases in your suite. If fewer than a third of your checks assert that something is refused, rejected, or handled as an error, the unusual cases are not covered.
6. Nothing has failed in a month
A suite that never fails is either protecting perfect code or protecting nothing, and the passing results cannot tell you which. In a team releasing real changes, a healthy suite catches something regularly, and each catch is the system doing its job.
The sign: look at the last thirty days of pipeline history. If there are no genuine failures, go and break something to check that the connection still works. It is a strange thing to have to verify and it is worth verifying.
Why this is the real risk
Adding evals badly is a specific kind of expensive. You pay the cost of building and maintaining the suite, you lose the caution that a team without checks naturally has, and you gain a signal that says everything is fine. Teams in that state release faster with less scrutiny, which is the worst possible combination if the checks test nothing.
So audit for trust rather than volume. A suite of fifteen checks that have each caught something real is worth more than 2,000 that have never failed. And when your engineers say they do not fully believe the passing build, take that seriously, because they are usually right and they are describing the failure mode nobody else can see. The metric side of that is in metrics for AI code quality, and the pipeline practices that prevent most of this is in evals in CI for coding agents.
Best for
- Auditing an inherited suite before you trust it
- Teams whose builds pass while customers keep finding defects
- Any suite that has grown quickly, especially one largely generated
Avoid if
- Do not treat a high check count as evidence of protection
- Do not respond to flakiness by re-running: remove it from the blocking tier
- Do not accept checks written in the same pass as the code as independent verification
Check before you decide
- Confirm each important check fails when you break the thing it protects
- Confirm at least a third of checks assert a refusal, rejection, or error path
- Confirm the suite has caught something real in the last month
Common questions
Can a passing test suite be worse than no test suite?
Yes. A team with no checks knows it is exposed and behaves carefully, while a team with passing checks believes it is covered and releases with less scrutiny. If the checks test nothing, that combination of speed and confidence is the most dangerous state a codebase can be in.
How do I tell whether a test is actually checking anything?
Break the code it protects on purpose and confirm the build fails. Assertions that cannot fail are common, usually because a mock returns the expected value, the setup fails without any error, or the assertion compares a value to itself after a round trip.
Why is one flaky test such a problem?
Because it teaches the team that a failed build might mean nothing, and that lesson applies to every other check in the suite. The real cost is the habit of re-running a failed build instead of reading what it says, after which a genuine failure gets the same treatment.
Is it a good sign if our tests never fail?
No. It is a reason to check. A suite that has caught nothing in a month is either protecting flawless code or protecting nothing, and passing results cannot distinguish those. Break something deliberately and confirm the pipeline notices.
What is a suite that describes the implementation rather than the requirement?
It is a suite made of checks that assert internal structure, so it breaks on every refactor even when behaviour has not changed. Teams respond by rewriting the tests to match the new structure, which means the suite is being edited to agree with whatever the code now does rather than checking a stable requirement, and at that point it only copies the code rather than checking it.
How can I tell if my eval suite is weighted toward normal use?
Count the negative cases: checks that assert something is refused, rejected, or handled as an error. If fewer than a third of the suite covers those, the unusual cases are not covered, and that missing coverage matters more for generated code because the most cited complaint in Stack Overflow's 2025 survey, at 66 percent, was output that is almost right but not quite, which almost always means correct in normal use and wrong in an unusual case.
Why is a suite written by the same pass that wrote the code untrustworthy?
Because the tests encode that same pass's interpretation of the requirement, including any misunderstanding it contains, so they pass by construction and prove internal consistency rather than correctness. Checking whether a test would still be the right check if the feature had been built a completely different way is a fast way to tell whether it came from the code instead of the specification.
What is the practical cost of a suite that gives false confidence?
It is worse than having no suite at all, because a team with no checks stays cautious while a team with passing checks believes it is protected and releases with less scrutiny. That combination of speed and unearned confidence is the most dangerous state a codebase can be in, since nobody is watching for the failures the weak checks were supposed to catch.
References
- [1] Anthropic, Best practices for Claude Code: a fresh context improves review because the model “won't be biased toward code it just wrote”, and a verification subagent means the agent doing the work is not the one grading it.
- [2] Stack Overflow 2025 Developer Survey, AI section: 66% of respondents cite “AI solutions that are almost right, but not quite” as their top frustration.
How a build like this runs
Related reading
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
The early warning signs your software project is in trouble
A software project rarely fails in one dramatic moment. It falls behind slowly, and the early signs are easy to explain away. Here are the ones to watch and what to do about each before it is too late.
What technical debt really is, and when to pay it back
Technical debt is not messy code. It is a deliberate choice to give up some quality for speed. Here is how to tell smart debt from reckless debt, and when to pay each one back.
More in Run it continuously
Running evals in CI when agents write the code
An eval that does not run on every change is only a note somebody wrote once. The value comes from blocking: no change reaches the main branch unless the relevant checks passed. That is simple to say and has a few real design decisions inside it, mostly about speed, about what an agent is allowed to do with a failed build, and about what happens on the day the check is wrong.
Metrics for AI code quality: what to watch instead of volume
Once code became cheap to produce, every metric based on how much of it you produce stopped telling you anything useful. What still means something is what comes back: how often a change breaks something, how long recovery takes, how many defects reach a customer, and how much of last month's work is being redone. Those four are still useful after the change, and the first two have ten years of research supporting them.
Agent debt: when generation outpaces review
Teams now generate most of their code with an agent and review almost none of it as carefully as before, because the agent is fast enough to make thorough review feel like the slowest step. The cost of that missing review appears later, in production, months after the pull request was merged, as an incident nobody can trace to a specific change.