How many tests does AI-generated code need?

"What coverage percentage should we be targeting?" It is the most common question we get about testing AI-written code, and it has the same problem as asking how many pages a good book should have.
Coverage counts which lines of code executed during a test run. It does not count whether anything was proven about them. You can execute a line and assert nothing, and a large share of the tests in most real codebases do exactly that.
So here is the unit we use instead, and how to know when you have enough.
Count protected promises, not covered lines
Your product makes promises. A user's data is only visible to their own organisation. A payment is either taken and recorded or neither. An export contains every row the filter selected and no others. A deleted account's tokens stop working immediately.
For each promise, ask one question: is there a check that would fail if this stopped being true?
That is the unit. It is harder to measure than a percentage and it is the only one that tells you anything. A codebase with fifteen checks covering fifteen real promises is safer than one with two thousand tests covering none of them, and we have seen both.
The exercise takes an afternoon. List your promises, honestly, including the ones nobody wrote down. Mark the ones with a real check behind them. The gaps are your work queue, in order of what a failure would cost. We wrote the full version of the method in how to write your first eval suite.
Why coverage targets cause harm
Not just imprecise. Actively counterproductive, for a reason that is about incentives rather than measurement.
The cheapest way to raise coverage is to test the code that is easiest to test. That means mappers, formatters, plain data objects, small pure functions with no dependencies. Those are also the parts of a codebase least likely to cause an incident. Meanwhile the expensive risks are in the code that is hardest to test: permission logic mixed into a service, a migration with side effects, an integration that needs a real database to exercise properly.
A team managing to a coverage number will produce a beautiful report and leave the dangerous parts untouched, because the dangerous parts cost ten times as much coverage per hour of work. This is not a hypothetical failure mode. It is what a coverage target does every time, unless somebody deliberately works against it.
Use coverage as a guide to where tests exist, if you like. Zero percent on a module tells you nobody has looked. As a target it sends effort to exactly the wrong code.
Where generated code needs more than usual
Some classes of check matter more when a machine is writing the implementation, and the evidence for this is specific rather than a feeling.
Veracode's spring 2026 update tested more than 150 language models across 80 coding tasks. Only 55 percent of generations produced secure code. Java came out at 29 percent. Cross-site scripting was handled correctly in roughly one case in seven. Syntax correctness, meanwhile, exceeded 95 percent, and the security number has stayed the same for about two years while the syntax number rose. Their summary is the sentence worth remembering: models have become excellent at writing code that compiles and have failed at writing code that is safe.
So security is not something to assume was handled. Give the common flaw classes explicit checks on the paths that matter: injection anywhere a string becomes a query, escaping wherever user text is rendered, authentication asserted by testing that the unauthenticated call is refused.
The second area is error paths. Generated code is strongest in normal use, because that is what most examples in any training corpus demonstrate. In Stack Overflow's 2025 survey, the leading frustration with AI tools, cited by 66 percent of respondents, was output that is almost right but not quite. Almost right nearly always means correct in the common case and wrong in an unusual one: an empty list, a duplicate submission, an expired token, a dependency that timed out.
A useful ratio, if you want one number: at least a third of your checks should assert that something is refused, rejected, or handled as an error. Most suites we audit are well under that, and it shows up in their incident history.
The full list of what is worth checking, in priority order, is in what to check in an eval suite.
The two tests of whether you have enough
Forget the count. Two questions answer this better.
Would the suite have caught your last three incidents? Go and look. Pull the three most recent things that reached a customer, and ask what check would have stopped each one. If none of them were preventable by an automated check, your risk is elsewhere and you should stop adding tests. If all three were, you know exactly what to write next.
Do your engineers believe a passing build means the change is safe? Ask them, plainly, and listen for uncertain answers. A team that says yes will use the pipeline as a real approval step and let an agent work unattended. A team that says no will re-run builds, review defensively, and quietly rely on manual checking, which means the whole suite is producing a routine with no real value instead of confidence.
If the answer is no, more checks will not help. The problem is trust, and trust is broken by specific things: assertions that cannot fail, flakiness, and tests that mirror the implementation instead of the requirement. Those failure modes and their warning signs are in when evals give false confidence.
Deleting tests is part of the job
A suite gets better by removing tests as often as by adding them. Three categories are worth removing as soon as you find them.
Anything flaky. A check that fails one run in twenty costs more than it saves, not because of the wasted runs but because of what the team learns: a failed check might mean nothing, so re-run it. That lesson then applies to every other check in the suite, including the honest ones.
Anything asserting internal structure. Private call counts, class existence, the internal order of operations. These break on every refactor, and a team that spends its time updating tests to agree with new code has a copy of the code rather than a check.
Anything cosmetic. Exact copy, exact spacing, exact colour. Assert that the element is present and reachable. Do not make ordinary design work pay that cost.
What good looks like
We are wary of numbers in this piece for a reason, so here is what it contains rather than how big it is. A healthy suite on a product we would be comfortable running has: every promise that would hurt if broken covered by at least one check that has actually failed in real use, at least a third of its checks on refusals and error paths, explicit assertions on permissions and on the security classes generated code handles badly, migrations checked against data that looks like production data, nothing flaky in the blocking tier, and a real failure caught somewhere in the last month.
Count the checks if you want. It will be a smaller number than you expect, and it will do more work than the two thousand tests it replaced.
Thanks to the teams who let us run the last-three-incidents exercise on their own suites. Nobody enjoys that hour and everybody leaves it with a better list of work.
Coverage tells you what your tests ran. Only your incident history tells you what they were worth.
Related guide: Eval-driven development: how to prove AI-written code works.
Sources
- Veracode, Spring 2026 GenAI Code Security update: more than 150 models across 80 tasks, 55 percent of generations secure, Java at 29 percent, syntax correctness above 95 percent and flat security pass rates over roughly two years.
- Stack Overflow 2025 Developer Survey, AI section: 66 percent cite AI output that is almost right but not quite as their leading frustration.


