Engineering

How many tests does AI-generated code need?

Editorial · Reveneau · August 22, 2026

How many tests does AI-generated code need?

"What coverage percentage should we be targeting?" It is the most common question we get about testing AI-written code, and it has the same problem as asking how many pages a good book should have.

Coverage counts which lines of code executed during a test run. It does not count whether anything was proven about them. You can execute a line and assert nothing, and a large share of the tests in most real codebases do exactly that.

So here is the unit we use instead, and how to know when you have enough.

Count protected promises, not covered lines

Your product makes promises. A user's data is only visible to their own organisation. A payment is either taken and recorded or neither. An export contains every row the filter selected and no others. A deleted account's tokens stop working immediately.

For each promise, ask one question: is there a check that would fail if this stopped being true?

That is the unit. It is harder to measure than a percentage and it is the only one that tells you anything. A codebase with fifteen checks covering fifteen real promises is safer than one with two thousand tests covering none of them, and we have seen both.

The exercise takes an afternoon. List your promises, honestly, including the ones nobody wrote down. Mark the ones with a real check behind them. The gaps are your work queue, in order of what a failure would cost. We wrote the full version of the method in how to write your first eval suite.

Why coverage targets cause harm

Not just imprecise. Actively counterproductive, for a reason that is about incentives rather than measurement.

The cheapest way to raise coverage is to test the code that is easiest to test. That means mappers, formatters, plain data objects, small pure functions with no dependencies. Those are also the parts of a codebase least likely to cause an incident. Meanwhile the expensive risks are in the code that is hardest to test: permission logic mixed into a service, a migration with side effects, an integration that needs a real database to exercise properly.

A team managing to a coverage number will produce a beautiful report and leave the dangerous parts untouched, because the dangerous parts cost ten times as much coverage per hour of work. This is not a hypothetical failure mode. It is what a coverage target does every time, unless somebody deliberately works against it.

Use coverage as a guide to where tests exist, if you like. Zero percent on a module tells you nobody has looked. As a target it sends effort to exactly the wrong code.

Where generated code needs more than usual

Some classes of check matter more when a machine is writing the implementation, and the evidence for this is specific rather than a feeling.

Veracode's spring 2026 update tested more than 150 language models across 80 coding tasks. Only 55 percent of generations produced secure code. Java came out at 29 percent. Cross-site scripting was handled correctly in roughly one case in seven. Syntax correctness, meanwhile, exceeded 95 percent, and the security number has stayed the same for about two years while the syntax number rose. Their summary is the sentence worth remembering: models have become excellent at writing code that compiles and have failed at writing code that is safe.

So security is not something to assume was handled. Give the common flaw classes explicit checks on the paths that matter: injection anywhere a string becomes a query, escaping wherever user text is rendered, authentication asserted by testing that the unauthenticated call is refused.

The second area is error paths. Generated code is strongest in normal use, because that is what most examples in any training corpus demonstrate. In Stack Overflow's 2025 survey, the leading frustration with AI tools, cited by 66 percent of respondents, was output that is almost right but not quite. Almost right nearly always means correct in the common case and wrong in an unusual one: an empty list, a duplicate submission, an expired token, a dependency that timed out.

A useful ratio, if you want one number: at least a third of your checks should assert that something is refused, rejected, or handled as an error. Most suites we audit are well under that, and it shows up in their incident history.

The full list of what is worth checking, in priority order, is in what to check in an eval suite.

The two tests of whether you have enough

Forget the count. Two questions answer this better.

Would the suite have caught your last three incidents? Go and look. Pull the three most recent things that reached a customer, and ask what check would have stopped each one. If none of them were preventable by an automated check, your risk is elsewhere and you should stop adding tests. If all three were, you know exactly what to write next.

Do your engineers believe a passing build means the change is safe? Ask them, plainly, and listen for uncertain answers. A team that says yes will use the pipeline as a real approval step and let an agent work unattended. A team that says no will re-run builds, review defensively, and quietly rely on manual checking, which means the whole suite is producing a routine with no real value instead of confidence.

If the answer is no, more checks will not help. The problem is trust, and trust is broken by specific things: assertions that cannot fail, flakiness, and tests that mirror the implementation instead of the requirement. Those failure modes and their warning signs are in when evals give false confidence.

Deleting tests is part of the job

A suite gets better by removing tests as often as by adding them. Three categories are worth removing as soon as you find them.

Anything flaky. A check that fails one run in twenty costs more than it saves, not because of the wasted runs but because of what the team learns: a failed check might mean nothing, so re-run it. That lesson then applies to every other check in the suite, including the honest ones.

Anything asserting internal structure. Private call counts, class existence, the internal order of operations. These break on every refactor, and a team that spends its time updating tests to agree with new code has a copy of the code rather than a check.

Anything cosmetic. Exact copy, exact spacing, exact colour. Assert that the element is present and reachable. Do not make ordinary design work pay that cost.

What good looks like

We are wary of numbers in this piece for a reason, so here is what it contains rather than how big it is. A healthy suite on a product we would be comfortable running has: every promise that would hurt if broken covered by at least one check that has actually failed in real use, at least a third of its checks on refusals and error paths, explicit assertions on permissions and on the security classes generated code handles badly, migrations checked against data that looks like production data, nothing flaky in the blocking tier, and a real failure caught somewhere in the last month.

Count the checks if you want. It will be a smaller number than you expect, and it will do more work than the two thousand tests it replaced.

Thanks to the teams who let us run the last-three-incidents exercise on their own suites. Nobody enjoys that hour and everybody leaves it with a better list of work.

Coverage tells you what your tests ran. Only your incident history tells you what they were worth.

Related guide: Eval-driven development: how to prove AI-written code works.

Sources

Common questions

How much test coverage does AI-generated code need?

Coverage is the wrong measure, because it counts which lines executed rather than which behaviours were proven, so a suite can reach a high percentage while asserting almost nothing. Count protected promises instead: for each thing your product guarantees, is there a check that would fail if the guarantee broke?

Does AI-generated code need more tests than hand-written code?

It needs more automated verification, though not necessarily more test files. The volume arriving is higher than a person can review, and the characteristic failure is code that is almost right, which is the hardest kind for a human to spot and one of the easiest for an assertion to catch.

What is a reasonable number of checks to start with?

Five to twenty per critical flow is a normal starting point, and small is deliberate. Each should cover something that would cause real damage, run fast enough that nobody avoids it, and have failed at least once in real conditions so the team knows it works.

Why is a high coverage target counterproductive?

Because the fastest way to raise coverage is to test the code that is easiest to test, which is rarely the code that causes incidents. Working to reach the number moves effort towards mappers, formatters, and small pure functions while permission logic mixed into a service and migrations with side effects stay unasserted, since those are ten times more expensive to cover per hour of work.

How do I know when a codebase has enough tests?

Ask whether the suite would have caught your last three incidents, and whether your engineers believe a passing build means the change is safe. Those two answers are more informative than any percentage, and if the second one is no, adding more checks will not fix it.

Should the AI write the tests as well as the code?

It can draft them, and the checks should come from the specification rather than the implementation. When one pass produces both the code and its tests, the tests encode that implementation's own interpretation of the requirement, so they pass while proving nothing about whether the requirement itself was met. Writing the promise as a sentence first keeps the check accurate about what it is testing.

What kinds of check matter most for generated code?

Permissions and tenancy, data invariants including migrations, contracts at every boundary, error paths, and the security classes that generated code handles worst. Veracode's spring 2026 testing found only about 55 percent of AI generations were secure while more than 95 percent compiled, so security needs its own assertions rather than an assumption.

Is it worth deleting tests?

Yes, for anything flaky, anything asserting internal implementation detail, and anything cosmetic that changes weekly. A check that fails intermittently teaches the team to re-run instead of read, which quietly devalues every other check in the suite. Tests that mirror private call counts or exact spacing break on every refactor and end up copying the code instead of checking it.

How should a suite grow over time?

From incidents and from the work in front of you. Write the check that would have caught each defect that reached users before you release the fix, and add one check for existing behaviour whenever you are about to change an area, so coverage goes where the code is actually changing.

Do fast unit tests or slow integration tests matter more?

Both, in different tiers. Fast unit-level checks give an agent or an engineer a feedback loop in minutes, while slower integration checks against real infrastructure are the ones that catch the expensive mistakes, so run the first on every push and the second before merge.

What if the code has no tests at all and the team has no time?

Cover the three flows where a silent failure would cost the most and stop there, which is usually two or three weeks rather than a quarter. A testing project with a coverage goal tends to stop making progress; three protected flows in the pipeline is a finished piece of work.

Does adding checks slow down AI-assisted development?

It speeds it up after the first week, because an agent with a check it can run finds its own failures and iterates without waiting for a person to notice. Without a check, the agent stops as soon as the work looks done and every mistake waits until a person has time to look at it.