Roll it out

Adding evals to a codebase that has none

Adding checks to a codebase that has none fails when it is run as a project with a coverage goal. It works when it is run as a series of small, motivated additions: cover the flows where failure is expensive, cover the areas you are about to change anyway, and turn every incident into a permanent check. Six weeks of that achieves more than three months of a testing project.

Published August 20, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Start by covering the three things that would cause the most harm, then stop and release.
  • Cover the area you are about to modify, just before you modify it. The suite grows as part of real work.
  • Characterisation checks come before refactors: capture what the code does now, then change it safely.
  • Every incident produces a check before it produces a fix. That rule alone builds a good suite in a year.
  • Do not backfill checks for code you are about to delete or replace.

Most codebases we are handed have some tests. Few have tests anyone trusts, and a growing share were built fast with AI tools and have almost nothing. The instinct in that situation is to launch a testing initiative with a coverage target. We have watched several of those and they share an ending: three weeks of enthusiasm, a large number of low-value tests around the easy code, and then the work is dropped, without any announcement, when a deadline arrives.

Here is the approach that still works when a product plan and its deadlines put pressure on the team.

Week one: find out what is actually true

Before writing anything, work out where the risk is. Three questions get you most of the way.

What has broken in the last six months? Read the incident history, the support escalations, the hotfixes. Real failure history is worth more than any theory about where the bugs are.

Where would a failure that nobody notices cost the most? Money, permissions, customer data, the numbers customers make decisions from, the integrations other systems trust.

What does the existing suite actually check? Pick five of the most important-looking tests and break the code they protect. If they still pass, you have learned something important about how much of your current comfort is real.

That is a day of work and it usually reorders everyone's priorities.

Weeks two and three: cover the top three flows

Take the three highest-cost flows and write the invariants down as sentences, then turn each into a check, using the method in how to write your first eval suite. Get them into the pipeline as blocking checks before you write anything else.

Two rules keep this from expanding into the project you are trying to avoid. Test at the boundary a caller uses rather than the internals, because the internals of untested code are usually complicated and mixed together, and testing them requires refactoring first. And accept slower, uglier checks here than you would write on a new build: an integration check that starts a real database and takes 40 seconds is fine if it proves that permissions work.

Then stop. Release something. The suite exists, it blocks merges, and it covers the parts that matter. That is a finished piece of work, and finishing it is what wins support for the next round.

From then on: cover what you change

Adopt one habit and let the suite grow without a project. Before changing an area, write one check for the behaviour that already exists there. Then make your change.

This is cheap, because you are already reading that code. It is well-motivated, because you are about to risk breaking it. And it concentrates coverage exactly where the code is changing, which is where the risk is. Over a year the covered share of your codebase ends up matching how much of it is in active use, which is the right result.

Before a refactor: characterisation checks

When you plan to restructure something, capture what it currently does first, including the odd behaviours. Run the real inputs through it, save the outputs as fixtures, and assert that the new version produces the same results.

You are asserting only that you are not changing the current behaviour by accident, whether or not that behaviour is correct. That distinction is what makes a large refactor safe, and it is the single highest-value technique on a legacy system. If a captured behaviour turns out to be a bug, that is a separate, deliberate change with its own check.

This is also the technique that makes an AI-assisted refactor viable at all. An agent can restructure a module aggressively when there is a fixture-based check proving the outputs did not change. Without one, nobody should be letting anything restructure that module, machine or human.

Forever: incidents become checks

The rule that matters most in the long run. When something reaches a customer that should not have, the check that would have caught it gets written before the fix is released.

It costs almost nothing in the moment, because you have just finished understanding the failure in detail. And its value grows over time: a year of this produces a suite that matches your product's real weaknesses, which no amount of upfront planning can achieve.

What to skip

Code you are about to delete or replace. Backfilling checks for a module that will be removed in two months is pure waste.

Generated boilerplate with no logic in it: mappers, plain data classes, config.

Anything cosmetic. Copy, spacing, colour choices that change for ordinary design reasons.

Aiming for a coverage number anywhere. It moves effort towards the code that is easy to test, which is almost never the code that causes your problems.

If the codebase was built with AI and nobody knows what is in it

This is now a common starting point, and it needs one addition to the plan above. Before adding checks, find out what the code actually does at the boundaries where it can cause harm: authentication, permissions, anything touching money or personal data. Generated code tends to be strong in normal use and weak exactly there, and Veracode's spring 2026 testing across more than 150 models found only 55 percent of generations produced secure code while syntax correctness ran above 95 percent [1].

So the first checks on an AI-built codebase are usually security and permission assertions rather than behaviour tests, because that is where the undiscovered problems concentrate. The wider process for deciding what to fix first in a codebase in that state, including what to fix in what order, is in what it takes to fix a vibe-coded app.

A worked walkthrough of week one

Say a team inherits an order management service with 40,000 lines of code, a test folder that has not been touched in eight months, and no one left on the team who wrote the original permission logic. Reading the incident history turns up three real problems from the last six months: one customer was briefly able to see another customer's order total in an API response, one refund was processed twice because of a retry with no idempotency check, and one report exporter silently dropped rows when a field contained a comma. None of those three needed advanced testing theory to find. They were sitting in a support ticket log the whole time.

Picking five of the existing tests and breaking the code they protect turns up a second finding: three of the five still pass after the underlying function is changed to always return an empty result. The tests were asserting that a call did not throw, not that it returned the right thing. That single afternoon of work produces a more accurate picture of the codebase's real risk than a week of reading the code cover to cover would have, because it is built from what actually went wrong rather than from a guess about what might.

Why the three-week limit matters more than the three flows

The specific flows chosen in weeks two and three matter less than the discipline of stopping after them. Teams that succeed at this almost always describe the same feeling partway through week two: a strong urge to also cover a fourth flow that looks almost as risky, and then a fifth. Resisting that urge is the actual skill being practiced here, because the goal of the first pass is not maximum coverage. It is proving, quickly and visibly, that a blocking check can exist in this codebase and survive contact with a real deadline. A team that ships three solid checks in three weeks has evidence the approach works. A team that is still going in week six, chasing an eighth flow, has quietly turned the staged plan back into the coverage project it was designed to avoid.

What a characterisation check catches that nobody expected

Characterisation checks are worth a longer look because their value is often larger than the refactor they were written to protect. Capturing the actual outputs of an old pricing function, rather than what the documentation says it should do, regularly turns up behaviour nobody remembers deciding on. A discount that only applies on the third Tuesday of certain months because of an off-by-one in a date calculation written years ago. A currency rounding rule that differs by half a cent depending on which code path a request happened to take. None of this shows up by reading the code, because reading code shows you what it is trying to do, and a characterisation check shows you what it is actually doing, which are not always the same thing after several years of small changes by different people.

This is also where a genuine decision point appears. Once a captured behaviour is identified as a bug rather than an intended quirk, the temptation is to fix it while you are already in the code doing the refactor. Resist that combination. Fix the refactor first, with the characterisation check proving nothing changed, ship that safely, and treat the bug fix as its own separate, deliberate change with its own eval added for the corrected behaviour. Combining the two means that if anything goes wrong after release, there are two candidate causes tangled into one change instead of one cause you can point to directly.

How this interacts with an agent doing the refactor itself

If a coding agent is doing the restructuring rather than a person, the characterisation check changes from a safety net into something closer to a specification. An agent asked to restructure a module with a fixture-based check already in place has a concrete target: keep producing the same fixture outputs while changing the internal shape however it judges is best. That is a well-posed task an agent can iterate against unattended. An agent asked to restructure the same module with no check in place has an ambiguous task with no way to know if it succeeded, and the honest answer to "did this refactor preserve behaviour" is that nobody, including the agent, can say. Writing the characterisation check first is what turns an unsupervisable request into a supervisable one, which is the same argument made about the eval suite generally in what is eval-driven development.

Best for

  • Legacy systems with tests nobody trusts
  • AI-built codebases where nobody can say what is guaranteed
  • Teams that need coverage to grow without pausing the product plan

Avoid if

  • Do not launch a testing project with a coverage target: those stop before they finish
  • Do not backfill checks for code you plan to delete or replace
  • Do not refactor untested legacy code before capturing its current behaviour

Check before you decide

  • Confirm five existing tests actually fail when you break what they protect
  • Confirm the incident history, rather than intuition, chose your first three flows
  • Confirm the new checks block merges from the week they are written

Common questions

How do I add tests to a codebase that has none?

Do not run it as a testing project with a coverage goal, because those stop before they finish. Read the incident history to find where failure actually happens, cover the three highest-cost flows with blocking checks over two or three weeks, then grow the suite by covering each area just before you change it.

What is a characterisation check?

It captures what code currently does, including its unusual behaviours, so you can restructure it without changing behaviour by accident. You run real inputs through the existing code, save the outputs as fixtures, and assert the new version matches, which is what makes a large refactor safe enough to give to an agent.

Should we aim for a coverage percentage on a legacy codebase?

No. A coverage target pulls effort towards the code that is easiest to test, which is rarely the code that causes incidents. Cover by cost of failure and by what you are about to modify instead.

Where do you start on a codebase built quickly with AI tools?

With security and permission assertions rather than behaviour tests. Generated code is usually competent in normal use and weakest at the boundaries where failures cause harm, and Veracode's spring 2026 testing found only 55 percent of AI generations were secure while more than 95 percent compiled.

Why do coverage-target testing initiatives usually fail?

Because they get run as a project rather than as motivated additions to real work, which produces a predictable pattern: a few weeks of enthusiasm, a large number of low-value tests around the code that was easiest to test, and the work being dropped, without any announcement, once a deadline arrives. Covering by cost of failure and by what is about to change avoids that pattern entirely.

How do I decide which three flows to cover first in a legacy codebase?

Read the incident history rather than guessing: what has actually broken in the last six months, where a failure nobody notices would cost the most, and what the existing suite actually checks once you break the code it claims to protect. That review is usually a day of work and it reliably reorders the team's assumptions about where the risk is.

Is it worth writing evals for code that is about to be deleted?

No. Backfilling checks for a module that is removed from the codebase within a couple of months is pure waste, along with adding checks to generated boilerplate with no real logic, such as mappers or plain data classes, and to anything cosmetic that changes for ordinary design reasons.

How does a characterisation check make an AI-driven refactor safe?

It captures the real inputs and outputs of the existing code as fixtures before anything changes, so an agent can restructure the module aggressively while the fixture-based check proves the outputs did not change. Without that proof, letting anything, machine or human, restructure an untested module is a guess rather than a safe change.