Eval-driven development: how to prove AI-written code works / Build the eval suite
How to write your first eval suite
You do not need a testing strategy document to start. You need one flow where a failure nobody notices would be expensive, a short list of sentences describing what must always be true about it, and each of those sentences turned into a check that runs on every change. That is a real eval suite, it takes a day or two, and it protects more than a month of trying to raise a coverage number.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Start with the one flow where a failure nobody notices costs the most, rather than with the code that is easiest to test.
- Write the invariants as plain sentences first. If you cannot state it, you cannot check it, and that is a finding about your spec.
- Make each check fail before you make it pass, so you know it is actually checking something.
- Connect it to the pipeline the same day. A check that does not block a merge is only a note.
- Add to the suite after real incidents. Every problem that reaches users becomes a permanent check.
The most common reason a team has no eval suite is that starting feels like it requires a plan for the whole codebase. It does not. Here is the version that works, in the order we do it.
Step 1: pick the flow where a failure nobody notices is most expensive
Pick the one where a wrong answer that nobody notices for a week does the most damage, even if it is harder to test or has less code than others.
For most products that is one of five things: authentication and permissions, anything that moves money, anything that writes or deletes customer data, the export or report that customers make decisions from, or the integration that other systems trust. Pick one. Just one.
The reason to choose by cost rather than by convenience is that the first suite has to win support inside the team. A suite that catches something real in its first fortnight gets extended. A suite that proves your date formatter works gets abandoned.
Step 2: write the invariants as sentences
Before any code, write down what must always be true. Plain sentences, the kind you would say out loud to a colleague. For a permissions boundary that might be:
A user can only read records belonging to their own organisation. A revoked session cannot perform any write. An admin action on another organisation's data is refused and logged. A deleted user's tokens stop working immediately, rather than at the next refresh.
Four sentences. You now have four evals, and you have also done the more valuable thing, which is discover whether anyone actually knows the answer. Almost every time we run this exercise, at least one sentence produces a disagreement in the meeting. That disagreement was already in your codebase, undecided, being resolved differently by whoever last changed it. Finding it costs an hour here and costs an incident later.
If the sentences are hard to write, the problem is earlier, in the specification. Writing specs an agent can verify is about fixing that part.
Step 3: turn each sentence into a check that fails first
Write the check, then run it against the current code and watch what happens.
If it fails, good. You have found a real gap and you have proof the check works. If it passes immediately, break something on purpose and confirm it fails. A check that has never failed is a check you cannot trust, and a surprising share of assertions in real codebases cannot fail at all: the mock returns the expected value, the setup skips the branch, the assertion compares a variable to itself.
Keep each check at the level the sentence was written at. If the sentence is about a user reading records, drive it through the same entry point a user does: the API handler or the service function, rather than three private methods. Testing at the level of the promise is what keeps the suite working through refactors.
Step 4: use real data rather than invented data
The cases that break software in production are usually different from the cases people invent at a desk. Names with apostrophes, empty arrays where the code expects one item, a timezone offset that pushes a date into the previous month, a string long enough to hit a column limit, a unicode emoji in a field someone assumed was ASCII.
Anthropic's guidance on eval design says the same thing about model evals and it applies exactly: mirror the real task distribution and do not forget the unusual cases, including irrelevant or nonexistent data, over-long input, and cases where even a human would hesitate [1]. Pull your fixtures from real traffic where you can, with the sensitive fields replaced. Ten real cases are worth more than a hundred invented ones.
Step 5: make it block the merge on the same day
This is the step teams skip and it is the step that makes the difference. A check that runs when someone remembers is documentation. Put it in the pipeline, make a failure block the merge, and do it before you move on to the next flow, because a suite that has never blocked anything will not be maintained.
If your pipeline is slow, split it: the fast checks on every push, the slow ones before merge. What matters is that no change reaches the main branch without the relevant checks having passed. The mechanics are in evals in CI for coding agents.
Step 6: add checks after incidents rather than from a big plan
Once the first flow is covered, avoid planning the rest. Adopt one rule instead: every time something reaches a user that should not have, the eval that would have caught it gets written before the fix.
That rule does two useful things. It keeps the suite proportional to your product's actual failure modes rather than to someone's idea of good coverage. And it makes the suite a record of everything you have learned from real failures, which is the most valuable document a team owns and the one nobody ever writes deliberately.
Do the same when you work on an area of the codebase for other reasons. Adding a feature to a module with no checks? Write one invariant for what is already there before you touch it. The suite grows as part of your normal work instead of needing a project of its own.
What a good first suite looks like
Small. Between five and twenty checks. All of them on things that would cause harm. All of them fast enough that nobody resents running them. Every one has failed at least once in real use, so the team knows they are real. None of them flaky, because the first flaky check is when the team starts to ignore the suite.
That is the whole thing. It does not require a framework decision, a coverage tool, or a new job title. It requires picking the flow that matters and writing down what must be true about it, which is work nobody can do for you and which turns out to be most of the value.
Where to go next
What to check in an eval suite is the checklist for the second and third flows. Adding evals to an existing codebase covers the staged approach for a large system with years of untested code. And if you want the case for spending the time, the business case for evals is written for the person approving it.
Common questions
How do I start an eval suite from nothing?
Pick the single flow where a failure nobody notices would cost the most, write down as plain sentences what must always be true about it, and turn each sentence into an automated check that you first watch fail. Then connect those checks to your pipeline so a failure blocks the merge, which usually takes a day or two in total.
How many checks should a first eval suite have?
Between five and twenty is normal, and small is the point. Each one should cover something that would cause real harm if it broke, run fast enough that nobody avoids it, and have failed at least once in real conditions so the team knows it works.
Why should a new check fail before it passes?
Because a check that has never failed might not be capable of failing. Mocks that return the expected value, setup code that skips the branch, and assertions that compare a variable to itself are common, and the only way to know an assertion is checking something real is to see it fail on purpose.
Where should test data come from?
From real traffic wherever possible, with sensitive fields replaced, because production data contains the kinds of values nobody invents at a desk: unusual names, empty collections, timezone boundaries, oversized strings, unexpected character sets. Ten real cases usually catch more than a hundred imagined ones.
What if we cannot write the invariant as a sentence?
That is the most useful finding the exercise produces, because it means nobody has decided what the correct behaviour is. The disagreement was already in your codebase being resolved differently by whoever last changed that area, and settling it costs an hour now against an incident later.
Should the first eval suite cover the easiest code or the riskiest code?
The riskiest code. Choose the one flow where a wrong answer that nobody notices for a week does the most damage, such as authentication, money movement, or customer data, rather than the code that happens to be simplest to test. A suite that catches something real in its first weeks gets extended; a suite proving a date formatter works gets abandoned.
How long does it take to make a new eval suite block merges?
The blocking step should happen the same day the checks are written. A check that runs only when someone remembers is documentation rather than a check that blocks bad changes, and a suite that has never blocked a merge tends not to get maintained, so connecting it to the pipeline is part of the same day of work as writing the first invariants.
What should drive which checks get added after the first suite?
Real incidents, rather than a plan drawn up in advance. Every time something reaches a user that should not have, the eval that would have caught it gets written before the fix is released. That rule keeps the suite proportional to the product's actual failure modes rather than to a general idea of good coverage.
Related reading
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
More in Build the eval suite
What to check in an eval suite: the seven things that matter
Coverage percentages tell you which lines of code ran during the tests, and nothing about which promises to users are protected. This is the list we work through instead: seven classes of check, ordered by how much damage they prevent, with a note on what is not worth automating. Most products need all seven eventually and only two or three of them on day one.
Writing specs an AI agent can verify
When a machine writes the implementation, the specification stops being a document people skim and becomes the actual input to the work. Vague specs used to produce slow projects. Now they produce large amounts of wrong code that looks right, fast. A spec that works has four parts: the behaviour stated as testable sentences, the boundaries named, the out-of-scope list written down, and an end-to-end check that proves the whole thing.
Using a model as a judge, without fooling yourself
Some things you want to check have no single correct string to compare against: the quality of an error message, whether a diff matches its spec, whether generated documentation is accurate. A model can grade those, and it is a genuinely useful eval when it is set up with two rules: the judge must be independent of the thing it grades, and the judge itself has to be checked against human judgment on a sample.