Eval-driven development: how to prove AI-written code works / Build the eval suite
What to check in an eval suite: the seven things that matter
Coverage percentages tell you which lines of code ran during the tests, and nothing about which promises to users are protected. This is the list we work through instead: seven classes of check, ordered by how much damage they prevent, with a note on what is not worth automating. Most products need all seven eventually and only two or three of them on day one.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Order the work by the damage each check prevents.
- Permissions and data invariants are the checks that most often do not exist and most often matter.
- Security-relevant paths need explicit checks, because generated code is measurably weakest there.
- Contract checks are what stop one team's refactor from breaking another team's product.
- Skip checks on things that change every week for cosmetic reasons. A check nobody trusts costs more than it saves.
A suite is better when it protects the things that would cause harm. Here are the seven classes we work through, in the order we usually add them.
1. Behaviour, at the level a caller sees it
The first class. For each significant flow, assert what a user or a calling system gets back for a valid request. Drive it through the same entry point the real caller uses, rather than through internal functions, so the check stays true through a refactor.
This is the least glamorous class and it catches the most, because it is what most changes accidentally alter.
2. Contracts at the boundary
Anywhere your code passes something to a system that did not build it: an API response shape, an event payload, a database schema, a file format, a webhook. These deserve explicit checks that assert the required fields exist, the types are what they claim, and nothing that used to be there disappeared without notice.
Boundary breakage is the most expensive kind, because the failure appears in someone else's system, hours later, in a way nobody can trace back to a diff. A schema assertion in your suite turns that into a failed build in the pull request that caused it.
3. Unusual cases and error paths
Normal use is the part that always works, because it is the part everyone tries. The value is in the rest: empty collections, missing optional fields, boundary numbers, duplicate submissions, expired tokens, a dependency returning an error, the same request arriving twice.
Anthropic's eval design guidance lists exactly this category for model evals and the same list applies to code: irrelevant or nonexistent input, over-long input, badly formed input, and cases ambiguous enough that a human would hesitate [1]. Generated code is particularly weak here, since the training signal is full of examples that handle the normal case and stop.
Also assert the error behaviour itself. Check that a bad request is refused with the right status, without leaking internal detail, and without a partial write left behind.
4. Permissions and tenancy
The class most often missing and most damaging when it fails. For every resource, assert who can read it, who can write it, and what happens to someone who should not be able to do either. If your product has organisations, teams, or tenants, assert that data from one cannot be reached from another, including through search, exports, aggregate counts, and error messages.
Write the negative cases explicitly. A check that an authorised user succeeds is worth much less than a check that an unauthorised one fails, and the second is the one nobody writes.
5. Data invariants and migrations
Things that must remain true about your stored data whatever the code does. A balance never goes negative. A record always has an owner. A soft-deleted row never appears in a list. Two rows never claim the same unique slot. Statuses only move in permitted directions.
Migrations deserve their own check: run the migration against a copy of data with the same structure as production, then assert the invariants still hold and the row counts match expectations. A migration is the change class where a mistake is hardest to undo, which makes it the one most worth blocking with a check.
6. Security-relevant paths
Give the common flaw classes their own explicit checks rather than trusting that they were handled. Injection at every place a string becomes a query or a command. Output escaping wherever user text is rendered. Authentication on every endpoint that should have it, asserted by testing that the unauthenticated call is refused. Secrets absent from logs and error responses.
The reason to be explicit is measurable. Veracode's spring 2026 testing across more than 150 models found only 55 percent of generations produced secure code, with Java at 29 percent and cross-site scripting handled correctly in roughly 13 to 15 percent of relevant cases, while syntax correctness ran above 95 percent [2]. That difference has stayed the same for two years. Generated code will compile and will not necessarily be safe, so safety needs its own assertions.
7. Performance thresholds where they are part of the promise
Not micro-benchmarks. Hard thresholds on the paths where slowness is a defect: the query that runs on every page load, the export that has to finish inside a request timeout, the job that must complete before the next one starts. Assert the number of queries too, since the most common performance regression in generated code is a loop that issues one query per row without anyone noticing.
What not to bother automating
Three things, and being deliberate about this is what keeps a suite trusted.
Anything cosmetic that changes weekly. Exact copy, exact spacing, exact colour. Assert that a required element is present and reachable, rather than that a heading reads exactly one way; otherwise the suite becomes a cost on ordinary design work.
Implementation detail. How many times a private method was called, whether a particular class exists, the internal order of operations. These break on refactors and teach the team that failures mean nothing.
Anything you cannot make deterministic without a lot of work. A check that fails one run in twenty is worse than no check, because it trains everyone to run it again and it devalues every other check. Either fix it or take it out of the blocking suite. When evals give false confidence covers this in full.
How to sequence it
Two or three classes on day one, chosen by what would cause the most harm: usually behaviour on your main flow, plus permissions if you have multiple tenants, plus data invariants if you handle money. Add the rest as the product and its failures show you where the risk actually is.
For a codebase that already exists with none of this, the staged approach is in adding evals to an existing codebase. For what the checks are worth in business terms, see the business case for evals.
Best for
- Behaviour, contracts, unusual cases, permissions, data invariants, security paths, and hard performance thresholds
- Any promise made to a system or a team that did not write the code
- Any invariant whose breach would be expensive or slow to notice
Avoid if
- Do not assert exact copy, spacing, or colour that changes for ordinary design reasons
- Do not assert private call counts or internal structure: those break on every refactor
- Do not keep a check you cannot make deterministic in the blocking suite
Check before you decide
- Confirm the negative permission cases are asserted, as well as the positive ones
- Confirm migrations are checked against data with the same structure as production before they run
- Confirm the security classes have their own checks rather than being assumed
Common questions
What should an eval suite actually check?
Seven classes, in order of the damage each prevents: user-visible behaviour, contracts at every boundary, unusual cases and error paths, permissions and tenancy, data invariants including migrations, security-relevant paths, and hard performance thresholds where slowness counts as a defect. Most products need two or three of these on day one and the rest as failures show them where the risk is.
Is test coverage a good measure of an eval suite?
No. Coverage counts which lines executed, whether or not any promise is protected, so it is possible to reach a high number while asserting almost nothing of value. Judge a suite by whether it would have caught your last three incidents.
Why do permission checks matter so much?
Because they are the class most often missing and among the most damaging when they fail, and because the useful cases are the negative ones. A check that an authorised user succeeds proves little; a check that an unauthorised user is refused, including through search, exports, and error messages, is the one that prevents a data exposure.
Should we write security checks if we already run a scanner?
Yes, for the paths that matter, because a scanner finds known patterns and an eval asserts your specific promise. Veracode's testing found AI-generated code was secure in only about 55 percent of generations while compiling more than 95 percent of the time, so the assumption that working code is safe code does not hold.
What should not go in an eval suite?
Anything cosmetic that changes weekly, anything that asserts internal implementation detail, and anything you cannot make deterministic. Each of those produces failures that mean nothing, and a suite that keeps giving false alarms trains the team to ignore every check.
What is a contract check and why does it matter?
A contract check asserts that anything your code passes to a system that did not build it, such as an API response, an event payload, or a webhook, keeps the required fields, the correct types, and no field that disappeared without notice. Boundary breakage is expensive because the failure appears in someone else's system hours later, and a schema assertion in your suite turns that into a failed build in the pull request that caused it.
How should migrations be checked before they run?
Run the migration against a copy of data with the same structure as production, then assert the data invariants still hold and the row counts match expectations. A migration is the change class where a mistake is hardest to undo, which makes it one of the most worthwhile classes of check to block with before it ever runs on real data.
Should performance be part of an eval suite?
Yes, but as hard thresholds on the paths where slowness counts as a defect, such as a query that runs on every page load or an export that must finish inside a request timeout, rather than as general micro-benchmarks. It is also worth asserting the number of queries a path issues, since the most common performance regression in generated code is a loop that issues one query per row.
References
- [1] Anthropic, Create strong empirical evaluations: edge cases to include, such as irrelevant or nonexistent input data, over-long input, poor or irrelevant input, and ambiguous test cases where even humans would struggle.
- [2] Veracode, Spring 2026 GenAI Code Security update: 55% of generations secure overall, Java at 29%, cross-site scripting handled correctly in roughly 13 to 15% of cases, syntax correctness above 95%.
Related reading
A practical pre-launch security review for a small team
You do not need perfect security to launch. You need to check the few basics that find most real problems, and to know when the risk is big enough to bring in a specialist.
How many tests does AI-generated code need?
The honest answer is not a number, and coverage percentages are the wrong unit. Here is the unit we use instead, and why a suite of fifteen checks can be worth more than two thousand.
More in Build the eval suite
How to write your first eval suite
You do not need a testing strategy document to start. You need one flow where a failure nobody notices would be expensive, a short list of sentences describing what must always be true about it, and each of those sentences turned into a check that runs on every change. That is a real eval suite, it takes a day or two, and it protects more than a month of trying to raise a coverage number.
Writing specs an AI agent can verify
When a machine writes the implementation, the specification stops being a document people skim and becomes the actual input to the work. Vague specs used to produce slow projects. Now they produce large amounts of wrong code that looks right, fast. A spec that works has four parts: the behaviour stated as testable sentences, the boundaries named, the out-of-scope list written down, and an end-to-end check that proves the whole thing.
Using a model as a judge, without fooling yourself
Some things you want to check have no single correct string to compare against: the quality of an error message, whether a diff matches its spec, whether generated documentation is accurate. A model can grade those, and it is a genuinely useful eval when it is set up with two rules: the judge must be independent of the thing it grades, and the judge itself has to be checked against human judgment on a sample.