Guide

Eval-driven development: how to prove AI-written code works

AI can write a feature in minutes. Proving that the feature does what you asked still takes real work, and that work is now the only thing that stops a fast team from releasing a broken product.

Published August 20, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Eval-driven development means the check comes first. You define how a change will be proven correct, then let the AI write code until the check passes.
  • Writing code now costs close to nothing, and checking it still takes work. That difference is why teams that adopt AI coding tools often feel faster while producing more work that has to be redone.
  • Evals and code review answer different questions. Review catches bad design once. An eval catches a behaviour change every time anyone changes the code, including next year.
  • An eval has to run automatically on every change to work as a check; otherwise it is only documentation. Its value comes from blocking the changes that fail.
  • The number worth tracking is how much of what you release has to be fixed or redone: defects that reach users, reverted changes, and rework.

Adding eval tests was the best engineering decision we have made. We say that as a company that generates one hundred percent of its code, which makes us exactly the sort of team that could release a large volume of software that looks right and is wrong. The evals are the reason we do not.

Here is the short version. When a machine writes the code, the writing stops being the hard part. What is left is deciding what "correct" means and proving it, on every change, without a person having to remember. That is what evals do, and building them first changes the order of the work: you write the check, then you let the AI write code until the check passes. This guide covers what evals are, why they work better than both manual review and traditional test-after habits for AI-written code, how to build a suite from nothing, how to run it in CI, and how to measure whether it is working.

What eval-driven development actually is

An eval is an automated check that decides whether a piece of software behaves the way the specification says. It runs, it produces a pass or a fail, and it needs no human judgment to interpret. That sounds like a test, and a unit test is one kind of eval. The word comes from AI research, where model quality is measured against a fixed set of cases with known answers, and it has a useful difference in emphasis: an eval is written to prove a requirement, where a test is often written to cover a function.

Eval-driven development is the practice of writing that check before the implementation, then treating the check as the definition of done. The specification says what should be true. The eval encodes it. The AI writes code, runs the eval, reads the failure, and tries again. A human reviews the diff once the check passes, which means the human is reading working code rather than checking whether it works at all.

The full definition, including where the idea comes from and how it differs from test-driven development, is in what is eval-driven development.

Why the slowest step changed

For thirty years, typing was the slow part of building software. It is not any more, and most teams have not reorganised around that.

The evidence that the change is real, and that it does not automatically help, is now strong. Google's DORA program surveyed nearly 5,000 technology professionals for its 2025 report on AI-assisted software development and found that 90 percent use AI at work and more than 80 percent believe it makes them more productive. The same study found AI adoption has a positive relationship with delivery throughput and a negative relationship with delivery stability [1]. More code, arriving faster, breaking more often.

The second piece is what happens to individual judgment. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. The developers using AI tools took 19 percent longer, and afterwards estimated that AI had made them 20 percent faster [2]. METR is careful to label the result as historical, since the tools have changed since then, and we are careful to repeat that. The lasting finding is that the people doing the work could not tell which way the number went. If a senior engineer's sense of speed is unreliable, a senior engineer's sense of quality is not going to be better, and neither belongs in the place where you decide whether a change is safe.

The third piece is what the generated code actually contains. Veracode's spring 2026 update tested more than 150 language models against 80 coding tasks and found that only 55 percent of generations produced secure code, a figure that has stayed almost the same for two years while syntax correctness rose above 95 percent [3]. Java came out at 29 percent secure. Cross-site scripting was handled correctly in roughly one generation in seven. Models have become excellent at writing code that compiles and have not become good at writing code that is safe.

And developers know something is wrong. In Stack Overflow's 2025 survey, 84 percent of respondents use or plan to use AI tools, while 46 percent said they actively distrust the accuracy of the output, against 33 percent who trust it. The single most cited frustration, at 66 percent, was "AI solutions that are almost right, but not quite" [4].

Almost right is the whole problem. Code that is almost right passes a quick read. It also works in a demo. It is exactly the kind of failure that a human reviewer, reading fast on a Friday, is worst at catching and that a machine check catches every time. We cover this in more detail in the verification gap in AI coding.

Evals come from AI research, and now they apply to code

Nobody in AI research releases a model without an eval set. When OpenAI wanted a trustworthy measure of whether models can fix real software issues, the answer was SWE-bench Verified: 500 real GitHub issues, each reviewed by human annotators, each paired with a test patch that decides whether a proposed fix counts [5]. No opinions, no demo, a fixed set of cases and a pass rate.

That is the same idea, applied to models instead of products. Anthropic's own guidance on designing evals gives three rules that apply directly to code: make them task-specific and include the unusual cases, automate the grading, and prefer more cases with less precise automated scoring over a handful of carefully hand-graded ones [6]. A large number of automated results is worth more than a small amount of human attention. That is a statement about your test suite as much as about a model.

The practical version, for anyone using a coding agent on a repository, is in Anthropic's engineering guidance for Claude Code, which names the failure directly. The trust-then-verify gap: the model produces a plausible-looking implementation that does not handle unusual cases. The fix it gives is one line long. Always provide verification, and if you cannot verify it, do not release it [7]. An agent stops when the work looks done, and without a check it can run, "looks done" is the only signal it has. Give it a check and it can keep trying until the check passes, without you.

That last point is the one most teams miss. An eval suite protects your users, and it is also the feedback the agent needs to do the job properly in the first place. A repository with good automated checks makes AI better at writing code, because the AI gets to find out it was wrong before you do.

The four jobs an eval suite has to do

Across the work we do, a useful suite does four things. If one is missing, the suite starts to feel like a formality.

It encodes the spec. The eval should assert what the feature must do for a user or a caller. If it asserts how the current code happens to be structured, it will break on every refactor and teach your team to ignore failures. That is how teams stop using their suites.

It fails clearly on the cases that cost money. Authentication, permissions, money movement, data deletion, anything irreversible. Here, what matters is a named list of risks rather than a coverage percentage. What to check in an eval suite is the list we work through.

It runs on every change, automatically, with the power to block. An eval that runs when someone remembers is not a check. See evals in CI for coding agents.

It is trusted. A suite with three known-flaky tests is worse than a suite with none, because the team learns to re-run failed builds instead of reading them. When evals give false confidence covers the ways a passing build can be wrong.

What this changes about a working week

The work in a day is different, and it is worth being concrete about it.

The spec gets more attention than it used to, because it is now an input to a machine rather than a document nobody reopens. Ambiguity in the spec turns directly into wrong code, quickly, in volume. Writing the eval alongside the spec is the cheapest way to find out that the spec was vague, since you cannot encode a requirement you cannot state. Writing specs an agent can verify is about that discipline.

The implementation step gets faster and less interesting. This is the part people expected AI to change and it did.

Review changes the most. A reviewer reading a diff that already passes a real eval suite is doing the thing humans are good at: asking whether this was the right thing to build, whether the design will cause problems in six months, whether a case is missing from the check itself. A reviewer reading an unverified diff is doing the thing humans are bad at, which is simulating a computer. Our full position on that division of labour is in evals vs tests vs code review, and the human half is covered in the companion guide on reviewing AI-generated code.

Where teams get this wrong

Four patterns show up again and again.

Aiming only for a coverage number. Coverage counts the lines of code that ran during tests, whether or not any check looked at the result. It is possible to reach 90 percent while proving almost nothing, and teams that aim for the number usually do.

Letting the model grade its own work. If the same agent writes both the implementation and the check, in the same run, you have a system that agrees with itself. Anthropic's own guidance suggests a fresh context or a second model precisely so that the thing doing the work is not the thing grading it [7]. We treat that as a rule: the check must come from the spec, before or independently of the code that satisfies it. LLM as judge for code review covers how to use a model as a grader without this failure.

Adding evals to everything at once. On an existing codebase this effort stops before it finishes, every time. Start at the parts where a failure costs the most and add to the suite after each real incident. Adding evals to an existing codebase is the staged version.

Measuring the wrong thing. Volume of code released stopped being a useful measure once code became cheap to produce. Defects that reach users, revert rate, and time to restore service still mean something. Metrics for AI code quality covers what we watch instead.

What to do first

If you want one thing to do this week, do this. Pick the flow in your product where a failure that nobody notices would be most expensive. Write down, in plain sentences, what must always be true about it. Turn each sentence into an automated check. Connect those checks to your pipeline so a failure blocks the merge. That is a real eval suite for the part of your system that matters most, and it is usually a day or two of work.

Then let real problems give you the next case. Every time something reaches a user that should not have, write the eval that would have caught it before you write the fix. A suite grown that way stays small, stays trusted, and covers exactly the failures your product actually has. The starting version is in how to write your first eval suite.

What a specification has to contain to be checkable

A spec written for a person can stay vague in places and still work, because the person filling gaps brings judgment the document does not need to spell out. A spec written for a machine cannot skip that step, because the machine has no judgment to fall back on. It will build exactly what the sentence says, including the parts the sentence left out by accident.

This changes what a good specification looks like. Every stated behavior needs a boundary that covers what the login form does when the password field is empty, when the network call times out, and when two requests arrive for the same account at once, in addition to the normal case. A sentence that reads clearly to a person and still leaves the failure path unstated is not finished, because there is no reader left to guess at it. Writing specs an agent can verify covers the concrete rewriting habit that catches this: read every requirement and ask what a test for it would actually assert, and if no test comes to mind, the requirement is not written down yet, only implied.

The upside is not extra paperwork. A specification precise enough to generate an eval from is also precise enough that two people reading it agree on what "done" means, which used to take a meeting and now takes reading the same sentence.

Why a suite that never fails is not protecting you

A team that runs its eval suite for a year and never sees it turn red usually assumes the code has been correct that whole time. The more common explanation is that the suite cannot fail, because it never asserts anything precise enough to catch a real mistake.

This shows up in a specific way. An eval that checks a function returns "a result" rather than a stated value passes whether the value is right or wrong. An eval that runs against a mock instead of the real dependency it is meant to protect will keep passing after that dependency's contract changes underneath it. Both look like coverage. Neither is protection, and a team relying on either finds out only when a customer does.

The test worth running on any suite you inherit or build is deliberate: break the behavior on purpose, in a way a user would notice, and confirm the suite turns red before you fix it back. If nothing fails, you have found a hole in your own safety net while it still costs nothing to patch. When evals give false confidence walks through the specific shapes this failure takes and how to catch each one before it matters.

Agent debt, and why it is different from technical debt

Technical debt is the cost of a shortcut a person took on purpose, usually because a deadline was closer than a clean solution. It accumulates at the speed a team writes code, which is a speed a person can feel and plan around.

Agent debt is different. A coding agent will happily generate weeks of plausible-looking implementation in an afternoon, and every line of it needs the same scrutiny a person's code would need, arriving far faster than any review process built around human writing speed was designed to handle. The debt is a backlog of unverified work that grew because generation and verification stopped moving at the same speed, not a shortcut anyone chose on purpose.

An eval suite is the only mechanism that scales with generation instead of falling behind it, because a check runs in seconds no matter how much code it is checking, while a person reading the same volume does not get faster. This is the direct, practical argument for building the suite before the volume of generated code outgrows what any reviewer could keep up with by hand. Agent debt: when generation outpaces review covers what this backlog looks like once it has already built up and how a team works it back down without stopping new work entirely.

Bringing a suite to a codebase that already exists

Almost nobody starts an eval suite on an empty repository. Most teams reading this guide are looking at a codebase that already runs, already has users, and has no checks that would catch a real mistake before it ships.

Trying to cover that codebase in one pass is the most common way this effort dies. A team commits to writing evals for everything, the work competes with every feature request for the same hours, and six weeks later the suite covers whatever was easiest to test first rather than what actually matters. The size of the gap is not the problem. Treating it as one project is.

The version that survives contact with a real roadmap is staged. Pick the single flow where a silent failure would cost the most, whether that is money movement, account access, or the one integration every customer depends on, and write checks for that flow first. Add to the suite every time a real incident happens, using the bug itself as the specification for the check that would have caught it. This grows a suite that tracks your actual risk instead of an arbitrary coverage target, and it fits inside a normal week instead of requiring one. Adding evals to a codebase that has none is the full staged version of this approach.

What to measure once the suite exists

Building the suite answers whether a change is correct before it ships. It does not by itself answer whether the practice is working, and that is a different question with a different, easy-to-get-wrong answer.

The tempting metric is volume: lines of code generated, pull requests merged, features shipped in a sprint. Every one of those got cheap to produce the moment a model started writing the code, which means every one of them stopped measuring effort and started measuring how fast the model typed. None of them says anything about whether what shipped was correct.

The metrics that still mean something are the ones tied to what happens after release: how often a defect reaches a real user, how often a change has to be reverted, how long it takes to restore service after something breaks. These move in the opposite direction from volume when a suite is doing its job, because catching a mistake before release is exactly what a good eval suite is for. Watching only volume while these get worse is how a team convinces itself that going faster is the same as doing well. Metrics for AI code quality covers the specific set worth tracking and why each one resists being gamed the way a coverage percentage does.

Where Reveneau fits

We generate all of our code, and a large eval suite has to prove every change against the specification before it reaches your branch. That is the whole reason we can work as fast as AI writes code without losing quality: the check is written first, it runs on every change, and a change that cannot pass it is not released.

We will not tell you that evals cut our defect rate by some percentage, because we have not run the controlled experiment that would let us say it, and a consultancy selling verification should not be the one publishing numbers it cannot show the working for. What we will tell you is what changed in practice: the failures we find now are design disagreements caught in a pull request, before a customer sees them.

This guide covers evals for code. The same practice applied to the answers an AI product gives has its own guides: why AI evals matter for the person who approves the budget, LLM evals for the method, AI agent evals for agents that take actions, and AI benchmarks vs your own evals for reading a model's published score.

If you are choosing a partner, asking whether we use AI tells you little, because everyone does. Ask what the check is, who wrote it, whether it runs on every change, and what happens when it fails. The business case for evals is written for the person who has to approve the time, and our AI development work is where we use this practice. If you have a codebase that was built fast and you are not sure what is true about it, talk to us and we will tell you honestly.

Explore the guide

Build the eval suite

How to write your first eval suite

You do not need a testing strategy document to start. You need one flow where a failure nobody notices would be expensive, a short list of sentences describing what must always be true about it, and each of those sentences turned into a check that runs on every change. That is a real eval suite, it takes a day or two, and it protects more than a month of trying to raise a coverage number.

What to check in an eval suite: the seven things that matter

Coverage percentages tell you which lines of code ran during the tests, and nothing about which promises to users are protected. This is the list we work through instead: seven classes of check, ordered by how much damage they prevent, with a note on what is not worth automating. Most products need all seven eventually and only two or three of them on day one.

Writing specs an AI agent can verify

When a machine writes the implementation, the specification stops being a document people skim and becomes the actual input to the work. Vague specs used to produce slow projects. Now they produce large amounts of wrong code that looks right, fast. A spec that works has four parts: the behaviour stated as testable sentences, the boundaries named, the out-of-scope list written down, and an end-to-end check that proves the whole thing.

Using a model as a judge, without fooling yourself

Some things you want to check have no single correct string to compare against: the quality of an error message, whether a diff matches its spec, whether generated documentation is accurate. A model can grade those, and it is a genuinely useful eval when it is set up with two rules: the judge must be independent of the thing it grades, and the judge itself has to be checked against human judgment on a sample.

Run it continuously

Running evals in CI when agents write the code

An eval that does not run on every change is only a note somebody wrote once. The value comes from blocking: no change reaches the main branch unless the relevant checks passed. That is simple to say and has a few real design decisions inside it, mostly about speed, about what an agent is allowed to do with a failed build, and about what happens on the day the check is wrong.

Metrics for AI code quality: what to watch instead of volume

Once code became cheap to produce, every metric based on how much of it you produce stopped telling you anything useful. What still means something is what comes back: how often a change breaks something, how long recovery takes, how many defects reach a customer, and how much of last month's work is being redone. Those four are still useful after the change, and the first two have ten years of research supporting them.

When evals give false confidence

A suite that catches nothing is worse than having no suite, because a team with no checks knows it is exposed and a team with passing checks believes it is covered. Six failure modes account for almost all of it, and each one has a specific sign you can look for this afternoon.

Agent debt: when generation outpaces review

Teams now generate most of their code with an agent and review almost none of it as carefully as before, because the agent is fast enough to make thorough review feel like the slowest step. The cost of that missing review appears later, in production, months after the pull request was merged, as an incident nobody can trace to a specific change.

Common questions

What is eval-driven development?

Eval-driven development is writing the automated check that defines correct behaviour before the code that satisfies it, then letting an AI coding agent work until that check passes. The check encodes a requirement from the specification, runs without human judgment, and stays in the pipeline as a permanent check on every future change.

How is an eval different from a unit test?

A unit test is one kind of eval. The difference is emphasis: a unit test is usually written to exercise a function, while an eval is written to prove a requirement from the specification, which means it can be a unit test, an integration test, a fixture comparison, a security check, or a model-graded rubric. In practice the useful shift is writing checks that describe what the product must do rather than how the current code is arranged.

Why does AI-written code need evals more than hand-written code?

Because it arrives faster than a person can read it, and because its failures are plausible rather than obvious. Veracode's spring 2026 testing found only 55 percent of AI code generations were secure while syntax correctness passed 95 percent, and the most common developer complaint in Stack Overflow's 2025 survey was output that is almost right but not quite. Automated checks catch that class of error reliably; a fast human read does not.

Does eval-driven development slow a team down?

It moves the effort rather than adding it. You spend more time before the implementation defining what correct means, and much less time afterwards debugging, reverting, and re-explaining. The 2025 DORA research found AI adoption raises throughput and lowers delivery stability at the same time, which is what a team paying that cost later instead of earlier looks like.

Can the AI write its own evals?

It can draft them, and it should not be the only thing that judges them. If the same agent writes the implementation and the check in the same pass, you get a system that agrees with itself. Anthropic's Claude Code guidance recommends a fresh context or a second model for review precisely so the thing doing the work is not the thing grading it, and the safest version is to derive the check from the specification before the implementation exists.

How many evals does a project need to start?

Fewer than most teams expect. Start with the one flow where a failure that nobody notices would be most expensive, write down what must always be true about it, and turn each of those sentences into a check. That is usually a day or two of work, and it protects more value than a month spent trying to raise a coverage percentage.

What should we measure to know if the evals are working?

Watch the problems that come back after release, rather than the amount of code released: defects that reach users, revert and rollback rate, and how long it takes to restore service after a bad change. Code volume stopped being a useful measure once code became cheap to generate, and a rising eval pass rate on a suite nobody trusts tells you nothing.

Do evals replace code review?

No, they change what review is for. Machine checks handle whether the code behaves correctly, which humans are slow and unreliable at, and that frees the reviewer to judge whether the change was the right idea, whether the design will still be good in a few years, and whether the eval itself is missing a case. Both halves are needed, and the review gets better when people are not doing the checking work a machine does better.

References