What is eval-driven development?
Eval-driven development is a way of building software where you write the automated check first, encode a requirement from the specification in it, and then let an AI coding agent write and rewrite the implementation until the check passes. The check is the deliverable your team owns and reviews. The code is what satisfies it. That reversed order matters more now than it did, because the code is no longer the expensive part.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- An eval is an automated check that decides, with no human judgment, whether software meets a stated requirement.
- Eval-driven means the check exists before the implementation and stays as a permanent check afterwards.
- It is close to test-driven development, with two differences: the checks come from the spec rather than from the code's structure, and the loop is run by an agent that can iterate in seconds.
- The eval is also the feedback signal the agent needs. Without a check it can run, an agent stops when the work looks done.
The word eval comes from AI research, where nobody releases a model without measuring it against a fixed set of cases with known answers. An eval set is that collection of cases, and the eval score is the pass rate. Apply the same idea to software instead of to a model and you get eval-driven development: a set of automated checks that encode what the product must do, run on every change, and decide pass or fail without anyone's opinion.
The definition, precisely
An eval has three properties. It is automated, so it runs without a person. It is deterministic in its verdict, so two engineers reading the result agree on what happened. And it encodes a requirement rather than an implementation, so it keeps working after a refactor.
Eval-driven development adds an ordering rule: the check comes first. You take a line from the specification, turn it into a check, watch it fail, and only then let anything write the code. The finished work is a diff plus a passing check that did not exist before. A reviewer can see both.
That ordering is the whole practice. Everything else in this guide is detail about how to do it well.
How it differs from test-driven development
Test-driven development, in its classic form, is red, green, refactor: write a failing unit test, make it pass, clean up. Eval-driven development is the same loop with two changes, and both come from who now writes the code.
The first change is what the check is about. Classic TDD developed around unit tests, which tend to follow the structure of the code: one test class per implementation class, one test per method. That works when a human is designing both together. It works badly when a machine is generating implementations, because the machine will freely rearrange the structure and the tests will all break for reasons that have nothing to do with behaviour. An eval asserts what a caller or a user must be able to rely on. It should not care how many classes exist behind it.
The second change is the speed of the loop. In TDD a human runs the loop, so the number of iterations is limited by patience. With an agent, the loop runs on its own: write, run the check, read the failure, try again, and repeat until the check passes. That changes the cost of writing a hard check. A check that would take a person an afternoon to satisfy is now worth writing, because the person is not the one satisfying it.
If your team already practises TDD, you are most of the way there. What changes is that the checks become the instructions you give to the machine, so they need to be readable as a statement of intent rather than as supporting code.
Why the check is also the instruction
There is a second reason to write the eval first, and it is the one teams underrate. The check is how the agent knows it is wrong.
Anthropic's engineering guidance for Claude Code puts it plainly: an agent stops when the work looks done, and without a check it can run, "looks done" is the only signal available, which means you have to do all the checking yourself. Give it something that produces a pass or a fail and the agent can keep trying until it passes [1]. The same document names the standard failure of skipping this, the trust-then-verify gap, where the model produces a plausible-looking implementation that does not handle unusual cases, and gives a one-line fix: always provide verification, and if you cannot verify it, do not release it.
So an eval suite does two jobs at once. It protects your users from a bad change, and it makes the AI meaningfully better at writing the code in the first place, because the AI gets to find out it was wrong before you do. A repository with real checks is a repository where an agent can work unattended. A repository without them is one where every mistake waits for a human to notice.
What counts as an eval
Broader than most people assume. Anything that yields an automatic verdict against a requirement qualifies:
A unit or integration test, which is the most common case. A fixture comparison, where you run a function over a saved input and diff the output against a saved expected result. A schema validation on an API response. A migration check that runs the migration against a copy of data with the same structure as production and asserts the invariants still hold. A performance assertion with a hard threshold. A security check for a specific class of flaw, which matters more than people think given that Veracode's spring 2026 testing found only 55 percent of AI code generations were secure while syntax correctness ran above 95 percent [2]. And, for the cases with no single right answer, a model-graded rubric, which is its own topic in LLM as judge for code review.
Anthropic's guidance on eval design is worth using directly here: be task-specific and include the unusual cases, structure the check so grading can be automated, and prefer many cases with less precise automated scoring over a few carefully hand-graded ones [3]. That last one is surprising and correct. A large number of automated results is worth more than a small amount of careful human attention, because the automated version runs again tomorrow.
What it is not
It is not a coverage target. Coverage measures which lines ran, whether or not any behaviour was proven, and a team aiming for a coverage number will produce tests that execute code without asserting anything useful about it.
It is not a replacement for review. It removes the part of review that humans are bad at, which is mentally simulating a computer, and leaves the part they are good at. We lay out that split in evals vs tests vs code review.
It is not the same as evaluating an AI product. If your product itself uses a model and you need to know whether its answers are good, that is a related discipline with different tools, covered in AI evaluation and guardrails for production. This guide is about proving that a code change does what the spec said, whoever or whatever wrote it.
Where to go next
If you want the argument for why this is now necessary rather than only neat, read the verification gap in AI coding. If you are convinced and want to start, go to how to write your first eval suite. If you are working out how this fits a codebase that already exists and has no checks worth trusting, start with adding evals to an existing codebase.
Common questions
What is an eval in software development?
An eval is an automated check that decides whether software meets a stated requirement, with a pass or fail verdict that needs no human interpretation. It can be a unit test, an integration test, a fixture comparison, a schema or security check, or a model-graded rubric, as long as it asserts a requirement rather than the current structure of the code.
Is eval-driven development the same as test-driven development?
It is TDD adapted for teams where a machine writes the implementation. The two differences are that the checks are derived from the specification rather than from the structure of the code, and that the write-run-fix loop is executed by an agent in seconds rather than by a person over an afternoon, which makes it worth writing harder checks.
Why write the eval before the code?
Two reasons. It forces you to state what correct means while you can still discover that the specification was vague, and it gives the coding agent a signal it can act on, since an agent with no check to run stops as soon as the work looks finished.
Do evals only apply to AI-written code?
No, the practice is good for any code, and it was good practice long before coding agents existed. What changed is the cost of skipping it: code now arrives far faster than a person can read it, so the manual review that used to catch the risk cannot keep up.
What are the three properties an eval must have?
An eval must be automated, so it runs without a person present. It must be deterministic in its verdict, so two engineers reading the result reach the same conclusion. And it must encode a requirement rather than the current structure of the implementation, so it keeps working after a refactor. A check missing any of the three is not a reliable eval, whatever it is called.
How do I turn a specification into an eval?
Take a line from the specification, write a check that fails against the current code, then let the implementation change until that check passes. The finished unit of work is a diff plus a passing check that did not exist before, which gives a reviewer both the change and the proof it does what was asked, in one artefact.
Does eval-driven development slow down the first version of a feature?
It adds a short step before implementation and removes a longer one afterwards. Writing the check first costs time up front, but it replaces the debugging, reverting, and re-explaining that follows a change nobody proved correct, and with an agent running the loop, the extra step usually saves more time than it costs within the same task.
What happens to the eval suite after the feature is released?
It stays in the pipeline as a permanent check rather than being archived once the feature is done. Every future change to that part of the codebase has to keep the check passing, which is what separates eval-driven development from a one-time test written to satisfy a code review and then forgotten.
References
- [1] Anthropic, Best practices for Claude Code: “Give Claude a way to verify its work”, and “Always provide verification (tests, scripts, screenshots). If you can't verify it, don't ship it.”
- [2] Veracode, Spring 2026 GenAI Code Security update: “only 55% of generation tasks result in secure code”, against syntax correctness above 95%.
- [3] Anthropic, Create strong empirical evaluations: task-specific with edge cases, automated grading, and “more questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals”.
Related reading
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
What good code review looks like when nobody wrote the code
With human code, the author is the first check and review is the second. With generated code, review is the only check. That one change alters most of what a reviewer should be doing.
More in Start here
The verification gap: why AI made writing code cheap and checking it expensive
Generating code got roughly one hundred times cheaper in three years. Reading it did not get cheaper at all, because a person still reads at the speed a person reads. That mismatch is the verification gap, and it explains why teams adopting AI tools often feel much faster while producing more work that has to be redone. The research on this is now strong enough to settle the question.
Evals vs tests vs code review: which one catches what
These three get treated as interchangeable quality activities and they are not. A unit test proves a function behaves. An eval proves a requirement holds. A human review judges whether the change was a good idea. Only the third can tell you the feature was pointless, and only the first two will still be checking next year when everyone who wrote it has left.