Evals vs tests vs code review: which one catches what
These three get treated as interchangeable quality activities and they are not. A unit test proves a function behaves. An eval proves a requirement holds. A human review judges whether the change was a good idea. Only the third can tell you the feature was pointless, and only the first two will still be checking next year when everyone who wrote it has left.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Unit tests check units of code. Evals check requirements. Review checks judgment. The three differ in what they check, and each needs the same care.
- Automated checks are the only part of quality that scales with code volume, because they read as fast as the machine runs.
- Human review is where design, naming, security choices, and whether the thing was worth building get decided.
- The most common mistake is having reviewers do the machine's job, which produces slow review and missed unusual cases at the same time.
Ask three engineers what the difference is between a test and an eval and you will get three answers, which usually means the words mean less than people think. Here is the distinction we use, and more importantly, the division of labour it implies.
Three different questions
A unit test asks: does this piece of code do what its author intended? It is written close to the implementation, runs in milliseconds, and is excellent at catching a regression in a function you refactored. It is also the check most likely to break for no good reason when the structure changes, because it is tied to the structure.
An eval asks: does the system satisfy this requirement? It is written from the specification rather than from the code. It can be implemented as a unit test, an integration test, a fixture diff, a schema assertion, a performance threshold, or a model-graded rubric. What makes it an eval is that if it fails, a stated promise about the product is broken, and if it passes, that promise holds regardless of how the code is arranged inside.
A code review asks: was this a good change? Is the design right, will it be maintainable, does the naming say what it means, is there a simpler version, does this belong in this service, is there a security consideration nobody wrote down, and, most valuable of all, should we be building this at all. None of that has a pass-fail verdict, which is exactly why it needs a person.
What each one is good at
Automated checks have four properties a person cannot match. They do not get tired, so the four-hundredth case gets the same attention as the first. They are exact, so they catch an off-by-one that a reader's brain corrects without noticing. They are repeatable, so a check written today runs unchanged in three years against code written by people who never met you. And they scale with volume, which is the property that matters most now that a change can be 3,000 lines produced in ten minutes.
Human review has properties no check can match. It can tell you the requirement itself was wrong. It can see that a change matches the exact words of the spec and goes against what the product is meant to do. It can notice that two systems are becoming more different over time, or that a beginner's coding pattern is spreading, or that the eval that just passed is asserting something too weak to matter. It includes accountability, which software cannot: a person can be asked why they approved this, and that question changes behaviour.
Where each one fails
Unit tests fail by testing the implementation. A suite that mirrors the class structure will break on every refactor, and a team that accepts that learns to treat failures as normal. That is the most expensive habit a codebase can acquire.
Evals fail by being too weak or too flaky. A check that asserts a response arrived, without asserting what was in it, only looks like checking. A check that fails one run in twenty teaches everyone to run it again. Both produce a passing result that means nothing, which we cover in when evals give false confidence.
Review fails by volume and by fatigue. Attention per line falls as the diff grows, and it falls fast. A reviewer given a large generated diff and no automated check has been handed an impossible job, and the honest outcome is an approval that means "this looks like something a competent person would write" rather than "I have verified this".
The division of labour we use
The rule is simple: anything with a right answer belongs to a machine, and everything else belongs to a person.
So behaviour, contracts, unusual cases, error paths, permission boundaries, data invariants, and performance thresholds go in the eval suite. They have right answers. Write them down once and let them run forever.
Design, scope, naming, dependency choices, whether the spec made sense, whether the eval is strong enough, and whether the change is worth its future maintenance cost go to the reviewer. Those need judgment, and a reviewer who is not busy simulating a computer has the attention to give them.
The order matters too. The automated check runs first and the human reads a diff that already passes. This also changes what the reviewer is looking at: working code, so the question becomes "is this right?" rather than "does this work?". Those are different reviews and the second one produces better software.
For the human half in detail, including what to look for in a generated diff, see how to review AI-generated code and our post on what good code review looks like.
A note on who is accountable
There is one thing neither an eval nor a linter can do: a check can only enforce a rule someone already thought of. When something reaches a customer that should not have, the useful question is who is responsible for making sure the missing test gets written. That responsibility has to belong to the team that released it. On our own builds it belongs to us, for the whole project and after release, rather than being handed off at the pull request. We wrote about that in who signs off on AI-written code.
Getting the mix right
If your review queue is slow and problems still reach production, you almost certainly have reviewers doing machine work. Move the mechanical checks into the suite and the queue speeds up while quality rises, which may sound unlikely and is simply what happens when each job moves to the party suited to it.
If your suite passes and your customers keep finding bugs, your checks are asserting the wrong things. Start from real incidents and write the eval that would have caught each one, as described in adding evals to an existing codebase.
A worked example of the split
Take a feature that lets a support agent refund a customer up to a stated limit without manager approval. Three kinds of work happen on that change, and each belongs to a different check.
The unit test covers the refund calculation itself: given an order total and a partial refund amount, does the function return the right remaining balance. It is written close to the code, runs in milliseconds, and would need rewriting if the function were split into two functions tomorrow, because it is pinned to today's structure.
The eval covers the requirement: a support agent cannot refund more than the stated limit without a second approval, regardless of how the refund is calculated internally. It is written from the policy, not from the function signature, so it survives a rewrite of the calculation logic, a change of programming language for that service, or a move from a monolith to a separate refunds service. The eval still asks the same question: was the limit enforced.
The review covers whether the limit itself is the right policy, whether logging the refund reason should be mandatory, whether a support agent should be able to see the customer's full order history while doing this, and whether this feature creates a new way for an employee to move money out of the business that nobody previously had to think about. None of those questions has a pass or fail answer, and none of them will be caught by any test, however well written.
What happens when the split is missing
Skip the eval and keep only the unit test and the review, and the limit enforcement lives entirely in a reviewer's attention on the day the pull request was opened. Six months later, someone refactors the refund service, the unit tests get rewritten to match the new structure because they were tied to it, and the limit check is quietly gone from the new version. Nothing failed. Nothing turned red. The requirement simply stopped being tested, because nothing was ever checking the requirement in the first place, only the implementation of the day it was written.
Skip the review and keep only automated checks, and a different failure appears. The eval suite might prove that the limit is enforced exactly as specified, and still say nothing about whether letting support agents issue refunds without approval was a reasonable policy for a company processing the volume this one does. A machine can prove a rule was followed. It cannot tell you the rule was a bad idea.
Where teams misjudge the boundary
The two most common mistakes go in opposite directions and are worth naming separately, because each looks like diligence while it is happening.
The first is writing every unit test as though it were an eval, meaning as though it must never need to change. Some unit tests exist precisely to pin down an implementation detail while it is being developed, and it is fine for those to break on a refactor. Trying to make all of them permanent produces a suite that resists every improvement to the code's internal structure, which teaches engineers to avoid refactoring rather than to trust their checks.
The second is treating a design discussion in a pull request comment as though it were an eval, meaning writing the concern down once and trusting everyone to remember it. A comment that says "make sure this never exceeds the limit" is not a check. It is a hope, aimed at a future reader who may never see it, on a line of code that may move. If a concern matters enough to write down, it is worth the extra step of turning it into an assertion the pipeline runs on every future change, rather than a note that ages out of relevance the moment the thread closes.
How the balance shifts as a team adopts AI tools more heavily
The proportions in this split are not fixed. As more of a codebase is generated rather than typed by hand, the argument for moving anything with a right answer into the eval suite gets stronger, because a coding agent will happily iterate against a check thousands of times in an afternoon and will not iterate meaningfully against a comment. The instruction an agent can act on is a failing assertion it can read and try to satisfy. A note asking for care is not something an agent, or a tired reviewer, reliably acts on. This is one more reason the eval suite is worth investing in ahead of review capacity, covered from the buyer's side in the business case for evals.
Best for
- Evals: anything with a right answer that must stay true on every future change
- Unit tests: fast feedback on a single piece of logic while it is being written
- Human review: design, scope, naming, security judgment, and whether to build it at all
Avoid if
- Do not use review as the main protection against behaviour bugs in large generated diffs
- Do not write evals that assert the current code structure, since they break on every refactor
- Do not keep a flaky check in the blocking suite: one unreliable test devalues all of them
Check before you decide
- Confirm every merge requires an automated check to pass, as well as an approval
- Confirm your reviewers see diffs that already pass, so their attention goes to judgment
- Confirm each defect that reaches users produces a new eval, as well as a fix
Common questions
What is the difference between an eval and a unit test?
A unit test checks that a piece of code does what its author intended, and is usually written against the code's structure. An eval checks that the system satisfies a requirement from the specification, so it can be implemented as a unit test or an integration test or a fixture comparison, and it stays valid when the internal structure changes.
Do evals replace human code review?
No. They take over the part of review that has right answers, such as behaviour, unusual cases, and contracts, and leave the part that needs judgment, such as design, scope, naming, and whether the change was worth making. Review usually gets both faster and more valuable once people are no longer doing the checking work that a computer does better.
Why does human review miss bugs in AI-generated code?
Because attention per line drops as a diff grows, and generated code is uniformly tidy, which reads as careful even when it is wrong in one specific place. The characteristic AI failure is code that is almost right, which is the hardest thing for a reader to notice and one of the easiest for an automated check to catch.
Should the same person write the eval and the code?
The eval should come from the specification and exist before the implementation, which in practice means it is written or reviewed independently of the code that satisfies it. If one pass produces both, you get a check that agrees with the implementation rather than one that tests the requirement.
What question does a code review answer that an eval cannot?
Whether the change was a good idea. A review can judge design, naming, maintainability, whether a case is missing from the check itself, and whether the feature should have been built at all, none of which has a pass or fail verdict. An eval only confirms that a stated requirement holds, so it cannot tell you the requirement was wrong.
Why do unit tests break more often than evals during a refactor?
Because a unit test is usually written close to the implementation and follows the structure of the code, so it tends to break whenever that structure changes even if the behaviour is unchanged. An eval is written from the specification, so it keeps passing through a refactor as long as the underlying requirement still holds.
How do you split work between an eval suite and a reviewer in practice?
Anything with a right answer, such as behaviour, contracts, unusual cases, permission boundaries, and data invariants, goes into the eval suite and runs automatically. Anything needing judgment, such as design, scope, naming, and whether the eval itself is strong enough, goes to a reviewer, and the automated check runs first so the reviewer reads a diff that already works.
Is a passing eval suite the same as an approved code review?
No. A passing suite proves the stated requirements hold, while an approved review confirms a person judged the change worth making and well designed. Both are needed, and treating a passing suite as sufficient approval removes the only check on whether the requirement itself was the right one.
Related reading
What good code review looks like when nobody wrote the code
With human code, the author is the first check and review is the second. With generated code, review is the only check. That one change alters most of what a reviewer should be doing.
Who signs off on AI-written code?
A test can only enforce a rule somebody already thought of. When AI writes most of the code, the question that decides whether a codebase stays trustworthy is who is accountable for the rule that was missing.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
More in Start here
What is eval-driven development?
Eval-driven development is a way of building software where you write the automated check first, encode a requirement from the specification in it, and then let an AI coding agent write and rewrite the implementation until the check passes. The check is the deliverable your team owns and reviews. The code is what satisfies it. That reversed order matters more now than it did, because the code is no longer the expensive part.
The verification gap: why AI made writing code cheap and checking it expensive
Generating code got roughly one hundred times cheaper in three years. Reading it did not get cheaper at all, because a person still reads at the speed a person reads. That mismatch is the verification gap, and it explains why teams adopting AI tools often feel much faster while producing more work that has to be redone. The research on this is now strong enough to settle the question.