Engineering

Adding eval tests was the best decision we made

Editorial · Reveneau · August 20, 2026

Adding eval tests was the best decision we made

We generate one hundred percent of the code we release. That is the company's whole proposition, and it means we are exactly the sort of team that could produce an enormous volume of confident, plausible, wrong software. The single change that stops that from happening is not a review policy and not a hiring standard. It is that we write the check before the code.

That decision has done more for the quality of our work than anything else we have tried, and the reasons are not the ones we expected at the start. Three things changed: what we actually hand over, what code review is for, and how much attention a specification deserves.

1. The check became the deliverable

The old order of a task was: read the ticket, write the code, add tests, open a pull request. Tests were the last step and everyone knew it, which is why they were usually written to agree with the code that already existed.

Now the first artefact is the check. Take the requirement, write the automated assertion that proves it, watch it fail against the current codebase, and only then let anything implement it. What we hand a reviewer is a diff plus a check that did not exist before and now passes.

The subtle part is what that does to ownership. The implementation is disposable. An agent can rewrite it five times in an afternoon, and often does, and nobody misses the version that got replaced. The check is not disposable. It is the durable statement of what the software promises, written by a person, reviewed by a person, and still running in three years when everyone involved has moved on.

Once you see the code as the cheap part and the check as the asset, a lot of decisions get easier. You stop arguing about test coverage as a percentage and start asking which promises are protected. We wrote up the full method in the guide on eval-driven development, and the practical starting version is in how to write your first eval suite.

2. Review stopped being verification

This is the change with the biggest effect on a normal working week.

A reviewer reading an unverified diff is doing two jobs at once. First, does this work? That means mentally executing the code against the cases that matter, which humans are slow and unreliable at, and which gets worse as the diff grows. Second, was this the right change? That is design, scope, naming, security, and whether the feature should exist. Only the second job needs a person.

When every diff arrives already passing real checks, the first job disappears. The reviewer opens a change knowing it behaves correctly and spends their entire attention on judgment. The review is faster and the comments are better. They are about the design of the change rather than about whether an unusual case was handled.

The research on why humans should not be the ones doing the verifying here is uncomfortable and worth reading. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. Working with AI tools, they took 19 percent longer, and afterwards estimated that AI had made them 20 percent faster. METR labels the result historical, since the tools have changed since then, and that caveat is fair. The durable finding is not the direction of the number. It is that skilled people were wrong about a large change in their own performance, in the direction that made the tool look better. If professional intuition cannot measure its own speed, it should not be the final check on correctness either.

We still keep a named engineer against every change, and we always will. What changed is that the engineer's name means something different now: not "I read this and it looked fine" but "I decided this was the right change, and here is the check that proves it does what we said".

3. The spec stopped being paperwork

The third change surprised us. When a machine implements what you wrote, the specification becomes an input rather than a document, and every vague sentence in it converts directly into confident wrong code, quickly, in volume.

A junior engineer handed a vague ticket comes back and asks a question. That was a feedback mechanism nobody designed and everybody depended on. An agent does not come back. It picks a plausible interpretation and commits to it, and you find out when you read the diff.

So we write specs differently. Each requirement is a sentence whose check is obvious: not "handle large exports properly" but "an export of up to 100,000 rows completes inside the 30 second timeout and returns a file whose row count matches the filtered query". That version is longer and it contains four decisions somebody was going to have to make anyway. Making them up front costs twenty minutes. Making them accidentally, inside an implementation, costs a release.

Anthropic's own engineering guidance for Claude Code reaches the same conclusion, and the line we quote most is the blunt one about the trust-then-verify gap: always provide verification, and if you cannot verify it, do not release it. The same document points out that an agent stops when the work looks done, so without a check it can run, looking done is the only signal available and you become the person who checks every result. Give it a check and the agent can check its own work. That matters beyond testing: it is the difference between a session you have to watch the whole time and one you can leave to run.

What it costs, honestly

Two things.

It costs time before the work rather than after it. Writing the invariants for a critical flow as plain sentences takes an hour or two, and it is genuinely harder than writing code. It is also where you discover that two people on the team believe different things about how permissions work, which is the sort of discovery that used to arrive as an incident.

And it costs discipline in one specific place: an agent optimises for the signal you give it, so if the signal is a green pipeline, then weakening a check is a valid strategy for reaching it. Nothing malicious, just the quickest way to reach the goal you stated. So changes to the check suite get reviewed separately from the change that motivated them, and a rise in skipped tests fails the build. It is the same reason nobody approves their own pull request.

What we will not claim

We are not going to tell you that this cut our defect rate by some percentage. We have not run the controlled experiment that would let us say it, and a company whose pitch is accountability should not publish a number it cannot show the calculation for. Any figure we invented would be exactly the kind of confident, unverifiable statement that this whole practice exists to prevent.

What we can say is what changed in kind. The problems we find now are design disagreements, raised in a pull request, argued about between two engineers. They are not behaviour surprises reported by a customer on a Tuesday morning. That is a qualitative claim, and it is true.

The outside numbers are worth more than ours anyway. Google's DORA program surveyed nearly 5,000 technology professionals in 2025 and found AI adoption has a positive relationship with delivery throughput and a negative relationship with delivery stability. More changes, arriving faster, breaking more often. Veracode's spring 2026 testing of more than 150 models across 80 tasks found only 55 percent of generations produced secure code, a figure flat for two years while syntax correctness climbed past 95 percent. Models got much better at writing code that compiles and no better at writing code that is safe. And in Stack Overflow's 2025 survey, the top frustration with AI tools, at 66 percent, was output that is almost right but not quite.

Almost right is the whole problem. Almost right passes a quick read, works in a demo, and is the exact failure a tired reviewer misses and a machine catches every time.

The version to copy

If you take one thing from this, take the steps rather than the ideas. Pick the flow in your product where a silent failure would be most expensive. Write down what must always be true about it, in sentences. Turn each sentence into a check. Add those checks to your pipeline so a failure blocks the merge. That is a day or two of work and it is a real eval suite for the part of your system that matters most.

Then let real failures decide the next check. When something reaches a customer that should not have, write the check that would have caught it before you write the fix. A suite grown that way stays small, stays trusted, and ends up covering exactly your product's real weaknesses.

Thanks to the teams who let us look at their passing pipelines and their incident histories side by side. The difference between those two records is where we learned the most.

Generating code is cheap now. Knowing it is correct never got cheaper, and that is the part worth building.

Related guide: Eval-driven development: how to prove AI-written code works.

Sources

Common questions

What is an eval test?

An eval test is an automated check that proves a stated requirement holds, with a pass or fail result that needs no human interpretation. It can be a unit test, an integration test, a fixture comparison, a schema assertion, or a model-graded rubric, and what makes it an eval is that it encodes a promise from the specification rather than the current shape of the code.

Why write the check before the code?

Because it forces you to state what correct means while it is still cheap to discover that nobody had decided, and because it gives a coding agent something to work against. An agent with no check to run stops as soon as the work looks finished, which means every mistake waits for a person to notice it.

Does writing evals first slow a team down?

Writing evals first moves the effort earlier rather than adding it. You spend more time defining correct behaviour before implementation and much less time debugging, reverting, and re-explaining afterwards, which is why teams that skip this step often feel fast while releasing more work that has to be redone. Deciding what correct means up front costs an hour or two, but discovering it by accident, inside a released feature, costs a release and often an incident.

How is this different from test-driven development?

It is TDD adapted to a machine doing the typing. The checks come from the specification rather than from the structure of the code, so they survive refactors, and the cycle of writing, running and fixing is done by an agent in seconds rather than by a person over an afternoon, which makes harder checks worth writing.

Can the AI write its own eval tests?

It can draft them, and it should not be the only thing judging them. If one pass produces both the implementation and the check, the check encodes that implementation's interpretation of the requirement, including any misunderstanding, so it passes without proving anything about the requirement itself.

What should the first eval tests cover?

The flow where a silent failure would cost the most, which is usually authentication and permissions, anything that moves money, or anything that writes or deletes customer data. Write down what must always be true about that flow as plain sentences, then turn each sentence into a check that blocks a merge when it fails.

How many eval tests does a project need to start?

A project needs fewer eval tests than most teams expect. Five to twenty checks on the things that would cause real harm, all fast, none flaky, each one having failed at least once in real conditions, is a stronger position than thousands of tests that have never failed. Start with the flow where a silent failure would cost the most, write down what must always be true about it, and add to the suite from there as real work shows new cases.

Do eval tests replace code review?

Eval tests do not replace code review, they change what review is for. Machine checks take over the questions with right answers, such as behaviour, unusual cases, and contracts, which frees the reviewer to judge design, scope, naming, and whether the change was worth making at all. A reviewer reading a diff that already passes real checks stops asking whether it works and starts asking whether it was the right change to make.

What happens when an agent tries to delete a failing test?

You need enforcement rather than instruction, because an agent optimises for the signal you gave it and a pipeline where every check passes is that signal. Review changes to the test suite separately from the change that motivated them, fail the build when the count of skipped tests rises, and keep thresholds in protected files.

Is a passing build proof that a change is safe?

Only if the checks can actually fail. Assertions that mock away the real code path, compare a value to itself, or silently skip are extremely common, so the useful habit is to break the code a check protects on purpose and confirm the build fails.

How do you know the eval suite is working?

Measure the results of your changes after release: change failure rate, how long recovery takes, and the share of defects your own checks catch before a customer does. A rising count of tests tells you nothing on its own. The better habit is to break the code a check protects on purpose now and then and confirm the build actually fails, since a check that cannot fail proves nothing.

Where should a team start if their codebase has no tests worth trusting?

Read the incident history to find where failure actually happens, cover the three highest-cost flows with blocking checks over two or three weeks, then stop and release. After that, add one check for existing behaviour each time you are about to change an area, so the suite grows along with real work instead of becoming a project of its own.