Adding eval tests was the best decision we made

We generate one hundred percent of the code we release. That is the company's whole proposition, and it means we are exactly the sort of team that could produce an enormous volume of confident, plausible, wrong software. The single change that stops that from happening is not a review policy and not a hiring standard. It is that we write the check before the code.
That decision has done more for the quality of our work than anything else we have tried, and the reasons are not the ones we expected at the start. Three things changed: what we actually hand over, what code review is for, and how much attention a specification deserves.
1. The check became the deliverable
The old order of a task was: read the ticket, write the code, add tests, open a pull request. Tests were the last step and everyone knew it, which is why they were usually written to agree with the code that already existed.
Now the first artefact is the check. Take the requirement, write the automated assertion that proves it, watch it fail against the current codebase, and only then let anything implement it. What we hand a reviewer is a diff plus a check that did not exist before and now passes.
The subtle part is what that does to ownership. The implementation is disposable. An agent can rewrite it five times in an afternoon, and often does, and nobody misses the version that got replaced. The check is not disposable. It is the durable statement of what the software promises, written by a person, reviewed by a person, and still running in three years when everyone involved has moved on.
Once you see the code as the cheap part and the check as the asset, a lot of decisions get easier. You stop arguing about test coverage as a percentage and start asking which promises are protected. We wrote up the full method in the guide on eval-driven development, and the practical starting version is in how to write your first eval suite.
2. Review stopped being verification
This is the change with the biggest effect on a normal working week.
A reviewer reading an unverified diff is doing two jobs at once. First, does this work? That means mentally executing the code against the cases that matter, which humans are slow and unreliable at, and which gets worse as the diff grows. Second, was this the right change? That is design, scope, naming, security, and whether the feature should exist. Only the second job needs a person.
When every diff arrives already passing real checks, the first job disappears. The reviewer opens a change knowing it behaves correctly and spends their entire attention on judgment. The review is faster and the comments are better. They are about the design of the change rather than about whether an unusual case was handled.
The research on why humans should not be the ones doing the verifying here is uncomfortable and worth reading. In a randomised controlled trial published in July 2025, METR gave 16 experienced open-source developers 246 real tasks in repositories they already knew well. Working with AI tools, they took 19 percent longer, and afterwards estimated that AI had made them 20 percent faster. METR labels the result historical, since the tools have changed since then, and that caveat is fair. The durable finding is not the direction of the number. It is that skilled people were wrong about a large change in their own performance, in the direction that made the tool look better. If professional intuition cannot measure its own speed, it should not be the final check on correctness either.
We still keep a named engineer against every change, and we always will. What changed is that the engineer's name means something different now: not "I read this and it looked fine" but "I decided this was the right change, and here is the check that proves it does what we said".
3. The spec stopped being paperwork
The third change surprised us. When a machine implements what you wrote, the specification becomes an input rather than a document, and every vague sentence in it converts directly into confident wrong code, quickly, in volume.
A junior engineer handed a vague ticket comes back and asks a question. That was a feedback mechanism nobody designed and everybody depended on. An agent does not come back. It picks a plausible interpretation and commits to it, and you find out when you read the diff.
So we write specs differently. Each requirement is a sentence whose check is obvious: not "handle large exports properly" but "an export of up to 100,000 rows completes inside the 30 second timeout and returns a file whose row count matches the filtered query". That version is longer and it contains four decisions somebody was going to have to make anyway. Making them up front costs twenty minutes. Making them accidentally, inside an implementation, costs a release.
Anthropic's own engineering guidance for Claude Code reaches the same conclusion, and the line we quote most is the blunt one about the trust-then-verify gap: always provide verification, and if you cannot verify it, do not release it. The same document points out that an agent stops when the work looks done, so without a check it can run, looking done is the only signal available and you become the person who checks every result. Give it a check and the agent can check its own work. That matters beyond testing: it is the difference between a session you have to watch the whole time and one you can leave to run.
What it costs, honestly
Two things.
It costs time before the work rather than after it. Writing the invariants for a critical flow as plain sentences takes an hour or two, and it is genuinely harder than writing code. It is also where you discover that two people on the team believe different things about how permissions work, which is the sort of discovery that used to arrive as an incident.
And it costs discipline in one specific place: an agent optimises for the signal you give it, so if the signal is a green pipeline, then weakening a check is a valid strategy for reaching it. Nothing malicious, just the quickest way to reach the goal you stated. So changes to the check suite get reviewed separately from the change that motivated them, and a rise in skipped tests fails the build. It is the same reason nobody approves their own pull request.
What we will not claim
We are not going to tell you that this cut our defect rate by some percentage. We have not run the controlled experiment that would let us say it, and a company whose pitch is accountability should not publish a number it cannot show the calculation for. Any figure we invented would be exactly the kind of confident, unverifiable statement that this whole practice exists to prevent.
What we can say is what changed in kind. The problems we find now are design disagreements, raised in a pull request, argued about between two engineers. They are not behaviour surprises reported by a customer on a Tuesday morning. That is a qualitative claim, and it is true.
The outside numbers are worth more than ours anyway. Google's DORA program surveyed nearly 5,000 technology professionals in 2025 and found AI adoption has a positive relationship with delivery throughput and a negative relationship with delivery stability. More changes, arriving faster, breaking more often. Veracode's spring 2026 testing of more than 150 models across 80 tasks found only 55 percent of generations produced secure code, a figure flat for two years while syntax correctness climbed past 95 percent. Models got much better at writing code that compiles and no better at writing code that is safe. And in Stack Overflow's 2025 survey, the top frustration with AI tools, at 66 percent, was output that is almost right but not quite.
Almost right is the whole problem. Almost right passes a quick read, works in a demo, and is the exact failure a tired reviewer misses and a machine catches every time.
The version to copy
If you take one thing from this, take the steps rather than the ideas. Pick the flow in your product where a silent failure would be most expensive. Write down what must always be true about it, in sentences. Turn each sentence into a check. Add those checks to your pipeline so a failure blocks the merge. That is a day or two of work and it is a real eval suite for the part of your system that matters most.
Then let real failures decide the next check. When something reaches a customer that should not have, write the check that would have caught it before you write the fix. A suite grown that way stays small, stays trusted, and ends up covering exactly your product's real weaknesses.
Thanks to the teams who let us look at their passing pipelines and their incident histories side by side. The difference between those two records is where we learned the most.
Generating code is cheap now. Knowing it is correct never got cheaper, and that is the part worth building.
Related guide: Eval-driven development: how to prove AI-written code works.
Sources
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (July 2025): 16 developers, 246 tasks, 19 percent longer with AI tools against a self-estimated 20 percent speed-up. METR labels the result historical.
- Google Cloud, Announcing the 2025 DORA report: nearly 5,000 respondents, 90 percent using AI at work, positive relationship with throughput and negative relationship with delivery stability.
- Veracode, Spring 2026 GenAI Code Security update: more than 150 models over 80 tasks, 55 percent of generations secure, syntax correctness above 95 percent.
- Stack Overflow 2025 Developer Survey, AI section: 84 percent using or planning to use AI tools, 46 percent distrust accuracy, 66 percent cite output that is almost right but not quite.
- Anthropic, Best practices for Claude Code: give the agent a check it can run, and the trust-then-verify gap.


