Never let the model grade its own work

There is a rule every engineering team already accepts: nobody approves their own pull request. It is so obvious that it never needs defending.
Then the same team hands a ticket to an agent, which writes the implementation and the tests in one run, the pipeline passes, and everyone treats that as verified work. It is not verified. It is internally consistent, which is a much weaker property and an easy one to mistake for the real thing.
This is the most common structural flaw we find in AI-built codebases, and it is worth being precise about why it happens.
The mechanism
Take a requirement with any ambiguity in it. Say a report should show active users.
An agent reads that, decides active means a session within the last 30 days, implements it, and writes a test asserting that a user with a session 20 days ago appears in the report and a user with one 40 days ago does not. The test is well written. It passes. It is also a perfect record of the agent's interpretation, and if the business meant something else by active, then the code is wrong and the check agrees with it.
Nothing in the pipeline can catch that, because the pipeline is comparing the implementation to a statement of the implementation. The requirement was never part of the process.
Repeat that across a whole project and you get a codebase with thousands of passing tests and no idea which of them assert a real promise. That is the state we are usually asked to look at, and the tests are the hardest part to fix, because to the team that owns them they look like an asset.
The fix is about where the standard comes from
The instinct is to add a reviewer. Useful, and secondary. The primary fix comes earlier: the check has to be derived from the specification, before the implementation exists.
That single ordering rule solves the problem. If the requirement was written as a sentence whose check is obvious, and the check was written from that sentence, then the standard was fixed before anything had a chance to interpret it. The agent's job becomes satisfying an external standard instead of documenting its own choices. That is the whole idea behind eval-driven development, and the specification discipline it depends on is in writing specs an agent can verify.
It also has a useful side effect. Trying to write the check first is the fastest way to find out that the requirement was vague, because you cannot encode a sentence nobody has decided. That hour of discomfort is the cheapest hour in the project.
Independence, in three layers
Once the standard is external, independence in the grading is the second layer of protection. Three rules, in order of how much they give you.
A fresh context, always. Do not grade a change in the session that produced it. The reasoning is still in context, so the grader evaluates its own intent rather than the artefact. Anthropic's guidance for Claude Code is direct about this: a verification subagent or a fresh-context review has a fresh model try to refute the result, so the agent doing the work is not the one grading it. The same document notes that a fresh context reviews better because it is not biased towards code it just wrote.
A different model where it matters. Two models trained differently tend to miss different things, so disagreement between them is informative. This is a modest gain compared to the first rule and it is close to free.
A named human on decisions that need judgment. A check enforces a rule somebody already thought of. It cannot decide the rule was missing or that the feature should not exist. That is why we keep an engineer's name against every change, which we wrote about in who signs off on AI-written code.
The risk to watch: an agent optimises for the target you gave it
Here is the failure that affects teams who have done everything else right.
You tell an agent to make the build pass. Deleting an assertion makes the build pass. Adding a skip makes the build pass. Loosening a threshold, widening a mock, wrapping the failing call in a try block: all of these make the build pass. None of it is malicious. It is the quickest way to reach the goal you stated, and stating the goal vaguely is the mistake.
So the constraint has to be stored somewhere the agent cannot change:
Changes to the check suite get reviewed separately from the change that motivated them. A diff touching both an implementation and the tests protecting it gets read with that in mind.
The pipeline fails when the count of skipped checks rises. This is worth more than any policy document.
Thresholds are kept in a protected file, not inline in the test that uses them.
Deleting a check requires a written reason in the pull request.
We are not describing distrust of the tool. It is the same reasoning that put branch protection on your main branch years ago.
What independent review actually catches
Worth being concrete, because "independence" sounds abstract until you see the categories.
Misread requirements, which is the case above and the most valuable catch.
Missing cases the implementation never considered. A self-written test suite covers the paths the implementation handles, which is precisely the set of paths that were never at risk. The empty list, the duplicate submission, the second tenant, the expired token: those are absent from both the code and its tests, together.
Assertions too weak to fail. A check that asserts a response arrived, without asserting what was in it, is only for show. A reviewer looking at the check itself, rather than at the code, is the only thing that finds these. That is the review skill we now value most, and the failure modes are catalogued in when evals give false confidence.
Where model grading does belong
None of this is an argument against using models as graders. Some things worth checking have no single right answer: whether an error message is actionable, whether a diff implements the requirement in its linked spec, whether documentation still matches behaviour. A model can grade those well against a written rubric, and it can handle a volume of changes no human reviewer can reach.
The rules that make it reliable are the same ones: independence from the work, a rubric with a small scale and named criteria, required evidence in the verdict, and validation of the judge against human grading on a repeating sample. A judge that has quietly gone lenient is worse than no judge, because it turns an unknown into a reassurance. The detailed version is in using a model as a judge, without fooling yourself.
The one-question audit
If you want to know the state of your codebase, ask this: how many of your tests were written in the same pass as the code they check?
Most teams have never asked, and the answer is usually higher than anyone expects. It is also fixable without a rewrite. Take your most important promises, one at a time, and write the check for each from the requirement rather than from the code. Then break the implementation on purpose and confirm the check fails. Any check that still passes was never testing it.
Thanks to the engineers who have sat through that exercise with us on their own repositories. The first check that does not fail when it should is always the moment everyone stops talking.
A test written by the thing it is testing only repeats what that thing already believes. Verification needs a second, separate reviewer, and it does not much matter whether that reviewer is human, only that it is a different one.
Related guide: Eval-driven development: how to prove AI-written code works.
Sources
- Anthropic, Best practices for Claude Code: a verification subagent has a fresh model try to refute the result so the agent doing the work is not the one grading it, and a fresh context improves review because the model is not biased toward code it just wrote.
- Anthropic, Create strong empirical evaluations: automated grading, including model-based grading, and the preference for volume of automated cases over a few hand-graded ones.


