Eval-driven development: how to prove AI-written code works / Build the eval suite
Using a model as a judge, without fooling yourself
Some things you want to check have no single correct string to compare against: the quality of an error message, whether a diff matches its spec, whether generated documentation is accurate. A model can grade those, and it is a genuinely useful eval when it is set up with two rules: the judge must be independent of the thing it grades, and the judge itself has to be checked against human judgment on a sample.
Published August 20, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Model grading is for the checks that have no exact right answer. Everything with a right answer should be a deterministic check.
- The model that wrote the code must not be the only thing grading it, and ideally not in the same context.
- A judge needs a written rubric with a small scale. Vague prompts produce graders that pass almost everything.
- Validate the judge against human review on a sample, repeatedly, because a judge that has become too lenient gives you false confidence rather than none.
- Keep model-graded checks advisory until they have proven reliable enough to block.
Deterministic checks cover more than people expect, and they do not cover everything. Whether an error message is actually helpful, whether a code comment describes what the function now does, whether a diff implements the requirement it claims: these have better and worse answers rather than one right answer. Model grading is the tool for that category, and it is easy to do in a way that produces wrong results that look trustworthy.
When to use it
Use a model judge when three things are true. The property you want to check is real and worth checking. It has no single correct output you could compare against. And a competent human reviewer could grade it consistently from a written rubric.
That third condition matters. If two of your own engineers would disagree about whether an example passes, a model will not fix that, because the problem is in the specification. Go and settle the disagreement first.
Common good uses: does this diff satisfy the requirement in the linked spec, is this error message actionable for a user, does this documentation match the current behaviour, is this migration missing a rollback path, does this change introduce a pattern that contradicts the conventions file.
The independence rule
If the same agent writes an implementation and grades it in the same run, you have built a system that agrees with itself, which gives you no real check. The reasoning that produced the change is still in context, so the grader is evaluating its own intent rather than the artefact.
Anthropic's guidance is explicit about the fix: a verification subagent, or a review in a fresh context, has a fresh model try to refute the result, so the agent doing the work is not the one grading it [1]. The same document notes a fresh context improves code review because the model is not biased towards code it just wrote.
So: separate context, always. A different model where you can. And where the grading is important, derive the rubric from the specification before the implementation exists, so the standard cannot change to fit the code. Our position on why this generalises beyond models is in never let the model grade its own work.
Writing a rubric that works
Four rules, learned through routine work.
Small scale. Pass or fail, or a three-point scale. Ask for a score out of ten and you will get sevens.
Named criteria. List what you are grading, one line each, in the order that matters. A judge asked whether something is "good" will tell you it is good.
Required evidence. Make the judge quote the specific line or requirement its verdict rests on. This does two things: it improves the verdict, and it makes a wrong verdict debuggable instead of mysterious.
One judgment per call. Do not ask a single call to assess correctness, style, security, and performance together. Separate checks, separate prompts, separate verdicts.
Anthropic's eval design guidance is useful support here as well: automate the grading, and prefer more cases with less precise scoring over a handful of hand-graded ones [2]. A judge that runs on 200 diffs a week with 90 percent agreement is worth more than a human spot-check of five.
Validating the judge
This is the step that separates a real check from one that only looks real, and it is the one that gets skipped.
Take a sample of the judge's verdicts, have a person grade the same cases independently, and compare. Do it when you build the judge and repeat it on a schedule, because three things change over time: your codebase, the model behind the API, and the population of changes being graded. A judge that has become lenient without anyone noticing is worse than no judge at all, because it turns an unknown into false reassurance.
Track two failure directions separately. False passes let bad changes through, which is the expensive one. False fails waste engineering time and, worse, teach the team to override the judge, after which it stops being a check.
Keep a small set of known-bad examples that the judge must always fail. If a change to the prompt or the model makes those pass, you have found a regression in your grader.
Where model grading fits in the suite
After the deterministic checks, as a second layer. Anything with an exact answer belongs in a normal assertion, because deterministic checks are faster, cheaper, and never argue. Model grading is for the layer above that, and it should start advisory: it comments, it does not block, until you have evidence from the validation sample that it agrees with your reviewers often enough to be trusted with a merge.
Cost and latency are real considerations too. A judge that adds several minutes to every pull request will be avoided. Run it on the changes where it is worth the cost, which usually means the ones touching the flows you care most about.
The honest limit
A model judge is a useful check and it is not a reviewer. It has nothing to lose if the product fails, no memory of the incident last quarter, and no accountability. It cannot tell you that a requirement was a bad idea, which is the most valuable thing a human review produces.
Use it to raise the minimum quality across a large volume of changes. Keep the final judgment a person's responsibility. That is the same division we draw in evals vs tests vs code review, and it is why our own process keeps the team accountable for every change however good the automated grading gets.
Best for
- Properties with no single right answer that a human could grade from a rubric
- Checking a diff against the requirement it claims to implement, in a fresh context
- Raising the minimum quality across a high volume of changes that human review cannot cover
Avoid if
- Do not use a model judge for anything a deterministic assertion can check
- Do not let the agent that wrote the code be the only grader, or grade in the same context
- Do not let a judge block merges before a validation sample shows it agrees with your reviewers
Check before you decide
- Confirm the judge is validated against human grading on a sample, on a repeating schedule
- Confirm false passes and false failures are tracked separately
- Confirm a set of known-bad examples that the judge must always fail
Common questions
What is LLM as judge?
It is using a language model to grade an output that has no single correct answer, against a written rubric, so the result becomes an automated check. It suits things like whether a diff satisfies its specification or whether an error message is actionable, and it is the wrong tool for anything a deterministic assertion can verify.
Can the same model write code and review it?
Not in the same run, and preferably not the same model. The reasoning that produced the change is still in context, so the grader evaluates its own intent rather than the artefact, which is why Anthropic's guidance recommends a fresh context or a separate subagent that tries to refute the result.
How do you know a model judge is accurate?
Sample its verdicts, have a person grade the same cases independently, and compare, then repeat on a schedule. Track false passes and false failures separately, and keep a fixed set of known-bad examples the judge must always fail so you can catch a regression in the grader itself.
Should a model judge be allowed to block a merge?
Only after it has proven reliable. Start advisory, gather agreement data against human reviewers, and promote it to a blocking check once the false-failure rate is low enough that engineers will not avoid it.
What makes a good rubric for a model judge?
A small scale such as pass or fail, or a three-point scale rather than a score out of ten, named criteria listed in the order that matters, a requirement that the judge quote the specific line or requirement its verdict rests on, and one judgment per call rather than grading correctness, style, and security together. Vague prompts asking whether something is good produce graders that rarely fail anything.
What is the difference between a false pass and a false fail from a model judge?
A false pass lets a bad change through undetected, which is the more expensive failure because it reaches production. A false fail wastes engineering time on a change that was actually fine, and repeated false fails teach the team to override the judge, at which point it stops functioning as a check at all. Both need to be tracked separately.
When is a model judge the wrong tool?
Whenever the property being checked has a single correct answer that a deterministic assertion could verify, since deterministic checks are faster, cheaper, and never argue. A model judge is also the wrong tool when two of your own engineers would disagree about whether an example passes, because that signals an unresolved specification problem rather than something a grader can settle.
How often should a model judge be re-validated?
On a repeating schedule, rather than only once at launch, because the codebase, the model behind the judge, and the population of changes being graded all change over time. A judge that has become lenient without anyone noticing produces false reassurance rather than an absence of signal, which is why a fixed set of known-bad examples the judge must always fail is worth keeping as a regression check on the grader itself.
References
- [1] Anthropic, Best practices for Claude Code: a verification subagent or fresh-context review “has a fresh model try to refute the result, so the agent doing the work isn't the one grading it”, and a fresh context improves review because the model is not biased toward code it just wrote.
- [2] Anthropic, Create strong empirical evaluations: automate grading where possible, including LLM-based grading, and prefer volume of automated cases over a few hand-graded ones.
Related reading
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
More in Build the eval suite
How to write your first eval suite
You do not need a testing strategy document to start. You need one flow where a failure nobody notices would be expensive, a short list of sentences describing what must always be true about it, and each of those sentences turned into a check that runs on every change. That is a real eval suite, it takes a day or two, and it protects more than a month of trying to raise a coverage number.
What to check in an eval suite: the seven things that matter
Coverage percentages tell you which lines of code ran during the tests, and nothing about which promises to users are protected. This is the list we work through instead: seven classes of check, ordered by how much damage they prevent, with a note on what is not worth automating. Most products need all seven eventually and only two or three of them on day one.
Writing specs an AI agent can verify
When a machine writes the implementation, the specification stops being a document people skim and becomes the actual input to the work. Vague specs used to produce slow projects. Now they produce large amounts of wrong code that looks right, fast. A spec that works has four parts: the behaviour stated as testable sentences, the boundaries named, the out-of-scope list written down, and an end-to-end check that proves the whole thing.