Build the eval suite

Using a model as a judge, without fooling yourself

Some things you want to check have no single correct string to compare against: the quality of an error message, whether a diff matches its spec, whether generated documentation is accurate. A model can grade those, and it is a genuinely useful eval when it is set up with two rules: the judge must be independent of the thing it grades, and the judge itself has to be checked against human judgment on a sample.

Published August 20, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Model grading is for the checks that have no exact right answer. Everything with a right answer should be a deterministic check.
  • The model that wrote the code must not be the only thing grading it, and ideally not in the same context.
  • A judge needs a written rubric with a small scale. Vague prompts produce graders that pass almost everything.
  • Validate the judge against human review on a sample, repeatedly, because a judge that has become too lenient gives you false confidence rather than none.
  • Keep model-graded checks advisory until they have proven reliable enough to block.

Deterministic checks cover more than people expect, and they do not cover everything. Whether an error message is actually helpful, whether a code comment describes what the function now does, whether a diff implements the requirement it claims: these have better and worse answers rather than one right answer. Model grading is the tool for that category, and it is easy to do in a way that produces wrong results that look trustworthy.

When to use it

Use a model judge when three things are true. The property you want to check is real and worth checking. It has no single correct output you could compare against. And a competent human reviewer could grade it consistently from a written rubric.

That third condition matters. If two of your own engineers would disagree about whether an example passes, a model will not fix that, because the problem is in the specification. Go and settle the disagreement first.

Common good uses: does this diff satisfy the requirement in the linked spec, is this error message actionable for a user, does this documentation match the current behaviour, is this migration missing a rollback path, does this change introduce a pattern that contradicts the conventions file.

The independence rule

If the same agent writes an implementation and grades it in the same run, you have built a system that agrees with itself, which gives you no real check. The reasoning that produced the change is still in context, so the grader is evaluating its own intent rather than the artefact.

Anthropic's guidance is explicit about the fix: a verification subagent, or a review in a fresh context, has a fresh model try to refute the result, so the agent doing the work is not the one grading it [1]. The same document notes a fresh context improves code review because the model is not biased towards code it just wrote.

So: separate context, always. A different model where you can. And where the grading is important, derive the rubric from the specification before the implementation exists, so the standard cannot change to fit the code. Our position on why this generalises beyond models is in never let the model grade its own work.

Writing a rubric that works

Four rules, learned through routine work.

Small scale. Pass or fail, or a three-point scale. Ask for a score out of ten and you will get sevens.

Named criteria. List what you are grading, one line each, in the order that matters. A judge asked whether something is "good" will tell you it is good.

Required evidence. Make the judge quote the specific line or requirement its verdict rests on. This does two things: it improves the verdict, and it makes a wrong verdict debuggable instead of mysterious.

One judgment per call. Do not ask a single call to assess correctness, style, security, and performance together. Separate checks, separate prompts, separate verdicts.

Anthropic's eval design guidance is useful support here as well: automate the grading, and prefer more cases with less precise scoring over a handful of hand-graded ones [2]. A judge that runs on 200 diffs a week with 90 percent agreement is worth more than a human spot-check of five.

Validating the judge

This is the step that separates a real check from one that only looks real, and it is the one that gets skipped.

Take a sample of the judge's verdicts, have a person grade the same cases independently, and compare. Do it when you build the judge and repeat it on a schedule, because three things change over time: your codebase, the model behind the API, and the population of changes being graded. A judge that has become lenient without anyone noticing is worse than no judge at all, because it turns an unknown into false reassurance.

Track two failure directions separately. False passes let bad changes through, which is the expensive one. False fails waste engineering time and, worse, teach the team to override the judge, after which it stops being a check.

Keep a small set of known-bad examples that the judge must always fail. If a change to the prompt or the model makes those pass, you have found a regression in your grader.

Where model grading fits in the suite

After the deterministic checks, as a second layer. Anything with an exact answer belongs in a normal assertion, because deterministic checks are faster, cheaper, and never argue. Model grading is for the layer above that, and it should start advisory: it comments, it does not block, until you have evidence from the validation sample that it agrees with your reviewers often enough to be trusted with a merge.

Cost and latency are real considerations too. A judge that adds several minutes to every pull request will be avoided. Run it on the changes where it is worth the cost, which usually means the ones touching the flows you care most about.

The honest limit

A model judge is a useful check and it is not a reviewer. It has nothing to lose if the product fails, no memory of the incident last quarter, and no accountability. It cannot tell you that a requirement was a bad idea, which is the most valuable thing a human review produces.

Use it to raise the minimum quality across a large volume of changes. Keep the final judgment a person's responsibility. That is the same division we draw in evals vs tests vs code review, and it is why our own process keeps the team accountable for every change however good the automated grading gets.

Best for

  • Properties with no single right answer that a human could grade from a rubric
  • Checking a diff against the requirement it claims to implement, in a fresh context
  • Raising the minimum quality across a high volume of changes that human review cannot cover

Avoid if

  • Do not use a model judge for anything a deterministic assertion can check
  • Do not let the agent that wrote the code be the only grader, or grade in the same context
  • Do not let a judge block merges before a validation sample shows it agrees with your reviewers

Check before you decide

  • Confirm the judge is validated against human grading on a sample, on a repeating schedule
  • Confirm false passes and false failures are tracked separately
  • Confirm a set of known-bad examples that the judge must always fail

Common questions

What is LLM as judge?

It is using a language model to grade an output that has no single correct answer, against a written rubric, so the result becomes an automated check. It suits things like whether a diff satisfies its specification or whether an error message is actionable, and it is the wrong tool for anything a deterministic assertion can verify.

Can the same model write code and review it?

Not in the same run, and preferably not the same model. The reasoning that produced the change is still in context, so the grader evaluates its own intent rather than the artefact, which is why Anthropic's guidance recommends a fresh context or a separate subagent that tries to refute the result.

How do you know a model judge is accurate?

Sample its verdicts, have a person grade the same cases independently, and compare, then repeat on a schedule. Track false passes and false failures separately, and keep a fixed set of known-bad examples the judge must always fail so you can catch a regression in the grader itself.

Should a model judge be allowed to block a merge?

Only after it has proven reliable. Start advisory, gather agreement data against human reviewers, and promote it to a blocking check once the false-failure rate is low enough that engineers will not avoid it.

What makes a good rubric for a model judge?

A small scale such as pass or fail, or a three-point scale rather than a score out of ten, named criteria listed in the order that matters, a requirement that the judge quote the specific line or requirement its verdict rests on, and one judgment per call rather than grading correctness, style, and security together. Vague prompts asking whether something is good produce graders that rarely fail anything.

What is the difference between a false pass and a false fail from a model judge?

A false pass lets a bad change through undetected, which is the more expensive failure because it reaches production. A false fail wastes engineering time on a change that was actually fine, and repeated false fails teach the team to override the judge, at which point it stops functioning as a check at all. Both need to be tracked separately.

When is a model judge the wrong tool?

Whenever the property being checked has a single correct answer that a deterministic assertion could verify, since deterministic checks are faster, cheaper, and never argue. A model judge is also the wrong tool when two of your own engineers would disagree about whether an example passes, because that signals an unresolved specification problem rather than something a grader can settle.

How often should a model judge be re-validated?

On a repeating schedule, rather than only once at launch, because the codebase, the model behind the judge, and the population of changes being graded all change over time. A judge that has become lenient without anyone noticing produces false reassurance rather than an absence of signal, which is why a fixed set of known-bad examples the judge must always fail is worth keeping as a regression check on the grader itself.