Start here

The three kinds of LLM eval: code checks, human review and model graders

An LLM eval is one of three kinds. A code check is a short program that returns pass or fail, and it suits any question with one right answer. Human review is a person reading the output, and it is the reference the other two are measured against. A model grader, often called LLM-as-a-judge, is a second model that returns a verdict, and it must be compared with human labels before its verdict counts. In a 2023 study by Zheng and co-authors, GPT-4 as grader agreed with human labelers on 85 percent of the votes that were not ties. The rule for choosing is to use the cheapest kind that can decide the question.

Published September 30, 2026. Editorial.

Key takeaways

  • The three kinds of LLM eval are a check written in code, review by a person, and a second model used as a grader. Anthropic's and OpenAI's documentation both list these three under different names.
  • Use a code check for every question that has one right answer. A code check returns the same verdict on every run and costs only computer time.
  • In the 2023 study by Zheng and co-authors, GPT-4 as grader agreed with human labelers on 85 percent of the votes that were not ties, and the human labelers agreed with each other on 81 percent. The test covered 80 chat questions.
  • The same study measured faults: when two answers swapped places, GPT-4 kept its verdict in 65.0 percent of cases, and Claude-v1 in 23.8 percent. Run every comparison twice with the order swapped.
  • Start by reading real outputs, then write code checks for what code can decide, and add a model grader only for the failure types that remain and only after comparing it with human labels.

Take a support assistant for an invented furniture shop. A customer asks where an order is, and the assistant writes a reply. Three questions can be asked about that one reply. The first is whether it contains the order number the customer gave, and a few lines of code answer that in a fraction of a second. The second is whether the delivery date in the reply matches the date in the order record, and code answers that too. The third is whether a customer who is already upset would read the reply as polite. That question needs a reader: a person, or a second model whose verdicts have been compared with what people said.

Those are the three kinds of LLM eval. An LLM is a large language model, an AI model that produces text, and an eval is a repeatable test of what the model produced. This page describes each kind, gives the published measurements of how far a model can be trusted as a grader, and ends with one rule for choosing between them. It is the first step of the method in LLM evals: how to measure whether an AI product works.

The vendor guides name the same three kinds in different words

Four vendor guides were read for this page on 30 September 2026. Three of them sort evals into the same three kinds. The fourth, Google Cloud's, lists two of the three and has no category for a person reading the output [6].

Guide Check written in code A person reads it A model grades it
Anthropic documentation [1] Code-based grading Human grading LLM-based grading
OpenAI documentation [2] Metric-based evals Human evals LLM-as-a-judge and model graders
Anthropic engineering post, 9 January 2026 [5] Code-based graders Human graders Model-based graders
Google Cloud documentation [6] Computation-based metrics No category on the page Rubric-based metrics

"LLM-as-a-judge" is the name most people search for. It means the third column: a model reads an output and returns a verdict on it. A rubric, the word Google Cloud uses, is a written list of criteria with a description of what counts as a pass.

All four guides are companies describing their own recommended practice. The measurements on this page come from the research papers cited below.

Code checks answer questions that have one right answer

A code check is a short program that inspects the output and returns pass or fail. Anthropic's documentation describes the plainest one: "Exact match evals measure whether the model's output matches a predefined correct answer, typically after normalizing whitespace and case." [1] Other code checks confirm that the output contains a required phrase or stays under a word limit. OpenAI's documentation adds "function call accuracy", which means checking that the model asked for the right action with the right details [2].

Anthropic's engineering post lists five strengths of code-based graders: "Fast", "Cheap", "Objective", "Reproducible" and "Easy to debug" [5]. The weakness it names is that a code check fails a valid answer that is worded differently from the pattern the check expects [5]. A check that looks for the exact words "arrives on 4 March" fails a correct reply that says "delivery is set for March 4".

Write a code check whenever the question can be put in this form: the output must contain, equal, or stay within a stated value. Hamel Husain, a consultant who writes about evals, calls these checks unit tests, the name programmers use for small automatic tests, and says the pass rate you require is a product decision [7].

Human review is the reference for the other two kinds

In human review a person reads the output and records a judgment. OpenAI's documentation says "Human judgment evals provide the highest quality but are slow and expensive", and it names a second problem: "Disagreement among human experts" [2]. Anthropic's documentation is more direct: "Most flexible and high quality, but slow and expensive. Avoid if possible." [1]

We read "avoid if possible" as advice about routine runs: a person cannot reread 500 outputs every time the prompt, the written instruction given to the model, is changed. The same guides still rely on people. Anthropic's engineering post lists, as a strength of human graders, that they are "Used to calibrate model-based graders", meaning the model grader is adjusted until its verdicts match the human ones [5].

You need human review at four points:

  • Before any criteria exist. Somebody reads real outputs to find out how the product fails. That step is error analysis.
  • To label the reference set. A model grader is checked against cases a person has already marked pass or fail.
  • On a sample after release. New requests arrive that no existing check was written for.
  • When the other two kinds cannot decide. That case goes to a person.

Writing the labelling guideline and measuring whether two reviewers agree are covered in human review and reviewer agreement.

A model can grade another model, within measured limits

In the third kind, a second model receives the output and a rubric and returns a verdict. Two papers measured how well this works.

Zheng and twelve co-authors tested it in 2023 on 80 questions answered by 6 chat models, with 58 human labelers who had expert-level knowledge [3]. Counting only the votes that were not ties, GPT-4 as the grader agreed with the human labelers 85 percent of the time. The human labelers agreed with each other 81 percent of the time [3]. The paper's summary is that the grader reached "over 80% agreement, the same level of agreement between humans" [3].

Liu and five co-authors published G-Eval in the same year. G-Eval is a grader prompt made of written criteria, a list of evaluation steps that the model writes for itself, and a form to fill in [4]. On SummEval, a public test set of summaries rated by people, the GPT-4 version reached an average Spearman correlation of 0.514 with the human ratings [4]. A Spearman correlation is a number from -1 to 1 that says how closely two rankings match, where 1 means the grader put the summaries in the same order the people did. The figure differed by criterion: 0.582 for coherence and 0.455 for fluency [4].

The two results use different statistics and different tasks, so they cannot be compared with each other. Read together they say that how far a model grader matches people depends on the task. Both papers tested the models of 2023, on chat answers and summaries. Your product is a third task, and the grader has to be measured on it. Shankar and four co-authors put the consequence in one sentence in 2024: "LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation." [8]

The faults that have been measured in model graders

The same two papers measured specific ways a model grader goes wrong.

Fault What was tested Result Source
Position bias: favours an answer because of its position Two answers shown twice, in swapped order Same verdict both times: GPT-4 65.0 percent, GPT-3.5 46.2 percent, Claude-v1 23.8 percent Zheng et al. [3]
Verbosity bias: favours the longer answer 23 answers made longer by repeating a list Grader fooled: Claude-v1 91.3 percent, GPT-3.5 91.3 percent, GPT-4 8.7 percent Zheng et al. [3]
Wrong answers marked correct 10 math questions, 20 judgments GPT-4 passed a wrong answer in 14 of 20 (70 percent); with the correct answer in the prompt, 3 of 20 (15 percent) Zheng et al. [3]
Preference for machine-written text Summaries by GPT-3.5 and by people Grader scored GPT-3.5 summaries higher, including where human judges preferred the human one Liu et al. [4]

Three fixes follow from these results. For a comparison of two answers, Zheng and co-authors say to "call a judge twice by swapping the order of two answers and only declare a win when an answer is preferred in both orders" [3]. Where a correct answer exists, put it in the grader's prompt, the change that cut the math failures from 14 to 3. OpenAI's documentation adds: "Use pairwise comparison or pass/fail for more reliability" [2]. Pairwise comparison means the grader is shown two outputs and asked which is better.

Zheng and co-authors also looked for a grader that prefers its own output, and concluded that their data was too limited to say whether that bias exists [3]. The sample code in Anthropic's documentation still carries the comment "Generally best practice to use a different model to evaluate than the model used to generate the evaluated output" [1], and our argument for that rule is in never let the model grade its own work.

Designing the grader is a separate subject. Using a model as a judge covers rubrics and known-bad examples, and LLM as judge vs a decision model compares a text-writing grader with a model that returns only a probability.

The three kinds compared

Code check Human review Model grader
What it can check Anything with one right answer: format, required content, numbers, the action chosen Any question a competent reader can judge, including questions nobody has written down yet Questions of judgment that have a written rubric: tone, completeness, whether the answer follows its source
Cost per case Computer time only The paid time of a person who knows the subject One paid call to a model for each question asked
Repeatability The same output gets the same verdict on every run Two reviewers can disagree on the same output The verdict can change between runs and when the grader model changes
When to use it On every change, for every case To set the criteria, to label the reference set, to review a sample after release On each release, for judged questions, after the grader has matched human labels

The repeatability row comes from the sources. Anthropic's engineering post lists "Non-deterministic" as a weakness of model-based graders: the same input can produce a different verdict on a second run [5].

Use the cheapest kind that can decide the question

Anthropic's documentation states the rule in one line: "choose the fastest, most reliable, most scalable method" [1]. Its engineering post gives the order: "We recommend choosing deterministic graders where possible, LLM graders where necessary or for additional flexibility, and using human graders judiciously for additional validation." [5] A deterministic grader is one that returns the same verdict every time, which in practice means code.

Apply the rule to one failure type at a time, in three steps.

  1. If a program can decide it, write the code check and stop.
  2. If the criterion can be written so that two people who know the subject would give the same verdict, try a model grader, and compare it with human labels before its verdict counts.
  3. If the criterion cannot be written down yet, the question stays with a person until it can.

Many questions that look like pure judgment split into two parts. Take "the reply states the refund policy correctly" for the invented furniture shop. One part is a code check: the sentence the reply quotes appears word for word in the policy document. The other part is judged: the quoted policy is the one that applies to this customer's question.

Which kind to start with

Start by reading. Until someone has read real outputs, there is no list of failures to write checks for. After that reading, write code checks for every failure type that code can decide, and add a model grader only for the failure types that remain. LLM eval metrics describes the scoring options for each.

Husain describes the same order as three levels of rising cost: unit tests at level 1, human and model evaluation at level 2, and A/B testing at level 3, where part of the live traffic is given the new version and compared with the rest [7]. Husain runs level 1 on every code change, level 2 on a set schedule, and level 3 only after a large product change [7].

How Reveneau applies this

All of Reveneau's code is written by AI, and every change must pass a large eval suite before release. We write that suite from the specification, before the code exists. The suite uses the same three kinds in the same order of preference. Any check with one right answer is written as code and never goes to a model. The checks that need judgment go to a grader: Reveneau grades them with Jev, TypeSafe AI's decision model, which returns a probability for a question written in advance. On our own suite the run is ten times faster than with our previous language-model grader. People write the criteria, label the cases the grader is checked against, and review the grades that are close to the pass threshold.

For an AI product we build for a client, we apply the rule on this page to each failure type: the cheapest of the three kinds that can decide the question. Reveneau, as a company, takes responsibility for the whole project through production and after release, and the eval suite is how we show that a change works before it reaches users. The scope of that work is described on our AI development service page.

Best for

  • Code checks: format, required content, numbers, and any question with one right answer
  • Model graders: judged questions with a written rubric that run on every release
  • Human review: setting criteria, labelling the reference set, sampling after release

Avoid if

  • A model grader has never been compared with human labels on your own outputs
  • The criterion is still unwritten: keep that question with a person
  • A person would have to reread every output on every change

Check before you decide

  • Each failure type is scored by the cheapest kind that can decide it
  • Every pairwise comparison runs twice with the two answers swapped
  • The model that grades is a different model from the one that wrote the output

Common questions

What are the three kinds of LLM eval?

The three kinds of LLM eval are a check written in code, review by a person, and a second model used as a grader. Anthropic's documentation calls them code-based grading, human grading and LLM-based grading, and OpenAI's documentation calls them metric-based evals, human evals and LLM-as-a-judge. A code check suits questions with one right answer, and the other two kinds handle questions of judgment.

What is LLM-as-a-judge?

LLM-as-a-judge is the practice of giving one model's output to a second model, together with written criteria, and asking the second model for a verdict. A 2023 paper by Zheng and twelve co-authors tested the practice on 80 questions answered by 6 chat models, with 58 human labelers as the comparison. The verdict is only useful after the grader has been compared with labels given by people on the same outputs.

Can one AI model reliably grade the output of another model?

One model can grade another within limits that have been measured. In the 2023 study by Zheng and co-authors, GPT-4 as grader agreed with human labelers on 85 percent of the votes that were not ties across 80 chat questions, while the human labelers agreed with each other on 81 percent. The G-Eval paper measured a different statistic on summaries, a correlation of 0.514 with human ratings, so the result depends on the task.

When does an LLM eval need human review?

An LLM eval needs human review at four points: before any criteria exist, when labelling the cases a model grader is checked against, when sampling outputs after release, and when the other two kinds cannot decide a case. Anthropic's engineering post of 9 January 2026 lists human graders as the ones used to calibrate model-based graders, which is why people stay part of the process.

Which of the three kinds of LLM eval should a team start with?

A team should start with a person reading real outputs, because until that happens there is no list of failures to write checks for. After that, write code checks for every failure a program can decide, and add a model grader for what remains. Hamel Husain describes the same order as three levels of rising cost and runs the first level on every code change.

Which of the three kinds of LLM eval costs the least to run?

A code check costs the least, because it uses only computer time and needs no paid model call and no reviewer. Anthropic's engineering post lists code-based graders as fast, cheap, objective and reproducible. A model grader costs one paid model call for each question asked, and human review costs the time of a person who knows the subject, which is why it is kept for work only people can do.

What goes wrong when a model is used as a grader?

A model used as a grader can favour an answer for its position or its length, and it can pass a wrong answer. Zheng and co-authors found that GPT-3.5 and Claude-v1 were each fooled by a lengthened answer in 91.3 percent of 23 cases, and that GPT-4 passed a wrong math answer in 14 of 20 judgments with a default prompt.

How can position bias in a model grader be reduced?

Position bias in a model grader is reduced by running each comparison twice with the two answers in swapped order, and counting a win only when the same answer is preferred both times. That is the fix Zheng and co-authors propose. In their test, GPT-4 kept the same verdict after a swap in 65.0 percent of cases and Claude-v1 in 23.8 percent.

Should the model that wrote an answer also grade that answer?

Use a different model to grade than the one that wrote the answer. Zheng and co-authors looked for graders that prefer their own output and concluded that their data was too limited to confirm the bias. Anthropic's documentation still marks a different grading model as generally best practice in its sample code, and the G-Eval paper found its grader scored machine-written summaries above human-written ones.

Are code checks enough for a product that writes free text?

Code checks alone are too narrow for a product that writes free text, and many judged questions still contain a part that code can check. Anthropic's engineering post notes that a code check fails valid answers worded differently from the expected pattern. A judged question often splits into a code part, such as whether a quoted sentence appears in the source document, and a smaller judged part for a grader or a person.

Do the vendor guides agree on the kinds of LLM eval?

Three of the four vendor guides read on 30 September 2026 agree on the same three kinds of LLM eval. Anthropic's documentation, Anthropic's engineering post and OpenAI's documentation each list a code kind, a human kind and a model kind. Google Cloud's evaluation overview lists computation-based and rubric-based metrics and has no category for human review.

What should I do after deciding how each question will be scored?

After deciding how each question will be scored, build the set of cases the evals will run on. For a model grader, the next step is to label cases by hand and compare the grader with those labels. Shankar and co-authors wrote in 2024 that evaluators generated by a model need further human validation.

References