Choose how to score

Human review in LLM evals: guidelines, labels and agreement between reviewers

Human labels are the reference that every model grader is checked against, so their quality sets the limit for every automatic score. Have a person who knows the field label outputs as pass or fail with a written reason, following a written guideline. Have a second person label the same sample separately, then measure agreement with Cohen's kappa, which removes the agreement expected by chance. In this page's invented example, two reviewers agree on 85 of 100 outputs and kappa is 0.571. The bands used to describe kappa, such as Landis and Koch's from 1977, are a convention, so set your own level by what a wrong label would cost.

Published September 30, 2026. Editorial.

Key takeaways

  • A model grader is accepted when its verdicts match human labels, so an error in the labels becomes an error in every score the grader produces afterwards.
  • Label each output pass or fail with a written reason, using a guideline that has one definition, real examples and hard cases for every failure type.
  • Cohen's kappa is observed agreement minus chance agreement, divided by 1 minus chance agreement. In this page's invented example: (0.85 - 0.65) / (1 - 0.65) = 0.571.
  • Percent agreement can rise while kappa falls. In the second invented example, agreement is 92 percent and kappa is 0.291, because the reviewers agree on only 2 of 10 failed outputs.
  • Hamel Husain and Shreya Shankar advise labelling 100 to 200 examples for each failure type before relying on a model grader, with 40 to 45 percent kept for one final test.

In a 2012 review of agreement statistics, Mary McHugh cites a study in which raters gave the same rating in 94.2 percent of cases, while kappa, a measure of agreement that removes the part expected by chance, was 0.555 on the same data [5]. The first number suggests raters who agree. The second says the raters reached 55.5 percent of the agreement that was possible beyond chance.

This matters for an AI product because every model grader is measured against labels that people gave. If two people read the same output and cannot agree whether it passes, a model grader tuned to match one of them is matching one opinion. This page, part of the guide to LLM evals for AI products, covers who should label, how to write the guideline, how to measure agreement with Cohen's kappa, and how many labels a model grader needs.

Human labels set the limit for every other check

A model grader is a language model that reads an output and returns a verdict. It is accepted for use when its verdicts match the verdicts a person already gave on the same outputs. Anthropic's guide to evals lists this as a strength of human graders, "Used to calibrate model-based graders", and lists their weaknesses as "Expensive" and "Slow" [1]. A 2024 study by Shreya Shankar and four co-authors states the dependency directly: "LLM-generated evaluators simply inherit all the problems of the LLMs they evaluate, requiring further human validation" [4].

So an error in the labels becomes an error in the grader, and then in every score the grader produces. When to use each type of check is covered in the three kinds of LLM eval.

Who should label eval data

The labeller should be the person who knows what a good answer is in the product's field: a support lead for a support assistant, a lawyer for a contract summariser. Hamel Husain's guide calls this person the principal domain expert, and warns against developers taking the role themselves or handing it to a convenient substitute such as their manager [2].

Husain and Shankar recommend that small and medium-sized companies appoint one domain expert who has the final say on quality [3]. They add that "When you do use multiple people, you'll need to measure their agreement using metrics like Cohen's Kappa, which accounts for agreement beyond chance" [3].

Our position is one person with the final say, plus a second person who labels a sample. The second labeller tests the guideline. If only its author can apply a rule, the rule is known to one person only, and it cannot yet be written into a grader's instructions.

How to write a labelling guideline

A labelling guideline is the written rule that turns "is this output good?" into a decision two people would make the same way. Write one section for each failure type found in error analysis. Each section holds:

  1. One sentence that defines a pass, and one that defines a fail.
  2. Real outputs that pass and real outputs that fail, each with the reason.
  3. The cases that are hard to decide, with the ruling for each.
  4. What to do when unsure: mark the item "unsure" and write why.
  5. A version number and a date.

Anthropic gives the test for a finished guideline: "A good task is one where two domain experts would independently reach the same pass/fail verdict" [1].

Expect the guideline to change. Shankar and co-authors observed nine industry practitioners and named what they saw criteria drift, meaning that the rules for a good output change as the reviewer reads more outputs: "users need criteria to grade outputs, but grading outputs helps users define criteria" [4]. Their participants "still refined their criteria upon further grading, even going back to change previous grades" [4]. Nine people in one session is a small study with no rate attached. The instruction it supports: when a rule changes, raise the version number and label the earlier items again under the new rule.

Use pass or fail with a written reason

Husain's instruction to the expert is short: "No complex scoring scales or multiple metrics. Just a clear pass or fail decision. In addition to the pass/fail decision, the domain expert should write a critique that explains their reasoning" [2]. Husain and Shankar give three problems with 1 to 5 scales: "the difference between adjacent points (like 3 vs 4) is subjective and inconsistent across annotators, detecting statistical differences requires larger sample sizes, and annotators often default to middle values to avoid making hard decisions" [3]. The wider comparison of scoring scales is in LLM eval metrics.

The written reason has two uses. A disagreement can be traced to one sentence of the guideline, and the reasons later become the worked examples inside a model grader's instructions.

How to measure agreement between two reviewers

Inter-annotator agreement, also called inter-rater reliability, is how often two people who label the same items separately give the same label. The simplest measure is percent agreement, and two people who pass most outputs will agree often by chance alone. McHugh's review gives the history: "In 1960, Jacob Cohen critiqued use of percent agreement due to its inability to account for chance agreement" [5]. The measure that corrects for chance agreement is named Cohen's kappa after that paper.

The formula, as McHugh prints it, is kappa = (Pr(a) - Pr(e)) / (1 - Pr(e)), "Where Pr(a) represents the actual observed agreement, and Pr(e) represents chance agreement" [5]. Kappa is 1 when the reviewers always agree and 0 when they agree only as often as chance predicts. Cohen's kappa covers two raters [5].

Here is an invented illustration. Two reviewers each label the same 100 outputs of a support assistant.

Reviewer B: pass Reviewer B: fail Total
Reviewer A: pass 70 10 80
Reviewer A: fail 5 15 20
Total 75 25 100

The arithmetic, step by step:

  1. Observed agreement: (70 + 15) / 100 = 0.85.
  2. Chance that both pass: 0.80 × 0.75 = 0.60. Chance that both fail: 0.20 × 0.25 = 0.05. Chance agreement: 0.60 + 0.05 = 0.65.
  3. Kappa: (0.85 - 0.65) / (1 - 0.65) = 0.20 / 0.35 = 0.571.

Now change the counts so that most outputs pass: both pass on 90, both fail on 2, and they split on 8 (4 each way). Observed agreement rises to (90 + 2) / 100 = 0.92. Each reviewer passes 94, so chance agreement is 0.94 × 0.94 + 0.06 × 0.06 = 0.8836 + 0.0036 = 0.8872. Kappa is (0.92 - 0.8872) / (1 - 0.8872) = 0.0328 / 0.1128 = 0.291.

Percent agreement went up and kappa went down. The counts show why: 10 outputs were failed by one reviewer or both, and the two agreed on only 2 of them. An eval exists to find failures, so report the table of counts together with kappa.

With three or more reviewers, McHugh lists two other measures: Fleiss kappa, an "adaptation of Cohen's kappa for 3 or more raters", and Krippendorff's alpha, "useful when there are multiple raters and multiple possible ratings" [5].

What counts as a good kappa

No rule settles this. Two of this page's sources reproduce the bands proposed by J. R. Landis and G. G. Koch in the journal Biometrics in 1977 [6][7]. We read the wording below in a 1998 technical report by Khaled El Emam [6]; the 1977 paper was not open to us.

Kappa Landis and Koch label [6] Worked example above
Below 0.00 Poor
0.00 to 0.20 Slight
0.21 to 0.40 Fair 0.291
0.41 to 0.60 Moderate 0.571
0.61 to 0.80 Substantial
0.81 to 1.00 Almost Perfect

Treat these words as a convention. El Emam reports that "Landis and Koch concede that their benchmark is arbitrary, but they nevertheless contend that it can serve as a useful guideline" [6]. A 2013 conference paper by Xie prints the same ranges with the top band labelled "excellent", and quotes a 2002 criticism that the approach "has no sound theoretical basis and can be positively misleading to investigators" [7].

McHugh proposes stricter labels for healthcare research: 0 to .20 None, .21 to .39 Minimal, .40 to .59 Weak, .60 to .79 Moderate, .80 to .90 Strong, above .90 Almost Perfect [5]. Under that table the 0.571 in the example is "Weak" where Landis and Koch call it "Moderate".

Klaus Krippendorff, writing in 2004 on the related measure alpha, gives the principle we follow: "An acceptable level of agreement below which data are to be rejected as too unreliable must be chosen depending on the costs of drawing invalid conclusions from these data" [8]. For scholarly work Krippendorff suggests .800, and .667 where tentative conclusions are acceptable [8].

Eugene Yan writes, from personal practice with no study behind it, "I often see human inter-rater reliability (Cohen's Kappa) being as low as 0.2 - 0.3" [9]. Choose your level before labelling starts, by what a wrong label would cost, and write it into the guideline.

Resolve disagreements by fixing the guideline

Husain and Shankar's rule for the process is "Have annotators label the same examples independently before they discuss them" [3]. After that:

  1. List every item where the two labels differ.
  2. Read both written reasons for each item.
  3. Sort each item into one of three causes: the guideline said nothing, the guideline could be read two ways, or one reviewer made a mistake.
  4. For the first two causes, write the missing rule and add the item to the guideline as an example.
  5. Measure kappa again on items neither reviewer has discussed.

A vote between reviewers settles one item. A changed guideline settles every later item of the same kind, which is why we prefer it.

How many labels a model grader needs

The figures here are advice from named practitioners, with no dataset behind them. Husain and Shankar: "Plan to label 100 to 200 examples for each failure mode" [3]. They split the labels three ways: 10 to 20 percent as examples that may appear in the grader's instructions, 40 to 45 percent for refining the grader, and 40 to 45 percent kept for one final test [3]. That final part is a held-out set: cases nobody looked at while writing the grader. They also ask for 30 to 50 passes and 30 to 50 fails in both of the last two parts [3]. Yan recommends "at least 50-100 failures out of 200+ total samples" and says to "Evaluate these evaluators on precision, recall, and Cohen's Kappa" [9].

Here is an invented illustration on 200 held-out outputs. The person failed 60. The grader failed 66, and 48 of those were among the person's 60. Both passed 122.

Measure Arithmetic Result
Precision: share of the grader's fails that the person also failed 48 / 66 0.727
Recall: share of the person's fails that the grader found 48 / 60 0.800
Observed agreement (48 + 122) / 200 0.850
Chance agreement (66/200 × 60/200) + (134/200 × 140/200) 0.568
Kappa (0.850 - 0.568) / (1 - 0.568) 0.653

The grader missed 12 real failures and raised 18 false ones. Whether that is acceptable depends on which mistake costs more in your product. The same check for a decision model is in calibrating Jev against your own human labels, and the argument for keeping the grader separate from the writer is in never let the model grade its own work.

How Reveneau applies this

At Reveneau all code is written by AI, and every change must pass a large eval suite written from the specification before the code exists. The judged checks in that suite are graded by Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than it was with our previous language-model grader. A grader's verdicts can be trusted as far as the labels it was checked against. So we write the labelling guideline before the first label, record each label as pass or fail with a reason, and ask a second person to label a sample, so that agreement can be measured and reported with its table of counts.

Reveneau, as a company, takes responsibility for the whole project through production and after release. For human review, that means the guideline, the labels and the agreement figures are kept with the eval suite, where a client can read them. To set this up for your own product, see AI development at Reveneau or contact us.

Best for

  • Checks where the right answer is a judgment: tone, completeness, whether a summary is fair
  • Building the labelled set that a model grader is checked against
  • Products where a wrong pass has a cost someone can name

Avoid if

  • The check has one exact answer that code can compare: use a code check
  • Nobody who knows the field can give time to labelling
  • The failure types have not been found yet: read outputs first

Check before you decide

  • The guideline has a version number and a date
  • Two people labelled the same sample separately before any discussion
  • Kappa is reported with the table of counts behind it
  • The grader's final test used labels nobody saw while writing it

Common questions

What is inter-annotator agreement?

Inter-annotator agreement is how often two people who label the same items separately give the same label. Another name for the same thing is inter-rater reliability. For an LLM eval, the items are outputs of the product and the labels are pass or fail. Percent agreement is the simplest measure, and Cohen's kappa is the measure that also removes the agreement two reviewers would reach by chance.

What is Cohen's kappa?

Cohen's kappa is a measure of agreement between two raters that subtracts the agreement expected by chance. The formula, as Mary McHugh's 2012 review prints it, is observed agreement minus chance agreement, divided by 1 minus chance agreement. Kappa is 1 when two reviewers always agree and 0 when they agree only as often as chance predicts. Jacob Cohen published the measure in 1960.

What is a good kappa for eval labels?

A good kappa has no fixed definition, and the common labels are a convention. Landis and Koch's 1977 bands, as reproduced by El Emam in 1998, call 0.41 to 0.60 Moderate and 0.61 to 0.80 Substantial. McHugh's 2012 table calls .40 to .59 Weak. Krippendorff's principle is to choose the acceptable level by the cost of drawing a wrong conclusion from the labels.

Who should label eval data for an AI product?

Eval data should be labelled by a person who knows what a good answer is in the product's field, such as a support lead for a support assistant. Hamel Husain calls this person the principal domain expert and warns against developers taking the role. Husain and Shreya Shankar recommend one expert with the final say for small and medium-sized companies, with agreement measured when several people label.

How do I write a labelling guideline for LLM outputs?

Write a labelling guideline as one section for each failure type. Each section holds a one-sentence definition of pass and of fail, real outputs that pass and fail with the reason, the hard cases with a ruling, and what to do when unsure. Add a version number and a date. Anthropic's test is that two domain experts would independently reach the same pass or fail verdict.

How do I check a model grader against human labels?

Check a model grader by running it on outputs a person has already labelled and that nobody used while writing the grader. Eugene Yan advises scoring it on precision, recall and Cohen's kappa. In this page's invented example of 200 outputs, the grader's precision, the share of its fails that the person also failed, is 48 / 66 = 0.727, its recall, the share of the person's fails that it found, is 48 / 60 = 0.800 and its kappa against the person is 0.653.

How many human labels does a model grader need?

Hamel Husain and Shreya Shankar advise 100 to 200 labelled examples for each failure type, with 10 to 20 percent used as examples in the grader's instructions, 40 to 45 percent for refining it and 40 to 45 percent kept for one final test. Eugene Yan recommends at least 50 to 100 failures among 200 or more samples. Both figures are practitioner advice with no study behind them.

Why is percent agreement misleading on its own?

Percent agreement counts agreement that would happen by chance. When most outputs pass, two reviewers often give the same label without holding the same definition of a failure. In this page's second invented example, two reviewers agree on 92 of 100 outputs and kappa is only 0.291, because they agree on 2 of the 10 outputs that either one failed. McHugh's review cites a study with 94.2 percent agreement and kappa of 0.555.

What should two reviewers do when they disagree on a label?

Two reviewers who disagree should compare their written reasons and find the cause in the guideline. Either the guideline said nothing about the case, or it could be read two ways, or one reviewer made a mistake. For the first two causes, write the missing rule and add the item as an example. Husain and Shankar advise that reviewers label independently before any discussion.

Why do the labelling rules keep changing while reviewers label?

Labelling rules change because reading outputs teaches reviewers what they care about. A 2024 study by Shreya Shankar and co-authors observed nine industry practitioners and named this criteria drift: people need rules to grade outputs, and grading outputs helps them define the rules. Participants went back and changed earlier grades. Give the guideline a version number, and label earlier items again when a rule changes.

Does a small team with one expert still need a second reviewer?

A small team can let one expert decide every label, which is what Husain and Shankar recommend for small and medium-sized companies. A second person labelling a sample is still worth the time, because it tests whether the guideline can be followed by someone other than its author. A rule that only one person can apply cannot yet be written into a model grader's instructions.

What do I use when more than two people label the same outputs?

Cohen's kappa covers two raters. For three or more, Mary McHugh's 2012 review lists Fleiss kappa, an adaptation of Cohen's kappa, and Krippendorff's alpha, which handles several raters and several possible ratings. For alpha in scholarly work, Krippendorff suggested .800, and .667 where tentative conclusions are acceptable, and wrote that the level should depend on the cost of invalid conclusions.

References