LLM evals: how to measure whether an AI product works
LLM evals measure whether a product built on a large language model does its job. The method has a fixed order: learn the three kinds of eval, read real outputs and name the failure types, build a test set from real usage, size the set for the difference you need to detect, choose a score for each failure type, and keep measuring after release. A 90 percent pass rate on 100 cases has a 95 percent confidence interval of 84.1 to 95.9 percent, by the formula in Evan Miller's 2024 paper, so a small set cannot confirm a small improvement. Each step in this guide is tied to a named paper or vendor guide.
Published September 30, 2026. Editorial.
Key takeaways
- An LLM eval is one of three kinds: a check written in code, review by a person, or a second model used as a grader. Use the cheapest kind that can decide the question.
- Read real outputs before choosing any metric. Hamel Husain and Shreya Shankar recommend at least 100 traces, and a 2024 study with nine practitioners found that grading criteria change as reviewers read more.
- No vendor page read on 30 September 2026 states a minimum test set size. Anthropic's engineering post suggests 20 to 50 tasks drawn from real failures as a start, which is that company's own advice.
- A 90 percent pass rate on 100 cases has a 95 percent confidence interval of 84.1 to 95.9 percent. On 1,000 cases the interval is 88.1 to 91.9 percent.
- Choose a metric for each failure type. A 2018 review of 284 correlations in 34 papers found that the evidence does not support using BLEU outside machine translation.
- In a product that searches documents before answering, test the search step and the answer separately, because a single score on the final answer cannot show which of the two failed.
- A model grader can only be as accurate as the human labels it is checked against. Have two people label the same cases and measure their agreement with Cohen's kappa.
- A hosted model can change under one name: a 2023 paper measured GPT-4 at 84 percent on one task in March and 51 percent in June. Rerun your evals on a schedule after release.
In 2023 three researchers, Lingjiao Chen, Matei Zaharia and James Zou, gave the same questions to GPT-4 in March and again in June. On one task, telling prime numbers from composite numbers, the March version was right 84 percent of the time and the June version 51 percent [1]. The service had the same name on both dates. Their conclusion was that a hosted model, one that a vendor runs and sells access to, can change its behaviour within a few months and has to be monitored continuously [1].
A product built on that model behaved differently in June, with no edit by the team that built it. The way to find that out is a test the team owns and can run again. That test is an eval. An LLM is a large language model, an AI model that produces text, and an LLM eval is a repeatable test of what a product built on one produces. OpenAI's documentation gives the reason ordinary software tests are insufficient: "Models sometimes produce different output from the same input, which makes traditional software testing methods insufficient for AI architectures." [2]
This guide is the method, in the order the work is done, and each step has its own page. It covers the output of an AI product such as a support assistant or a document search. Five related subjects have their own guides: the business reasons in why AI evals matter, testing AI-written code in eval-driven development, grader design with a decision model in evals with Jev, products that take actions in AI agent evals, and public model scores in AI benchmarks vs your own evals.
The method in order
| Step | The question it answers | Where it is explained |
|---|---|---|
| 1. Know the three kinds of eval | Which check can decide this question at the lowest cost | The three kinds of LLM eval |
| 2. Read real outputs | How this product fails, and how often | Error analysis |
| 3. Build the test set | Which cases every version is run on | How to build an LLM eval dataset |
| 4. Size the set | How far the pass rate can be trusted | How many test cases an LLM eval needs |
| 5. Choose a score for each failure type | What is counted, and where that count misleads | LLM eval metrics |
| 6. Split a search-based product in two | Whether the search step failed or the answer did | RAG evaluation |
| 7. Check the human labels | Whether two reviewers agree | Human review and reviewer agreement |
| 8. Keep measuring after release | What changed once real users arrived | Offline evals vs online evals |
Each section below gives the main point of one step, with one fact from a named source. The linked pages hold the detail.
Three kinds of eval: code checks, human review and model graders
Every eval is one of three kinds. Anthropic's documentation names them code-based grading, human grading and LLM-based grading [3]. A code check is a short program that returns pass or fail. Human review is a person reading the output. A model grader, often called LLM-as-a-judge, is a second model that returns a verdict.
Each has its use. A code check costs only computer time and gives the same verdict on every run, so it suits any question with one right answer: the reply contains the order number, or the date matches the record. Human review can judge anything a competent reader can judge, and it is slow. A model grader is fast on questions of judgment, and its verdicts count only after they have been compared with human labels.
One measurement of model graders comes from Zheng and twelve co-authors in 2023. On 80 chat questions, GPT-4 as grader agreed with human labelers on 85 percent of the votes that were not ties, while the labelers agreed with each other on 81 percent [4]. The same paper found that when two answers swapped places, GPT-4 kept its verdict in 65.0 percent of cases [4]. A grader whose verdict depends on the order of the answers has to be checked on your own outputs. A related rule is argued in never let the model grade its own work.
The rule, in the words of Anthropic's documentation, is to "choose the fastest, most reliable, most scalable method" [3]. Use the cheapest kind that can decide the question. The three kinds of LLM eval compares them in a table and lists the measured faults of model graders.
Read the outputs before choosing a metric
The first piece of work is reading. Take a random sample of real outputs, write a short note on each one that failed, group the notes into failure types, and count each type. Only then decide what to measure. This step is called error analysis.
Hamel Husain and Shreya Shankar, who teach a course on evals, recommend reviewing at least 100 traces and stopping when new traces show no new failure type [5]. A trace is the full record of one request. The number is their own working guide, with no dataset behind it.
The research behind the step is a 2024 study by Shankar and four co-authors. In a qualitative study (one that observes what people do and reports it in words) with nine industry practitioners, they found that people changed their grading criteria as they graded more outputs [6]. They named this criteria drift, meaning the criteria change as the reading goes on: "users need criteria to grade outputs, but grading outputs helps users define criteria" [6]. Criteria written before any reading therefore miss failures that the reading would have shown.
A ready-made score skips this step and measures a property somebody else chose. OpenAI's documentation lists "Overly generic metrics" among common mistakes in eval design [2]. Error analysis: read the outputs before you choose a metric gives the four steps, with a worked table of failure types and the check each one leads to.
Build the test set from real usage
A test set is the list of cases that every version of the product is run on. Each case holds an input, the context the model was given, the expected behaviour, and tags that say which failure type it tests. A set whose expected answers a person has confirmed is often called a golden dataset.
Cases should come from what users do. OpenAI's documentation lists the possible sources: "Consider synthetic eval data, domain-specific eval data, purchased eval data, human-curated eval data, production data, and historical data." [2] Synthetic data means cases written by a model. They help cover rare requests, and Husain and Shankar state their limit: "Synthetic data cannot tell you how common a failure is in production." [5]
On size, no vendor page read for this guide on 30 September 2026 states a minimum. The numbers that exist are suggestions from named people and companies. Anthropic's engineering post of 9 January 2026 says 20 to 50 simple tasks drawn from real failures are enough to begin [7]. Husain and Shankar write: "A purpose-built eval set often grows to 100 or more examples." [5]
Two more rules apply. Include cases where a behaviour should happen and cases where it should stay absent [7]. Keep part of the set away from the people who write the prompts, the written instructions given to the model, so that each prompt is tested on cases its writer never saw. How to build an LLM eval dataset from real usage covers sources, synthetic cases, personal data and versioning.
Size the set for the difference you need to detect
A pass rate from a small set fits a wide range of true values. Evan Miller's 2024 paper on eval statistics gives the formula for pass or fail scores. The standard error measures how far the pass rate would move if the test were repeated on a fresh set of cases. It is the square root of p × (1 - p) / n, where p is the pass rate and n is the number of cases, and the 95 percent confidence interval is the pass rate plus and minus 1.96 standard errors [8]. A confidence interval is the range in which the true pass rate plausibly lies.
| Cases | Standard error | 95 percent interval for a 90 percent pass rate |
|---|---|---|
| 100 | 0.0300 | 84.1 to 95.9 percent |
| 1,000 | 0.0095 | 88.1 to 91.9 percent |
The working for the first row: 0.9 × 0.1 / 100 = 0.0009, the square root of 0.0009 is 0.03, and 1.96 × 0.03 = 0.0588, which is 5.88 percentage points on each side of 90. The second row uses the same steps with n = 1,000.
A version that scores 92 percent on 100 cases is inside the interval for a 90 percent pass rate, so the 2 point difference could be chance. Miller recommends comparing two versions on the same cases and analysing the difference case by case, and answering each case several times when the output varies between runs [8]. The paper's own example arrives at 969 cases to detect a 3 point difference, under assumptions the paper itself calls fictional [8]. How many test cases an LLM eval needs has the full table and the paired comparison.
Choose a metric for each failure type
A metric is the rule that turns an output into a number. The main families are exact match and other checks in code, word-overlap scores such as BLEU, similarity scores that compare meaning, rubric scores given by a person or a model, and pairwise comparison, where a judge is shown two outputs and picks the better one. A rubric is a written list of criteria.
Each family was built for one purpose. BLEU was published in 2002 by four researchers at IBM to score machine translation against human reference translations, by counting shared word sequences [9]. A 2018 review by Ehud Reiter covered 284 correlations, each a measured match between BLEU and human ratings, reported in 34 papers. It concluded that the evidence "does not support using BLEU outside of MT, for evaluation of individual texts, or for scientific hypothesis testing" [10]. MT is machine translation. A support assistant's reply has no reference translation, so a BLEU score says little about whether the reply is correct.
For judged questions, Husain and Shankar prefer pass or fail to a 1 to 5 scale: "Binary evaluations force clearer thinking and more consistent labeling." [5]
Our position is that a metric is chosen for one failure type at a time. One overall score hides the failure types inside an average. In an invented illustration, a support assistant can average 4.3 out of 5 for helpfulness while one reply in ten gives a wrong delivery date. LLM eval metrics has a table of what each metric counts and when it fails.
Test the search step and the answer separately in a RAG system
RAG stands for retrieval-augmented generation: a search step fetches documents and the model writes its answer from them. The name comes from a 2020 paper by Patrick Lewis and eleven co-authors [11].
Two different things can fail. The search step can fetch the wrong passages, or the model can answer wrongly from the right ones. A single score on the final answer cannot tell you which of the two happened, so each step gets its own test.
For the answer, the Ragas paper of 2023 defines three measures that need no reference answer written by a person: faithfulness, answer relevance and context relevance [12]. Faithfulness asks whether the answer is supported by the fetched text. A model splits the answer into statements and checks each one against that text, and the score is the number of supported statements divided by the total number of statements [12]. An answer with 10 statements, 8 of them supported, scores 8 / 10 = 0.8.
The paper tested the measures on WikiEval, a set built from 50 Wikipedia pages and labelled by two people. Agreement with those human labels was 0.95 for faithfulness, 0.78 for answer relevance and 0.70 for context relevance [12]. The authors make the Ragas tool, and they name context relevance as the hardest of the three to evaluate [12]. RAG evaluation explains each measure and maps what the user sees to the step that failed.
Human labels set the limit for every other measure
A model grader is checked against labels given by people, so the grader can only be as accurate as those labels. Good labelling has four parts: a written guideline, a pass or fail label with a written reason, two people labelling the same cases independently, and a measure of how often they agree.
Plain percent agreement overstates how well two reviewers agree, because two people who both pass most cases will often agree by chance. Cohen's kappa corrects for that. As Mary McHugh's 2012 review prints it, kappa = (Pr(a) - Pr(e)) / (1 - Pr(e)), where Pr(a) is the agreement observed and Pr(e) is the agreement expected by chance [13]. McHugh cites a study in which percent agreement was 94.2 percent and kappa was 0.555 on the same data [13].
The usual names for kappa values are the bands proposed by Landis and Koch in 1977, in which 0.41 to 0.60 is "Moderate" and 0.61 to 0.80 is "Substantial". We read those bands in a secondary source, a 1998 technical report by Khaled El Emam, which also records that Landis and Koch conceded their benchmark is arbitrary [14]. Treat the bands as a convention.
When two reviewers disagree, correct the guideline and label again. Husain and Shankar add a rule for the process: reviewers label the same examples independently before they discuss them [5]. Human review in LLM evals works through a kappa calculation and the number of labels needed to check a model grader.
Keep measuring after release
An offline eval runs a fixed set of cases before release. An online eval scores live traffic after release. LangChain's documentation for its LangSmith product defines the first this way: "Offline evaluations target examples from datasets: curated test cases with reference outputs that define what "good" looks like." [15] It defines the second as evaluation of "real production traces without reference outputs" [15].
Each one finds a different class of problem. The offline set shows whether a change made known cases better or worse, and it holds only the requests somebody put into it. Online measures reach the new requests, and they report a problem after users have already seen it. The change that Chen and co-authors measured in GPT-4 is the reason to rerun the offline set on a schedule, including in months when your own code stayed the same [1].
The two form a cycle. A failure found online becomes a new offline case, the offline run confirms the fix, and the online measure confirms it on live traffic [15]. Husain and Shankar suggest reading 100 or more fresh traces in each review cycle of 2 to 4 weeks [5].
Tools change as well. OpenAI's documentation, read on 30 September 2026, says of its hosted Evals platform: "Evals will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026." [2] Keep your cases and labels in files you own, so that a tool closing takes none of them away. Offline evals vs online evals compares the two and lists the open-source eval tools with their licences.
How Reveneau applies this method
All of Reveneau's code is written by AI, and every change must pass a large eval suite before release. We write that suite from the specification, the written description of what the product must do, before the code exists. For an AI product, we build the suite by the method in this guide. We read real outputs with a person on the client's side who knows the correct answers, and we write the failure types into the specification. We score each failure type with the cheapest of the three kinds of eval that can decide it, and we report every pass rate with its interval and the number of cases behind it.
Reveneau grades the judged checks in its suite with Jev, TypeSafe AI's decision model, which returns a probability for a question written in advance. On our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release, and that includes the measuring after release. Our AI development service page describes the work.
Explore the guide
Start here
The three kinds of LLM eval: code checks, human review and model graders
An LLM eval is one of three kinds. A code check is a short program that returns pass or fail, and it suits any question with one right answer. Human review is a person reading the output, and it is the reference the other two are measured against. A model grader, often called LLM-as-a-judge, is a second model that returns a verdict, and it must be compared with human labels before its verdict counts. In a 2023 study by Zheng and co-authors, GPT-4 as grader agreed with human labelers on 85 percent of the votes that were not ties. The rule for choosing is to use the cheapest kind that can decide the question.
Error analysis: read the outputs before you choose a metric
Error analysis is the step that comes before any metric: read a sample of real outputs, write a short note on each failure, group the notes into failure types, and count them. Only then choose what to measure. Hamel Husain and Shreya Shankar recommend reading at least 100 traces and stopping when new ones show no new failure type. A 2024 study by Shankar and four co-authors, with nine practitioners, found that reviewers changed their criteria as they read more outputs, which is why criteria written before the reading miss failures. The reader should be a person who knows what a correct answer is.
Build the test set
How to build an LLM eval dataset from real usage
An eval is a repeatable test of an AI product's output. For a product built on a large language model (LLM), the eval dataset is a list of real requests, each stored with the context the model received and a written statement of what a correct response must do. Build it from your own logs, support tickets and past requests, add difficult requests on purpose, and use model-written cases only to fill gaps. Keep one part of the set away from the people who write the prompts, meaning the instructions given to the model. Replace personal data before a case is stored, and give every change a version number. Anthropic's engineering team suggests 20 to 50 cases to start. The vendor documentation read on 30 September 2026 states no minimum.
How many test cases an LLM eval needs: sample size and error bars
A pass rate of 90 percent measured on 100 test cases fits a true rate anywhere from 84.1 to 95.9 percent, using the simple 95 percent interval. On 1,000 cases the same result narrows to 88.1 to 91.9 percent. The number of cases needed to evaluate a large language model (LLM) product therefore depends on the size of the difference you want to detect. A set of 20 to 50 cases, the starting size Anthropic suggests, can show a large failure. Stating a rate to within 3 points takes 385 cases at a 90 percent pass rate. Detecting a move from 90 to 92 percent takes 1,666 cases by the formula in the NIST statistics handbook, and the comparison is best made with both versions run on the same cases.
Choose how to score
LLM eval metrics: exact match, similarity scores, rubrics and pairwise comparison
Metrics for evaluating a product built on a large language model (LLM) belong to five families. Exact match and code assertions check an output that has one right form. BLEU and ROUGE count the words an output shares with a reference text written by a person. BERTScore compares meaning instead of exact words. A rubric score is a judgement by a person or a model against written criteria. Pairwise comparison asks which of two outputs is better. Each family answers a different question, so choose one metric for each failure type you found when reading real outputs, and report each one as its own pass rate. A single overall score hides which failure type changed.
RAG evaluation: test the search step and the answer separately
A RAG system (retrieval-augmented generation) has two steps: a search step fetches passages from your documents, and a language model writes an answer from them. Either step can cause a wrong answer, and the repair is different for each, so test them separately. Test the search step in code by checking that the expected passage came back and at which position. Score the answer for faithfulness and answer relevance, and the fetched passages for context relevance, as the 2023 Ragas paper defines them. On the paper's own test of 50 Wikipedia pages, those measures agreed with human annotators at 0.95, 0.78 and 0.70.
Human review in LLM evals: guidelines, labels and agreement between reviewers
Human labels are the reference that every model grader is checked against, so their quality sets the limit for every automatic score. Have a person who knows the field label outputs as pass or fail with a written reason, following a written guideline. Have a second person label the same sample separately, then measure agreement with Cohen's kappa, which removes the agreement expected by chance. In this page's invented example, two reviewers agree on 85 of 100 outputs and kappa is 0.571. The bands used to describe kappa, such as Landis and Koch's from 1977, are a convention, so set your own level by what a wrong label would cost.
Common questions
What does an LLM eval measure in an AI product?
An LLM eval measures whether the output of a product built on a large language model meets a stated expectation, on a set of cases that can be run again. OpenAI's documentation explains why ordinary software tests are insufficient here: models sometimes produce different output from the same input. An eval therefore reports a pass rate over many cases instead of one pass or one fail.
What is the first step in evaluating a product built on an LLM?
The first step in evaluating a product built on an LLM is to read a random sample of real outputs and write down how each failed one failed. Hamel Husain and Shreya Shankar recommend reviewing at least 100 traces. A 2024 study by Shankar and four co-authors found that reviewers change their criteria as they read, so criteria written before the reading miss failures.
In what order should a team build its LLM evals?
A team should build its LLM evals in this order: read real outputs and name the failure types, build a test set from real usage, size the set, choose a score for each failure type, check the human labels, and keep measuring after release. Anthropic's engineering post of 9 January 2026 says 20 to 50 tasks drawn from real failures are enough for a first set.
Why are ordinary software tests insufficient for a product built on an LLM?
Ordinary software tests assume that the same input always gives the same output, and a large language model can give different outputs for one input. OpenAI's documentation says models sometimes produce different output from the same input, which makes traditional testing methods insufficient. An LLM eval handles this by running many cases, sometimes each case several times, and reporting a rate with its interval.
Is a pass rate from 100 test cases precise enough to compare two versions?
A pass rate from 100 test cases is precise enough to show a large difference and too imprecise for a small one. By the formula in Evan Miller's 2024 paper, a 90 percent pass rate on 100 cases has a 95 percent confidence interval of 84.1 to 95.9 percent. A second version scoring 92 percent is inside that interval, so the difference could be chance.
Is there an official minimum size for an LLM eval dataset?
There is no official minimum size for an LLM eval dataset on any vendor page read for this guide on 30 September 2026. The figures that exist are suggestions from named people and companies. Anthropic's engineering post suggests 20 to 50 tasks to start, and Husain and Shankar write that a purpose-built set often grows to 100 or more examples.
Can a general metric such as BLEU show whether my AI product works?
A general metric such as BLEU cannot show whether an AI product works, unless the product is a translator. BLEU was published in 2002 to score machine translation against human reference translations. A 2018 review by Ehud Reiter, covering 284 correlations in 34 papers, found that the evidence does not support using BLEU outside machine translation or on individual texts.
Why does an AI product need retesting when nothing in it has changed?
An AI product needs retesting because the hosted model underneath it can change while the product's own code stays the same. A 2023 paper by Chen, Zaharia and Zou measured GPT-4 at 84 percent accuracy on one task in March 2023 and 51 percent in June 2023, under the same service name. Running the same eval set on a schedule shows such a change.
Does this guide cover evals for code written by AI?
The LLM evals guide covers evals for the output of an AI product, such as a support assistant or a document search. Evals for code written by AI are a separate subject with their own Reveneau guide on eval-driven development. The two share the three kinds of check that Anthropic's documentation names: code-based grading, human grading and grading by a model.
Do I need to buy an eval tool before I start?
An eval tool is optional at the start, because the first steps need only saved outputs and a person who knows the correct answers to read them. Tools also change: OpenAI's documentation says its hosted Evals platform becomes read-only on 31 October 2026 and shuts down on 30 November 2026. Keep your cases and labels in files you own.
Does the same method work for an AI agent that takes actions?
The same method works for an AI agent as a starting point, and an agent needs further tests as well. Reading outputs, naming failure types and building a test set from real usage all apply. Anthropic's advice to start with 20 to 50 tasks comes from a post written for agents. Reveneau's separate guide on AI agent evals covers the steps an agent takes.
How does Reveneau use LLM evals when building an AI product?
Reveneau writes the eval suite from the specification before the code exists, and every change must pass it before release. For an AI product, the failure types found by reading real outputs go into that specification. Reveneau grades the judged checks with Jev, TypeSafe AI's decision model, and on its own suite the run is ten times faster than with its previous language-model grader.
References
- [1] Chen, Zaharia and Zou, How is ChatGPT's behavior changing over time? (arXiv, submitted 18 July 2023, abstract read): GPT-4 at 84 percent accuracy in March 2023 and 51 percent in June 2023 on identifying prime against composite numbers, and the conclusion that a hosted model needs continuous monitoring.
- [2] OpenAI, Evaluation best practices (API docs, undated, read 30 September 2026): why traditional software testing is insufficient; the list of dataset sources; overly generic metrics as a mistake; the notice that the hosted Evals platform becomes read-only on 31 October 2026 and shuts down on 30 November 2026.
- [3] Anthropic, Define success criteria and build evaluations (Claude Platform Docs, undated, read 30 September 2026): code-based, human and LLM-based grading, and the rule to choose the fastest, most reliable, most scalable method.
- [4] Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv, submitted 9 June 2023, NeurIPS 2023 Datasets and Benchmarks Track): 80 questions; 85 percent agreement between GPT-4 and humans and 81 percent between humans on non-tie votes; GPT-4 consistent in 65.0 percent of cases when two answers swap order.
- [5] Hamel Husain and Shreya Shankar, AI Evals: Everything You Need to Know (page dated 18 September 2026): at least 100 traces; the limit of synthetic data; a purpose-built set often grows to 100 or more examples; binary labels; independent labelling before discussion; 100 or more fresh traces per review cycle of 2 to 4 weeks.
- [6] Shankar et al., Who Validates the Validators? (arXiv, submitted 18 April 2024): criteria drift, found in a qualitative study with nine industry practitioners.
- [7] Anthropic, Demystifying evals for AI agents (engineering blog, 9 January 2026): 20 to 50 simple tasks drawn from real failures as a starting set; test both the cases where a behaviour should occur and where it should not.
- [8] Evan Miller, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arXiv, submitted 1 November 2024): the standard error for pass or fail scores (Equation 2), the 95 percent interval of 1.96 standard errors (Equation 3), paired differences, resampling answers, and the 969-question example with parameters the paper calls fictional.
- [9] Papineni, Roukos, Ward and Zhu, BLEU: a Method for Automatic Evaluation of Machine Translation (ACL, July 2002): BLEU scores a machine translation against human reference translations by counting matching word sequences.
- [10] Ehud Reiter, A Structured Review of the Validity of BLEU (Computational Linguistics, September 2018, abstract read): 284 correlations in 34 papers; the evidence does not support using BLEU outside machine translation, for individual texts, or for scientific hypothesis testing.
- [11] Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv, submitted 22 May 2020, NeurIPS 2020): the paper that named retrieval-augmented generation.
- [12] Es, James, Espinosa-Anke and Schockaert, Ragas: Automated Evaluation of Retrieval Augmented Generation (arXiv, submitted 26 September 2023, EACL 2024 System Demonstrations): the three measures, the faithfulness formula, the WikiEval set built from 50 Wikipedia pages, agreement with human labels of 0.95, 0.78 and 0.70, and context relevance as the hardest to evaluate.
- [13] Mary L. McHugh, Interrater reliability: the kappa statistic (Biochemia Medica, 15 October 2012): the kappa formula, and a cited study with 94.2 percent agreement and kappa of 0.555 on the same data.
- [14] Khaled El Emam, Benchmarking Kappa for Software Process Assessment Reliability Studies (ISERN-98-02, 1998): a secondary source that reproduces the bands proposed by Landis and Koch in 1977 and records that those authors called their benchmark arbitrary. The 1977 paper's own table was not read for this guide.
- [15] LangChain, Evaluation concepts (LangSmith docs, undated, read 30 September 2026): the definitions of offline and online evaluation and how failures found online become offline test cases.
Related reading
What a 90 percent pass rate on 50 test cases tells you
A pass rate of 90 percent measured on 50 test cases fits a true rate anywhere from 81.7 to 98.3 percent, so a result of 92 percent on the same 50 cases cannot be told apart from it. This post shows the arithmetic step by step and lists what to report next to every pass rate.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.