Why AI evals matter: the test that shows whether an AI product works / Decide with evals
How to read an AI eval report if you are not an engineer
A pass rate means little until you know what it was measured on. A result of 92 percent on 50 cases has a 95 percent confidence interval of 84.5% to 99.5%. The same 92 percent on 1,000 cases has an interval of 90.3% to 93.7%. To read an AI eval report, check six things: the number of cases and where they came from, the range around the pass rate, the pass rate for each group of cases, the list of failures, who or what did the grading and how many times each case ran, and whether the same set was run on the previous version. This page shows the arithmetic and ends with a checklist to print.
Published September 30, 2026. Editorial.
Key takeaways
- A result of 46 passes out of 50 cases has a 95% range of 84.5% to 99.5%. A result of 920 out of 1,000 has a range of 90.3% to 93.7%. Both are reported as 92 percent.
- The formula is pass rate plus and minus 1.96 times the standard error, where the standard error is the square root of pass rate times (1 minus pass rate) divided by the number of cases.
- An average can hide a failing group. In an invented illustration, 95 percent on 900 ordinary questions and 65 percent on 100 refund questions still add up to 92 percent overall.
- A fall from 92 to 90 percent on 200 cases is 4 cases and is within the range that chance produces. Ask which cases changed from pass to fail.
- No single number is a good pass rate. Ask for the list of failed cases and judge them by what each failure would cost.
A report arrives before a release meeting. Its first line says the support assistant passed 92 percent of its eval cases. That 92 percent can come from 46 passes out of 50 cases or from 920 passes out of 1,000, and those two results support different decisions.
This page teaches a reader who is not an engineer to read such a report. It gives six things to look for, the arithmetic behind the first of them, and a checklist to print. Every report figure on this page is an invented illustration unless a source is cited beside it. An eval is a repeatable test: a set of inputs, a written expectation for each one, and a way to score the output. What an AI eval is has the full definition.
Start with what the pass rate was measured on
A pass rate is the number of cases passed, divided by the number of cases run. Before you read the percentage, find four facts in the report:
- how many cases were run
- where the cases came from: real user questions, or questions the team wrote
- the date of the run
- the exact model version
NIST, the United States standards body, asks for this in its AI Risk Management Framework: "Test sets, metrics, and details about the tools used during TEVV are documented" [7]. TEVV is its short form for test, evaluation, verification and validation. A report without these four facts cannot be checked by anyone.
Why 92 percent of 50 cases is weaker evidence than 92 percent of 1,000
The cases in an eval are a sample of all the questions users could send. A different sample would give a different pass rate. The confidence interval says how different. In plain words, it is the range that is likely to contain the true pass rate.
The formula uses the standard error, a number that says how far the pass rate could move if a different set of cases had been used. It has two steps, and both are given in a 2024 paper by Evan Miller [1]:
- Standard error = the square root of (pass rate × (1 minus pass rate) ÷ number of cases).
- 95% confidence interval = pass rate plus and minus 1.96 × standard error [1] [2].
NIST's statistics handbook calls this form "the confidence expression most frequently used" [3]. Here is the arithmetic for 46 passes out of 50:
- 0.92 × 0.08 = 0.0736
- 0.0736 ÷ 50 = 0.001472
- The square root of 0.001472 is 0.03837
- 1.96 × 0.03837 = 0.0752, which is 7.5 percentage points
- 92.0 minus 7.5 is 84.5, and 92.0 plus 7.5 is 99.5
| Cases | Passed | Standard error | 95% range |
|---|---|---|---|
| 50 | 46 | 3.8 points | 84.5% to 99.5% |
| 200 | 184 | 1.9 points | 88.2% to 95.8% |
| 1,000 | 920 | 0.9 points | 90.3% to 93.7% |
The range for 50 cases is 15.0 points wide. The range for 1,000 cases is 3.4 points wide. With 50 cases the true rate could be 85 percent or 99 percent, and the report cannot say which. Four times the cases cuts the range in half: from 50 to 200 cases the margin goes from 7.5 points to 3.8.
Two cautions. NIST's handbook says this simple form "may not be accurate enough" when the sample or the number of failures is small, and it presents a second method, named after Wilson, first [3]. For 46 passes out of 50 the Wilson method gives 81.2% to 96.8%. Both methods lead to the same reading: 50 cases leave a wide range.
The second caution is about related cases. Anthropic reports from its own work that when questions come in related groups, the correct standard error on popular evals can be more than three times the simple one [2]. Fifty rewordings of five questions give less evidence than 50 separate cases.
How many cases are enough depends on the width you can accept. Rearranging the same formula for a 92 percent pass rate: a margin of 5 points needs 3.8416 × 0.0736 ÷ 0.0025 = 114 cases, and a margin of 2 points needs 3.8416 × 0.0736 ÷ 0.0004 = 707 cases, both rounded up. The method is in how many test cases an LLM eval needs, where LLM means large language model.
A lower score than last time may be no change
Suppose last month's report said 92 percent on 200 cases and this month's says 90 percent on the same 200. The first range is 88.2% to 95.8%. The second, computed the same way, is 85.8% to 94.2%. The two ranges share everything from 88.2% to 94.2%, and the difference is 4 cases out of 200. A fall of this size is within what chance produces on a set this small.
Miller's paper recommends comparing two runs question by question instead of comparing the two averages [1]. So ask one thing: which cases passed last time and fail now? Four new failures spread across topics suggest chance. Four new failures that all concern refunds are a finding, whatever the average did. What happens when the AI model changes covers one cause of a real fall.
An average can hide a failing group
The best known measured case comes from face analysis software. A 2018 study of three commercial systems found that "darker-skinned females are the most misclassified group (with error rates of up to 34.7%)", while "The maximum error rate for lighter-skinned males is 0.8%" [4]. Two existing test sets the authors examined were 79.6% and 86.2% lighter-skinned subjects [4], so a strong overall score on them gave little information about the smaller group. A 2019 paper on medical imaging describes the same pattern: "overall performance of a cancer detection model may be high, but the model still consistently misses a rare but aggressive cancer subtype" [5]. Both studies concern image systems. The arithmetic applies to any average.
Here is an invented illustration. A support assistant is run on 1,000 cases. It passes 855 of 900 ordinary questions (95 percent) and 65 of 100 refund questions (65 percent). The total is 920 of 1,000, which is 92 percent. The report's first line is the same as before, and 35 of every 100 refund answers are wrong. Ask for the pass rate per group, and for the number of cases in each group.
The failures decide what a good pass rate is
No single number is a good pass rate. NIST's framework "does not prescribe risk tolerance" [7], and the right figure depends on what the failures are. Ask for the list of failed cases, with the input and the output of each.
Then sort them. A clumsy sentence is one type of failure. A wrong refund rule or an invented source is another, and public AI failures and the tests that target each one shows what that type has cost. One failure of the second type can stop a release at a 99 percent pass rate. Using evals to decide a release explains how to set the pass line before the run.
Who or what did the grading, and how many times each case ran
A grader is the program or person that scores an output. Anthropic's engineering blog lists three types: "code-based, model-based, and human" [6]. When a model does the grading, ask how often it agrees with a person on a sample that people labelled first. Human review and reviewer agreement gives the method, and never let the model grade its own work explains why the grader must be separate from the product.
Then ask how many times each case ran. The same post says: "Because model outputs vary between runs, we run multiple trials to produce more consistent results" [6]. A report should state whether "pass" means the case passed once or passed every time. The two figures differ. In Anthropic's example, a product that succeeds on 75 percent of single runs passes all three of three runs 42 percent of the time, because 0.75 × 0.75 × 0.75 = 0.421875, which rounds to 42 percent [6]. Agent reliability across repeated runs covers this in detail.
Whether the set was run before, and whether the product was adjusted to fit it
Ask two things about the history of the set.
First, was the same set run on the previous version? Only then can two reports be compared.
Second, did the team adjust the product while looking at these same cases? If so, the score on them is higher than the score on new questions will be. The protection is a held-out set: a group of cases kept away from the people adjusting the product and used only for the final score. A 2024 paper by Kapoor and four co-authors reported on public tests of AI agents, which are AI programs that take steps on their own: "many agent benchmarks have inadequate holdout sets, and sometimes none at all" [8]. A benchmark is a public test set. Ask whether a held-out set exists, how many cases it has, and what the product scored on it. When evals give false confidence lists other ways a passing score can mislead.
Each number on the report and what to ask next
| Number on the report | What it tells you | What to ask next |
|---|---|---|
| Pass rate | The share of cases passed | On how many cases, and where did they come from? |
| Number of cases | How wide the range is | Are the cases separate, or rewordings of a few? |
| 95% range | The range likely to contain the true rate | Does the low end still meet the pass line? |
| Pass rate per group | Which type of question fails | How many cases are in the smallest group? |
| Runs per case | Whether a pass means once or every time | What is the rate when every run must pass? |
| Grader agreement with people | How far to trust the grader | On how many labelled cases was it measured? |
| Change since the last run | The direction of movement | Which cases changed from pass to fail? |
A checklist to print
- How many cases were run, and on what date?
- Which exact model version and which prompt (the written instructions given to the model) were tested?
- Where did the cases come from: real users or the team?
- What is the 95% range around the pass rate?
- Are the cases separate, or several wordings of the same question?
- What is the pass rate for each group of cases?
- Which cases failed, and what did the product answer?
- Which failures would cost money or cause harm?
- Who or what graded the outputs, and how was the grader checked?
- How many times was each case run, and does "pass" mean every time?
- Was this same set run on the previous version, and which cases changed?
- Is there a held-out set, and what was the score on it?
The same questions work on a supplier's report. How to evaluate a vendor's eval suite gives more detail, and why AI evals matter gives the wider argument.
How Reveneau reports eval results
At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. An eval report from Reveneau states the items on this page's checklist: the number of cases and where they came from, the 95% range beside every pass rate, the pass rate per group, the full list of failed cases, the grader, the number of runs per case, and the model version. We write the pass line before the run.
The judged checks in our suite are graded with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We check that grader against human labels, and the report says so.
Reveneau, as a company, takes responsibility for the whole project through production and after release. See AI development at Reveneau or contact us.
Best for
- A product owner or executive who must approve a release from an eval report
- A buyer who has received an eval report from a supplier
- Anyone who wants to check a reported pass rate with a calculator
Avoid if
- You are designing the eval set yourself: use the LLM evals guide for the method
- You need to compare two models statistically: the paper by Evan Miller cited here gives the paired method
Check before you decide
- Recompute one 95% range from the report's own pass rate and case count
- Ask for the file of failed cases and read ten of them
- Confirm the model version in the report matches the version in production
Common questions
What is a good pass rate for an AI product?
The right pass rate for an AI product depends on what the failures are and what each one would cost, so no single figure is correct. NIST's AI Risk Management Framework says it does not prescribe risk tolerance, so the organisation sets the line. Ask for the list of failed cases: one invented source or one wrong refund rule can stop a release at a pass rate of 99 percent.
How many test cases are enough to trust an eval result?
The number of test cases you need depends on how wide a range you can accept around the pass rate. Using the standard formula for a 92 percent pass rate, a margin of 5 percentage points needs 114 cases and a margin of 2 points needs 707 cases. With 50 cases the 95 percent range runs from 84.5% to 99.5%, a width of 15.0 points.
What is a confidence interval in plain words?
A confidence interval is the range that is likely to contain the true pass rate, given that the test cases are only a sample of the questions users could send. For a 95 percent interval, take the pass rate and add and subtract 1.96 times the standard error, as Evan Miller's 2024 paper and Anthropic's summary of it state. A result of 920 passes out of 1,000 gives 90.3% to 93.7%.
What should an AI eval report contain?
An AI eval report should contain the number of cases and where they came from, the date and the exact model version, the 95 percent range beside the pass rate, the pass rate for each group of cases, the list of failed cases, the grader and how it was checked, and the number of runs per case. NIST's framework asks that test sets, metrics and tools are documented.
What does it mean when an eval score went down?
A lower eval score can be a real fall or chance, and the size of the set tells you which is more plausible. A fall from 92 to 90 percent on 200 cases is 4 cases, and the two 95 percent ranges, 88.2% to 95.8% and 85.8% to 94.2%, overlap. Ask which cases passed last time and fail now. New failures on one topic are a finding.
Why is 92 percent on 50 cases weaker evidence than 92 percent on 1,000 cases?
A pass rate of 92 percent on 50 cases is weaker because a small sample leaves a wide range for the true rate. The standard error for 50 cases is 3.8 percentage points, which gives a 95 percent range of 84.5% to 99.5%. For 1,000 cases the standard error is 0.9 points and the range is 90.3% to 93.7%. Four times the cases cuts the range in half.
How can a high average pass rate hide a serious problem?
A high average pass rate can hide a group of cases that fails often, because a large group that passes makes up most of the average. The 2018 Gender Shades study of three commercial face analysis systems found error rates of up to 34.7% for darker-skinned females and at most 0.8% for lighter-skinned males. Ask for the pass rate per group and the number of cases in each.
Who or what grades the answers in an AI eval?
Answers in an AI eval are graded by code, by a model, or by a person. Anthropic's engineering blog names those three grader types. Code checks exact things, such as whether a cited source exists. A model or a person judges things with no exact answer, such as tone. When a model grades, ask how often it agrees with people on a sample that people labelled first.
How much work is it to check the numbers in an eval report myself?
Checking one range in an eval report takes five steps on a calculator. Multiply the pass rate by 1 minus the pass rate, divide by the number of cases, take the square root, multiply by 1.96, then add and subtract the result from the pass rate. For 46 passes out of 50 the steps give 0.0736, 0.001472, 0.03837 and 0.0752, so the range is 84.5% to 99.5%.
What is a held-out set and why should I ask about one?
A held-out set is a group of test cases kept away from the people adjusting the product and used only for the final score. Ask about one because a product adjusted while looking at its test cases scores higher on them than on new questions. A 2024 paper by Kapoor and co-authors reported that many public agent benchmarks have inadequate holdout sets, and sometimes none at all.
Does a pass in an eval report mean the product got the case right every time?
A pass can mean the case passed once or passed on every run, and the report should say which. The two figures differ. In an example from Anthropic's engineering blog, a product that succeeds on 75 percent of single runs passes all three of three runs 42 percent of the time, because 0.75 × 0.75 × 0.75 is 0.421875, which rounds to 42 percent. Ask how many times each case ran.
Do these questions apply when I buy an AI product from a supplier?
Yes. The same twelve checklist questions apply to a supplier's eval report, and a buyer has more reason to ask them because the buyer did not see the tests being written. Ask the supplier for the number of cases, the 95 percent range, the pass rate per group, the failed cases and the model version. A supplier with a real eval set can answer all five.
References
- [1] Evan Miller, Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arXiv preprint, 1 November 2024): the standard error formula for pass or fail scores (Equation 2), the 95% interval as the mean plus and minus 1.96 standard errors (Equation 3), and the recommendation to compare two runs on question-level paired differences.
- [2] Anthropic, A statistical approach to model evaluations (19 November 2024): the 95% interval is the mean score plus and minus 1.96 times the standard error; Anthropic's own finding that clustered standard errors on popular evals can be over three times the simple ones.
- [3] NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.1, Confidence intervals (read 30 September 2026): the simple interval for a proportion, called the confidence expression most frequently used; the caution for small samples; the Wilson method.
- [4] Buolamwini and Gebru, Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification (Proceedings of Machine Learning Research, volume 81, 2018): error rates of up to 34.7% for darker-skinned females and a maximum of 0.8% for lighter-skinned males across 3 commercial systems; test sets that were 79.6% and 86.2% lighter-skinned subjects.
- [5] Oakden-Rayner, Dunnmon, Carneiro and Re, Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging (arXiv preprint, 27 September 2019): a model with a high overall score can still consistently miss a rare subtype.
- [6] Anthropic, Demystifying evals for AI agents (9 January 2026): three types of grader; multiple trials because outputs vary between runs; a 75% per-trial success rate gives 42% for passing all three of three trials.
- [7] NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023): MEASURE 2.1 on documenting test sets, metrics and tools; the statement that the framework does not prescribe risk tolerance.
- [8] Kapoor, Stroebl, Siegel, Nadgir and Narayanan, AI Agents That Matter (arXiv preprint, 1 July 2024): many agent benchmarks have inadequate holdout sets, and sometimes none at all.
Related reading
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
What to do with a grader that gives no reason
A Jev grade is a probability with no paragraph. That is enough for a check you wrote and calibrated yourself, and wrong for a new kind of failure, a design judgment, or anything a regulator wants explained in words.