What a 90 percent pass rate on 50 test cases tells you

Start with an invented example. A team builds a support assistant for an invented furniture shop. The assistant runs on a large language model (LLM), and the team tests it with an eval: a fixed set of test cases, each one a customer question with a written statement of what a correct answer must do. The set has 50 cases. The current prompt, the written instruction given to the model, passes 45 of them. A new prompt passes 46. The weekly report says the pass rate rose from 90 to 92 percent, and a product manager is asked to approve the new prompt because of that sentence.
The sentence leaves out two facts. The first is the number of cases. The second is the range around the result. By the formula in Evan Miller's 2024 paper "Adding Error Bars to Evals", a pass rate of 90 percent on 50 cases fits a true pass rate anywhere from 81.7 to 98.3 percent. The new prompt's 92 percent is inside that range. A result of 85 percent would be inside it too, and so would 97.
Our position is that a pass rate should always be reported with its count of cases and its range. This post works through the arithmetic for the furniture shop's 50 cases. The table for larger sets, and each formula as its source prints it, are in our guide on how many test cases an LLM eval needs.
The 50 cases are a sample of every request the product will receive
The shop's customers will send thousands of different questions. The 50 cases represent all of them. A different 50 from the same customers would hold a different mix of easy and difficult questions, and the same prompt would pass a different number. The 90 percent is an estimate of the true pass rate, which is the rate the prompt would reach across every request.
Two terms from statistics describe how far such an estimate can be from the truth.
The standard error is the typical distance between a measured pass rate and the true one. Miller's paper gives it for scores that are pass or fail: the square root of p × (1 - p) / n, where p is the pass rate written as a fraction and n is the number of cases.
A confidence interval is the range of true rates that fits the measurement. The paper builds the 95 percent version by adding and subtracting 1.96 standard errors, and Anthropic's summary of the paper states the rule in words: "A 95% confidence interval can be calculated from the SEM by adding and subtracting 1.96 × SEM from the mean score." SEM is short for standard error of the mean. The 95 percent describes the method: across many repeated tests, ranges built this way contain the true rate 95 times in 100.
The statistics handbook of the US National Institute of Standards and Technology (NIST) prints the same interval with a symbol in place of 1.96, and calls it "the confidence expression most frequently used".
The working for 90 percent on 50 cases
- p × (1 - p) = 0.90 × 0.10 = 0.09
- 0.09 / 50 = 0.0018
- The square root of 0.0018 is 0.042426, so the standard error is 4.24 points.
- 1.96 × 0.042426 = 0.08316, a margin of 8.3 points.
- 90 - 8.3 = 81.7 and 90 + 8.3 = 98.3.
We computed every figure in this post with the Python programming language. The same five steps for the new prompt, 46 passes out of 50, give a standard error of 3.84 points, a margin of 7.5 points, and a range of 84.5 to 99.5 percent.
| Result | Margin | Simple interval | Wilson interval |
|---|---|---|---|
| 45 of 50 (90 percent) | 8.3 points | 81.7 to 98.3 | 78.6 to 95.7 |
| 46 of 50 (92 percent) | 7.5 points | 84.5 to 99.5 | 81.2 to 96.8 |
| 90 of 100 (90 percent) | 5.9 points | 84.1 to 95.9 | 82.6 to 94.5 |
| 900 of 1,000 (90 percent) | 1.9 points | 88.1 to 91.9 | 88.0 to 91.7 |
The last column uses the Wilson method, a second formula on the same NIST page. NIST warns that when the sample or the number of failures is small, limits from the simple formula "may not be accurate enough for some applications". It names two advantages of the Wilson method: "its worth does not strongly depend upon the value of n and/or p", and "the lower limit cannot be negative". For a set of 50 cases with 5 failures, we report the Wilson range, which is 78.6 to 95.7 percent.
Under either formula, the range for 45 passes out of 50 is more than 16 points wide.
Why 90 percent and 92 percent on 50 cases cannot be told apart
On 50 cases, each case is worth 2 points. The whole difference between the two prompts is one case. The two simple ranges, 81.7 to 98.3 and 84.5 to 99.5, overlap from 84.5 to 98.3.
By the sample size formula in a second section of the NIST handbook, finding a real move from 90 to 92 percent four times in five takes 1,666 cases.
Miller's paper offers a better comparison than two separate pass rates. Its fourth recommendation is to analyse "the question-level paired differences": run both versions on the same cases and take the difference on each case. To continue the invented example, suppose the new prompt passes 3 cases that the old one failed and fails 2 cases that the old one passed. Each fixed case counts as plus 1, each broken case as minus 1, and the other 45 count as 0.
- Mean difference: (3 - 2) / 50 = 0.02, which is 2 points.
- Variance, a measure of how spread out the 50 differences are: (5 - 50 × 0.0004) / 49 = 0.10163.
- Standard error: the square root of (0.10163 / 50) = 0.04508.
- Margin: 1.96 × 0.04508 = 0.0884, which is 8.8 points.
- The 95 percent interval on the difference runs from minus 6.8 to plus 10.8 points.
That interval includes zero. The result fits a new prompt that is 10.8 points better, and it also fits one that is 6.8 points worse.
The paired run also produces the list of the 5 cases that changed, and reading 5 cases takes minutes. If the 2 newly broken cases are refund questions, the product manager now knows what the new prompt costs, and the 2 point gain alone could never have shown it.
What 50 cases can show
A set of 50 cases is a reasonable place to begin. Anthropic's engineering team writes that "20-50 simple tasks drawn from real failures is a great start", because early changes to a product have large effects and "small sample sizes suffice". That is the company's own advice.
The arithmetic agrees. Suppose a change makes the furniture shop's assistant pass 30 cases out of 50, a rate of 60 percent. The same formula gives a margin of 13.6 points and a range of 46.4 to 73.6 percent. The top of that range is 8.1 points below the bottom of the range for 45 of 50. A set of 50 cases detects a 30 point fall. A 2 point gain needs a larger set, or a reading of the cases that changed.
Every range in this post depends on two conditions. The first is that Miller's formula treats the cases as a random draw from all possible requests. If the team chose its 50 cases by hand because they were difficult, the range describes difficult cases only, and the customers' true pass rate is a separate question. Our guide on how to build an LLM eval dataset explains how to keep a random sample and a hand-chosen group as two tagged groups, each with its own pass rate.
The second condition is that someone has read the failures. Hamel Husain writes that "your pass rate is a product decision, depending on the failures you are willing to tolerate". Five failed cases about delivery dates and five failed cases that promise a refund the shop does not offer produce the same 90 percent. Reading and naming failures comes first, as the guide on error analysis before metrics sets out.
What to report with every pass rate
Miller's paper gives the instruction in one sentence: "We suggest reporting the standard error of the mean alongside (beneath) the mean when reporting eval scores." For the furniture shop's weekly report, the invented line would read:
New prompt: 46 of 50 passed, 92 percent (95 percent Wilson interval 81.2 to 96.8). Old prompt: 45 of 50, 90 percent (78.6 to 95.7). Same 50 cases, dataset version 7, one run per case. Paired difference: plus 2 points (minus 6.8 to plus 10.8). Cases that changed: 3 fixed, 2 broken, listed below.
That line contains five things:
- the count of cases and the count that passed
- the 95 percent interval, by the Wilson method when the set is small or failures are few
- how many times each case was run, because an LLM can give a different answer to the same question. Anthropic's summary recommends "resampling answers from the same model several times, and using the question-level averages as the question scores"
- the version of the dataset, so both results are known to come from the same cases
- for a comparison, the paired interval and the list of cases that changed
A narrower range takes more cases. Eugene Yan states the rule: "to reduce the margin of error by half, we need to quadruple the sample size". At 90 percent the margin is 8.3 points on 50 cases and 4.2 points on 200. A margin of 3 points takes 385 cases, and a margin of 2 points takes 865. More cases and more runs per case also mean more grading, and our post on the real price of an LLM judge covers how that cost multiplies.
A reader who receives a pass rate with nothing next to it should ask two questions before deciding anything: how many cases, and what is the range? The guide on how to read an AI eval report lists the rest, and the full method sits in our guide to LLM evals. For software written by AI, where the unit being counted is a test instead of a graded answer, see how many tests AI-generated code needs.
Where Reveneau fits
All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an AI feature, the judged part of that suite is a set of cases like the furniture shop's. We report each pass rate with its count of cases and its 95 percent interval. When two versions are compared, we run both on the same cases and list the cases that changed. We grade the judged checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader, which shortens the wait when a set grows or each case runs several times. Reveneau, as a company, takes responsibility for the whole project through production and after release. To plan an eval set for your product, contact Reveneau.
Sources
- Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations: Evan Miller, arXiv, 1 November 2024. The standard error for pass or fail scores, the 95 percent interval with 1.96, paired differences, and the advice to report the standard error.
- A statistical approach to model evaluations: Anthropic, 19 November 2024. The interval rule in words, and answering each question several times.
- NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.1, Confidence intervals: US National Institute of Standards and Technology, undated, read 30 September 2026. The simple interval, its limits on small samples, and the Wilson method.
- NIST/SEMATECH e-Handbook of Statistical Methods, section 7.2.4.2, Sample sizes required: US National Institute of Standards and Technology, undated, read 30 September 2026. The sample size formula behind the 1,666 figure.
- Demystifying evals for AI agents: Anthropic engineering blog, 9 January 2026. The advice to start with 20 to 50 tasks drawn from real failures.
- Your AI Product Needs Evals: Hamel Husain, 29 March 2024. The pass rate as a product decision.
- Product Evals in Three Simple Steps: Eugene Yan, November 2025. Halving the margin takes four times the sample.


