Engineering

What a 90 percent pass rate on 50 test cases tells you

Editorial · Reveneau · October 2, 2026

What a 90 percent pass rate on 50 test cases tells you

Start with an invented example. A team builds a support assistant for an invented furniture shop. The assistant runs on a large language model (LLM), and the team tests it with an eval: a fixed set of test cases, each one a customer question with a written statement of what a correct answer must do. The set has 50 cases. The current prompt, the written instruction given to the model, passes 45 of them. A new prompt passes 46. The weekly report says the pass rate rose from 90 to 92 percent, and a product manager is asked to approve the new prompt because of that sentence.

The sentence leaves out two facts. The first is the number of cases. The second is the range around the result. By the formula in Evan Miller's 2024 paper "Adding Error Bars to Evals", a pass rate of 90 percent on 50 cases fits a true pass rate anywhere from 81.7 to 98.3 percent. The new prompt's 92 percent is inside that range. A result of 85 percent would be inside it too, and so would 97.

Our position is that a pass rate should always be reported with its count of cases and its range. This post works through the arithmetic for the furniture shop's 50 cases. The table for larger sets, and each formula as its source prints it, are in our guide on how many test cases an LLM eval needs.

The 50 cases are a sample of every request the product will receive

The shop's customers will send thousands of different questions. The 50 cases represent all of them. A different 50 from the same customers would hold a different mix of easy and difficult questions, and the same prompt would pass a different number. The 90 percent is an estimate of the true pass rate, which is the rate the prompt would reach across every request.

Two terms from statistics describe how far such an estimate can be from the truth.

The standard error is the typical distance between a measured pass rate and the true one. Miller's paper gives it for scores that are pass or fail: the square root of p × (1 - p) / n, where p is the pass rate written as a fraction and n is the number of cases.

A confidence interval is the range of true rates that fits the measurement. The paper builds the 95 percent version by adding and subtracting 1.96 standard errors, and Anthropic's summary of the paper states the rule in words: "A 95% confidence interval can be calculated from the SEM by adding and subtracting 1.96 × SEM from the mean score." SEM is short for standard error of the mean. The 95 percent describes the method: across many repeated tests, ranges built this way contain the true rate 95 times in 100.

The statistics handbook of the US National Institute of Standards and Technology (NIST) prints the same interval with a symbol in place of 1.96, and calls it "the confidence expression most frequently used".

The working for 90 percent on 50 cases

  1. p × (1 - p) = 0.90 × 0.10 = 0.09
  2. 0.09 / 50 = 0.0018
  3. The square root of 0.0018 is 0.042426, so the standard error is 4.24 points.
  4. 1.96 × 0.042426 = 0.08316, a margin of 8.3 points.
  5. 90 - 8.3 = 81.7 and 90 + 8.3 = 98.3.

We computed every figure in this post with the Python programming language. The same five steps for the new prompt, 46 passes out of 50, give a standard error of 3.84 points, a margin of 7.5 points, and a range of 84.5 to 99.5 percent.

Result Margin Simple interval Wilson interval
45 of 50 (90 percent) 8.3 points 81.7 to 98.3 78.6 to 95.7
46 of 50 (92 percent) 7.5 points 84.5 to 99.5 81.2 to 96.8
90 of 100 (90 percent) 5.9 points 84.1 to 95.9 82.6 to 94.5
900 of 1,000 (90 percent) 1.9 points 88.1 to 91.9 88.0 to 91.7

The last column uses the Wilson method, a second formula on the same NIST page. NIST warns that when the sample or the number of failures is small, limits from the simple formula "may not be accurate enough for some applications". It names two advantages of the Wilson method: "its worth does not strongly depend upon the value of n and/or p", and "the lower limit cannot be negative". For a set of 50 cases with 5 failures, we report the Wilson range, which is 78.6 to 95.7 percent.

Under either formula, the range for 45 passes out of 50 is more than 16 points wide.

Why 90 percent and 92 percent on 50 cases cannot be told apart

On 50 cases, each case is worth 2 points. The whole difference between the two prompts is one case. The two simple ranges, 81.7 to 98.3 and 84.5 to 99.5, overlap from 84.5 to 98.3.

By the sample size formula in a second section of the NIST handbook, finding a real move from 90 to 92 percent four times in five takes 1,666 cases.

Miller's paper offers a better comparison than two separate pass rates. Its fourth recommendation is to analyse "the question-level paired differences": run both versions on the same cases and take the difference on each case. To continue the invented example, suppose the new prompt passes 3 cases that the old one failed and fails 2 cases that the old one passed. Each fixed case counts as plus 1, each broken case as minus 1, and the other 45 count as 0.

  1. Mean difference: (3 - 2) / 50 = 0.02, which is 2 points.
  2. Variance, a measure of how spread out the 50 differences are: (5 - 50 × 0.0004) / 49 = 0.10163.
  3. Standard error: the square root of (0.10163 / 50) = 0.04508.
  4. Margin: 1.96 × 0.04508 = 0.0884, which is 8.8 points.
  5. The 95 percent interval on the difference runs from minus 6.8 to plus 10.8 points.

That interval includes zero. The result fits a new prompt that is 10.8 points better, and it also fits one that is 6.8 points worse.

The paired run also produces the list of the 5 cases that changed, and reading 5 cases takes minutes. If the 2 newly broken cases are refund questions, the product manager now knows what the new prompt costs, and the 2 point gain alone could never have shown it.

What 50 cases can show

A set of 50 cases is a reasonable place to begin. Anthropic's engineering team writes that "20-50 simple tasks drawn from real failures is a great start", because early changes to a product have large effects and "small sample sizes suffice". That is the company's own advice.

The arithmetic agrees. Suppose a change makes the furniture shop's assistant pass 30 cases out of 50, a rate of 60 percent. The same formula gives a margin of 13.6 points and a range of 46.4 to 73.6 percent. The top of that range is 8.1 points below the bottom of the range for 45 of 50. A set of 50 cases detects a 30 point fall. A 2 point gain needs a larger set, or a reading of the cases that changed.

Every range in this post depends on two conditions. The first is that Miller's formula treats the cases as a random draw from all possible requests. If the team chose its 50 cases by hand because they were difficult, the range describes difficult cases only, and the customers' true pass rate is a separate question. Our guide on how to build an LLM eval dataset explains how to keep a random sample and a hand-chosen group as two tagged groups, each with its own pass rate.

The second condition is that someone has read the failures. Hamel Husain writes that "your pass rate is a product decision, depending on the failures you are willing to tolerate". Five failed cases about delivery dates and five failed cases that promise a refund the shop does not offer produce the same 90 percent. Reading and naming failures comes first, as the guide on error analysis before metrics sets out.

What to report with every pass rate

Miller's paper gives the instruction in one sentence: "We suggest reporting the standard error of the mean alongside (beneath) the mean when reporting eval scores." For the furniture shop's weekly report, the invented line would read:

New prompt: 46 of 50 passed, 92 percent (95 percent Wilson interval 81.2 to 96.8). Old prompt: 45 of 50, 90 percent (78.6 to 95.7). Same 50 cases, dataset version 7, one run per case. Paired difference: plus 2 points (minus 6.8 to plus 10.8). Cases that changed: 3 fixed, 2 broken, listed below.

That line contains five things:

  • the count of cases and the count that passed
  • the 95 percent interval, by the Wilson method when the set is small or failures are few
  • how many times each case was run, because an LLM can give a different answer to the same question. Anthropic's summary recommends "resampling answers from the same model several times, and using the question-level averages as the question scores"
  • the version of the dataset, so both results are known to come from the same cases
  • for a comparison, the paired interval and the list of cases that changed

A narrower range takes more cases. Eugene Yan states the rule: "to reduce the margin of error by half, we need to quadruple the sample size". At 90 percent the margin is 8.3 points on 50 cases and 4.2 points on 200. A margin of 3 points takes 385 cases, and a margin of 2 points takes 865. More cases and more runs per case also mean more grading, and our post on the real price of an LLM judge covers how that cost multiplies.

A reader who receives a pass rate with nothing next to it should ask two questions before deciding anything: how many cases, and what is the range? The guide on how to read an AI eval report lists the rest, and the full method sits in our guide to LLM evals. For software written by AI, where the unit being counted is a test instead of a graded answer, see how many tests AI-generated code needs.

Where Reveneau fits

All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an AI feature, the judged part of that suite is a set of cases like the furniture shop's. We report each pass rate with its count of cases and its 95 percent interval. When two versions are compared, we run both on the same cases and list the cases that changed. We grade the judged checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader, which shortens the wait when a set grows or each case runs several times. Reveneau, as a company, takes responsibility for the whole project through production and after release. To plan an eval set for your product, contact Reveneau.

Sources

Common questions

What does a 90 percent pass rate on 50 test cases tell you?

A 90 percent pass rate on 50 test cases tells you that the true pass rate lies somewhere between 81.7 and 98.3 percent, using the simple 95 percent interval in Evan Miller's 2024 paper on eval statistics. The Wilson method in the NIST statistics handbook gives 78.6 to 95.7 percent for the same result. Fifty cases can show a large failure. Stating the rate to within 3 points takes 385 cases.

Can I tell 90 percent from 92 percent on 50 test cases?

No, 90 percent and 92 percent on 50 test cases cannot be told apart. The first is 45 passes and the second is 46, a difference of one case. Their simple 95 percent intervals are 81.7 to 98.3 percent and 84.5 to 99.5 percent, which overlap from 84.5 to 98.3. By the sample size formula in the NIST handbook, detecting a real move from 90 to 92 percent with an 80 percent chance takes 1,666 cases.

What is a standard error in everyday words?

A standard error is the typical distance between a pass rate measured on a sample of test cases and the true pass rate across every request. For pass or fail scores, Evan Miller's paper gives it as the square root of p × (1 - p) / n, where p is the pass rate and n is the number of cases. At 90 percent on 50 cases, the standard error is 4.24 points.

What is a 95 percent confidence interval on a pass rate?

A 95 percent confidence interval on a pass rate is the range of true rates that fits the measured result. Evan Miller's paper builds it by adding and subtracting 1.96 standard errors from the measured rate. The 95 percent describes the method: across many repeated tests, ranges built this way contain the true rate 95 times in 100. For 45 passes out of 50, the range is 81.7 to 98.3 percent.

Where does the value 1.96 in the interval come from?

The value 1.96 is the multiplier that Evan Miller's paper prints for a 95 percent interval, and Anthropic's summary of the paper states the same rule: add and subtract 1.96 standard errors from the mean score. The NIST statistics handbook writes the same interval with a symbol in place of the number. Multiply the standard error by 1.96 to get the margin, which is 8.3 points for 90 percent on 50 cases.

How do I compare two prompt versions on a small eval set?

Compare two prompt versions by running both on the same cases and looking at the difference on each case, which Evan Miller's paper recommends over comparing two summary scores. In this post's invented example, a new prompt fixes 3 of 50 cases and breaks 2, and the 95 percent interval on the difference runs from minus 6.8 to plus 10.8 points. That range includes zero, so read the 5 changed cases before deciding.

When should I use the Wilson interval for a pass rate?

Use the Wilson interval for a pass rate when the set is small or the failures are few. The NIST statistics handbook says the simple formula may not be accurate enough in those two conditions. With 49 passes out of 50, the simple formula gives an upper limit of 101.9 percent, which is impossible, and the Wilson interval is 89.5 to 99.6 percent. For 45 of 50, Wilson gives 78.6 to 95.7 percent.

How many test cases do I need to state a pass rate within 3 points?

Stating a pass rate within 3 points takes 385 test cases when the rate is 90 percent, and a margin of 2 points takes 865. The figures come from rearranging the interval formula in Evan Miller's paper. Eugene Yan states the general rule: halving the margin takes four times the sample. At 90 percent, the margin is 8.3 points on 50 cases and 4.2 points on 200.

Are 50 test cases enough to decide a release?

Fifty test cases are enough when the question is whether a change caused a large failure. Anthropic's engineering team suggests 20 to 50 tasks drawn from real failures as a start, on the grounds that early changes have large effects. A fall from 45 passes to 30 out of 50 gives a range of 46.4 to 73.6 percent, which lies entirely below the first range. A 2 point gain on 50 cases is one case and could be chance.

What should a report show next to an eval pass rate?

A report should show five things next to an eval pass rate: the count of cases and passes, the 95 percent interval, the number of runs per case, the dataset version, and for a comparison the paired interval with the list of cases that changed. Evan Miller's paper suggests printing the standard error beneath every eval score. A reader given a bare percentage should ask how many cases are behind it.