Build the test set

How many test cases an LLM eval needs: sample size and error bars

A pass rate of 90 percent measured on 100 test cases fits a true rate anywhere from 84.1 to 95.9 percent, using the simple 95 percent interval. On 1,000 cases the same result narrows to 88.1 to 91.9 percent. The number of cases needed to evaluate a large language model (LLM) product therefore depends on the size of the difference you want to detect. A set of 20 to 50 cases, the starting size Anthropic suggests, can show a large failure. Stating a rate to within 3 points takes 385 cases at a 90 percent pass rate. Detecting a move from 90 to 92 percent takes 1,666 cases by the formula in the NIST statistics handbook, and the comparison is best made with both versions run on the same cases.

Published September 30, 2026. Editorial.

Key takeaways

  • At a measured pass rate of 90 percent, the simple 95 percent interval runs from 84.1 to 95.9 percent on 100 cases and from 88.1 to 91.9 percent on 1,000 cases.
  • The standard error of a pass rate is the square root of p × (1 - p) / n, and the 95 percent interval is the rate plus and minus 1.96 standard errors. Evan Miller's paper prints both.
  • Halving the margin takes four times as many cases: at 90 percent the margin is 8.3 points on 50 cases and 4.2 points on 200.
  • Detecting a move from 90 to 92 percent with an 80 percent chance takes 1,666 cases by the NIST formula. To compare two versions, run both on the same cases and compute the interval on the case-by-case differences.
  • When the set is small or failures are few, use the Wilson interval, which NIST presents first: 49 passes out of 50 gives 89.5 to 99.6 percent, where the simple formula gives an impossible upper limit of 101.9 percent.

Two versions of a prompt, the written instruction given to a large language model (LLM), are tested on the same 100 cases. The first passes 90 and the second passes 92. The team releases the second because 92 is larger than 90. This is an invented illustration, and the arithmetic on this page shows its weakness: on 100 cases, a measured pass rate of 90 percent fits a true rate anywhere from 84.1 to 95.9 percent. A difference of 2 points is inside that range.

An eval is a repeatable test of an AI product's output. This page gives each formula as its source prints it and computes every figure. It is one step of the method in the guide to LLM evals.

A pass rate is an estimate with a range around it

The cases in an eval set are a sample of all the requests the product will receive. A different sample of the same size would give a different pass rate. The standard error is the typical size of that difference: how far a measured rate usually is from the true one.

Evan Miller's 2024 paper "Adding Error Bars to Evals" gives the formula for scores that are pass or fail [1]:

standard error = square root of ( p × (1 - p) / n )

Here p is the measured pass rate written as a fraction (0.90 for 90 percent) and n is the number of cases. The paper then gives the 95 percent confidence interval: the measured rate plus and minus 1.96 standard errors [1].

A confidence interval is the range of true rates that fits the measurement. At the 95 percent level, the method that builds the range captures the true rate in 95 of every 100 repeated tests. An error bar is that interval drawn on a chart, as a short line above and below the measured point.

The handbook of statistical methods from the US National Institute of Standards and Technology (NIST) writes the same interval with a symbol in place of 1.96, and calls it "the confidence expression most frequently used" [3].

The 95 percent interval at six set sizes

The table applies the formula to a measured pass rate of 90 percent. The fourth column is the simple interval above. The fifth is the Wilson interval, which the next section explains. Intervals are in percent.

Cases (n) Standard error (SE) Margin (1.96 × SE) Simple interval Wilson interval
50 4.24 points 8.3 points 81.7 to 98.3 78.6 to 95.7
100 3.00 points 5.9 points 84.1 to 95.9 82.6 to 94.5
200 2.12 points 4.2 points 85.8 to 94.2 85.1 to 93.4
500 1.34 points 2.6 points 87.4 to 92.6 87.1 to 92.3
1,000 0.95 points 1.9 points 88.1 to 91.9 88.0 to 91.7
2,000 0.67 points 1.3 points 88.7 to 91.3 88.6 to 91.2

The working for the row with 200 cases:

  1. p × (1 - p) = 0.90 × 0.10 = 0.09
  2. 0.09 / 200 = 0.00045
  3. The square root of 0.00045 is 0.021213, so the standard error is 2.12 points.
  4. 1.96 × 0.021213 = 0.04158, a margin of 4.2 points.
  5. 90 - 4.2 = 85.8 and 90 + 4.2 = 94.2.

One rule shows in the table. Going from 50 to 200 cases cuts the margin from 8.3 to 4.2 points, and going from 500 to 2,000 cuts it from 2.6 to 1.3. Eugene Yan states it: "to reduce the margin of error by half, we need to quadruple the sample size" [5]. Yan's own example applies the same arithmetic to a failure rate: 3 percent defects has a margin of 2.4 points on 200 samples and 1.7 points on 400 [5].

Where the simple formula is less accurate

NIST names two conditions in which limits built this way "may not be accurate enough for some applications": a small number of failures, or a small sample [3].

Suppose 49 of 50 cases pass, a rate of 98 percent. The simple formula gives a margin of 3.9 points and an interval of 94.1 to 101.9 percent. A pass rate above 100 percent is impossible. NIST raises this objection about a lower limit below zero: "A confidence limit approach that produces a lower limit which is an impossible value for the parameter for which the interval is constructed is an inferior approach." [3] The same arithmetic produces the impossible upper limit here.

NIST presents the Wilson method first. NIST names two advantages: "its worth does not strongly depend upon the value of n and/or p", and "the lower limit cannot be negative" [3]. For 49 of 50, the Wilson interval is 89.5 to 99.6 percent. In the table, the lower end of the Wilson interval is 3.0 points below the simple one at 50 cases and 0.1 points below it at 2,000.

Report the Wilson interval whenever the set is small or failures are few, the two conditions NIST names.

How many cases for the margin you want

Rearrange the formula to give the number of cases: n = 1.96 squared × p × (1 - p) / margin squared, with the margin written as a fraction. At a pass rate of 90 percent:

  • a margin of 5 points needs 139 cases (3.8416 × 0.09 / 0.0025 = 138.3, rounded up)
  • a margin of 3 points needs 385 cases
  • a margin of 2 points needs 865 cases
  • a margin of 1 point needs 3,458 cases

The vendor documentation read for this guide on 30 September 2026 sets no minimum size. The published numbers are named people's own suggestions, and each answers a different question. Anthropic's engineering team suggests 20 to 50 simple tasks drawn from real failures as a start, because early changes have large effects and "small sample sizes suffice" [6]. Hamel Husain and Shreya Shankar write that a purpose-built set "often grows to 100 or more examples", and that a check with one rule "may need only a few examples that Pass and a few that Fail" [7]. Miller suggests that new evals "should contain at least 1,000 questions", a figure tied to detecting a 3 point difference between two models under inputs the paper calls fictional [1].

A set of 20 to 50 cases, the size Anthropic suggests, will show a failure that affects a large share of requests, and finding such failures is the purpose of error analysis. Stating a rate to within 3 points takes 385 cases.

Is a 2 point improvement real?

NIST gives a formula for the smallest sample that can detect a change in a rate [4]:

N ≥ ( ( z1 × square root of (p0 × (1 - p0)) + z2 × square root of (p1 × (1 - p1)) ) / d ) squared

Here p0 is the old rate, p1 the new rate and d the difference between them. The value z1 sets how often chance alone will look like a change, and z2 sets how often a real change will be found. For a 5 percent chance of the first error, counted in both directions, z1 is 1.96, the value Miller's paper prints [1]. For an 80 percent chance of finding a real change, z2 is 0.8416, which we computed from the standard normal distribution, the bell-shaped curve used in statistics, with the statistics library of the Python programming language. As a check, the same two values put into Miller's own worked example, (1.96 + 0.8416) squared × (1/9) / 0.03 squared, return 969, the figure the paper prints [1].

For a change from 90 to 92 percent:

  1. The square root of (0.90 × 0.10) is 0.3, and 1.96 × 0.3 = 0.588.
  2. The square root of (0.92 × 0.08) is 0.27129, and 0.8416 × 0.27129 = 0.22832.
  3. (0.588 + 0.22832) / 0.02 = 40.816
  4. 40.816 squared is 1,665.95, so 1,666 cases.

The same formula gives 239 cases for a change from 90 to 95 percent. It treats the old rate as known. On 1,666 cases, a real move from 90 to 92 is found four times in five.

Compare two versions on the same cases

Miller's fourth recommendation is to analyse "the question-level paired differences" when two models are compared [1]. A paired comparison runs both versions on the same cases and takes the difference on each case: plus 1 where only the new version passes, minus 1 where only the old one passes, 0 where they agree. The paper's standard error for this is the square root of (variance of the differences / n), where variance measures how spread out the values are, and the interval is the mean difference plus and minus 1.96 standard errors [1].

An invented illustration: a new prompt fixes 9 cases out of 200 and breaks 5. The second column repeats it at ten times the size.

Step 200 cases: 9 fixed, 5 broken 2,000 cases: 90 fixed, 50 broken
Mean difference 4 / 200 = 0.02 40 / 2,000 = 0.02
Variance of the differences (14 - 200 × 0.0004) / 199 = 0.06995 (140 - 2,000 × 0.0004) / 1,999 = 0.06963
Standard error square root of (0.06995 / 200) = 0.0187 square root of (0.06963 / 2,000) = 0.0059
Margin (1.96 × SE) 3.7 points 1.2 points
95 percent interval minus 1.7 to plus 5.7 points plus 0.8 to plus 3.2 points

On 200 cases the interval includes zero, so the 2 point gain may be chance. On 2,000 cases it excludes zero. Pairing helps because two versions tend to pass and fail the same cases: Anthropic reports, from its own analysis, correlations between 0.3 and 0.7 in question scores between leading models [2]. A correlation measures how closely two sets of scores move together. The more two versions agree, the smaller the variance of their differences. Pairing also produces the list of cases that changed. Read those 14 cases before deciding. Both runs must use the same version of the set, as the page on building an eval dataset explains.

Run each case more than once

An LLM can give a different output for the same input. One run per case therefore mixes two sources of variation: which cases you chose, and which output the model produced this time. Miller's third recommendation is to reduce the second [1], and Anthropic's summary gives the method: "we recommend resampling answers from the same model several times, and using the question-level averages as the question scores" [2]. In practice, run each case several times, record the share of runs that passed, and use those shares as the scores.

Miller's worked example has 198 questions. Raising the number of runs per question from 1 to 10 "reduces the Minimum Detectable Effect from 13.2% to 7.5%" [1]. The minimum detectable effect is the smallest difference the test can find.

Repeating a case measures the same request again. New kinds of request come only from new cases. For products that act in several steps, the page on agent reliability across repeated runs covers the subject in full.

Cases that come in groups give less information

Ten questions about one document tell you less than ten questions about ten documents, because a model that misreads the document fails all ten together. Miller's second recommendation is to compute clustered standard errors when questions are drawn in related groups [1]. A clustered standard error treats each group as the unit of the sample. The paper's Table 4 compares the two calculations on three public test sets: on DROP the clustered standard error is 1.34 against a simple one of 0.44, a ratio of 3.05. On MGSM the ratio is 1.88, and on RACE-H it is 1.10 [1].

For a product eval, tag every case with its group: the document, the conversation, the customer. Where groups exist, the intervals in the first table are too narrow. Sample one case per group, or use the clustered calculation in Miller's paper.

What to report with every pass rate

Miller's advice is short: "We suggest reporting the standard error of the mean alongside (beneath) the mean when reporting eval scores." [1] For a product team, that becomes five items next to every rate:

  • the count of cases and the count that passed, such as 90 of 100
  • the 95 percent interval, using Wilson when the set is small or failures are few
  • how many times each case was run
  • the version of the dataset
  • for a comparison, the paired interval and the list of cases that changed

The page on how to read an AI eval report shows how a reader outside engineering checks them. For AI-written code the unit of counting differs, and the post on how many tests AI-generated code needs covers it.

How Reveneau sizes an eval set

At Reveneau we choose the size of an eval set from the decision it has to support. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an AI feature, the judged part of that suite is a set of cases like the ones on this page. We report each pass rate with its count of cases and its 95 percent interval, and when two versions of a prompt are compared, we run both on the same cases and report the paired interval.

We grade the judged checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. A faster grader shortens the wait when each case runs several times. Reveneau, as a company, takes responsibility for the whole project through production and after release, so every number we report is one we expect to be asked about later. To size an eval set for your own product, contact Reveneau.

Best for

  • Reporting a pass rate to someone who will decide a release
  • Comparing two prompt or model versions on the same cases
  • Deciding how many cases to add before the next measurement

Avoid if

  • You are still finding failure types: read the outputs first and compute a rate later
  • Cases are grouped by document or conversation and you plan to use the simple formula unchanged

Check before you decide

  • Each pass rate is shown with its count of cases and its 95 percent interval
  • Comparisons use the same cases and the same dataset version
  • The Wilson interval is used when the set is small or failures are few
  • The number of runs per case is stated

Common questions

How many eval cases do I need for an LLM product?

The number of eval cases depends on the difference you need to detect. At a measured pass rate of 90 percent, the simple 95 percent margin is 5.9 points on 100 cases and 1.9 points on 1,000. Anthropic's engineering team suggests 20 to 50 cases to start, which is enough to show a large failure and too few to state a rate within a few points.

What is an error bar on an eval score?

An error bar is the range of true pass rates that fits a measured eval score, drawn as a line above and below the point on a chart. Evan Miller's paper gives the usual version: the measured rate plus and minus 1.96 standard errors, which is a 95 percent confidence interval. For 90 percent on 200 cases, that range runs from 85.8 to 94.2 percent.

Is a 2 point improvement in an eval score real?

A 2 point improvement can be trusted only when the set is large enough to separate it from chance. By the sample size formula in the NIST handbook, detecting a move from 90 to 92 percent with an 80 percent chance takes 1,666 cases. On 100 cases the 95 percent interval for a 90 percent pass rate runs from 84.1 to 95.9, so a 2 point move is inside the range.

Should I run each eval case more than once?

Yes, when the model's output changes between runs. Anthropic's summary of Evan Miller's paper recommends answering each question several times and using the per-question averages as the scores. In the paper's example with 198 questions, going from 1 run to 10 runs per question reduced the smallest detectable difference from 13.2 percent to 7.5 percent. Repeated runs measure the same requests again, so coverage still comes from new cases.

What is a paired comparison in an eval?

A paired comparison runs two versions on the same eval cases and analyses the difference on each case: plus 1 where only the new version passes, minus 1 where only the old one passes, 0 where they agree. Evan Miller's paper recommends this over comparing two summary scores. The standard error is the square root of the variance of those differences divided by the number of cases.

What is a standard error in plain words?

A standard error is the typical distance between a measured pass rate and the true rate across all possible requests. For pass or fail scores, Evan Miller's paper gives it as the square root of p times (1 minus p) divided by n. At 90 percent on 200 cases the standard error is 2.12 points, and the 95 percent margin is 1.96 times that, which is 4.2 points.

Why is the simple interval formula less accurate for small eval sets?

The simple interval formula is an approximation, and the NIST handbook says it may not be accurate enough when the sample or the number of failures is small. With 49 of 50 cases passing, the simple formula gives an upper limit of 101.9 percent, which is impossible. NIST presents the Wilson method first. For 49 of 50, the Wilson interval is 89.5 to 99.6 percent.

How much does it cost to halve the margin on a pass rate?

Halving the margin on a pass rate takes four times as many cases. Eugene Yan states the rule, and the table on this page shows it: at 90 percent, the margin is 8.3 points on 50 cases and 4.2 points on 200. Stating a rate to within 3 points takes 385 cases at a 90 percent pass rate, and a margin of 1 point takes 3,458.

Are 20 to 50 eval cases enough to start with?

Yes, 20 to 50 cases are enough to find large failures. That is Anthropic's own advice, given on the grounds that early changes have large effects, so small samples suffice. A set that size cannot state a rate within a few points: at 90 percent on 50 cases, the simple 95 percent interval runs from 81.7 to 98.3 percent. Start small and grow the set as the differences you care about shrink.

What goes wrong when eval cases come in groups?

Grouped eval cases, such as ten questions about one document, carry less information than the same number of unrelated cases, so the simple formula reports a range that is too narrow. Evan Miller's paper compares both calculations on three public test sets: on DROP the clustered standard error was 1.34 against a simple one of 0.44, a ratio of 3.05. Tag each case with its group.

What should I report next to an eval pass rate?

Report the count of cases and passes, the 95 percent interval, the number of runs per case and the dataset version next to every eval pass rate. Evan Miller's paper suggests printing the standard error beneath each score. For a comparison between two versions, add the paired interval and the list of cases that changed, so a reader can judge whether the difference could be chance.

Does sample size matter for an eval check written in code?

Sample size matters less for a check with one rule. Hamel Husain and Shreya Shankar write that such a check may need only a few examples that pass and a few that fail, because the aim is to confirm the rule works. Sample size matters when you report a rate, such as the share of answers a grader marks correct, or when you compare two versions.

References