LLM evals: how to measure whether an AI product works / Build the test set
How to build an LLM eval dataset from real usage
An eval is a repeatable test of an AI product's output. For a product built on a large language model (LLM), the eval dataset is a list of real requests, each stored with the context the model received and a written statement of what a correct response must do. Build it from your own logs, support tickets and past requests, add difficult requests on purpose, and use model-written cases only to fill gaps. Keep one part of the set away from the people who write the prompts, meaning the instructions given to the model. Replace personal data before a case is stored, and give every change a version number. Anthropic's engineering team suggests 20 to 50 cases to start. The vendor documentation read on 30 September 2026 states no minimum.
Published September 30, 2026. Editorial.
Key takeaways
- A golden dataset is a set of test cases whose expected answers a person has confirmed. For most product answers, a written expectation is a better target than one reference answer.
- Each case holds four parts: the input as it arrived, the context the model received, the expected behaviour in words, and tags for source, date and failure type.
- Keep a random sample of real traffic and a hand-chosen group of difficult cases as two tagged groups, and report their pass rates separately.
- Model-written cases widen coverage. Hamel Husain and Shreya Shankar state the limit: synthetic data cannot tell you how common a failure is in production.
- Replace personal data when a case enters the set, reserve a part that prompt writers never open, and store the dataset version with every result.
Take a support assistant for an invented furniture shop. The team tests it with twelve questions written in one afternoon: opening hours, delivery cost, how to return a sofa. All twelve pass. In the first week after release, customers ask about one order split across two deliveries, paste an order number with a typing mistake in it, and write half a message in Spanish. The twelve questions contained none of these.
That example is an invented illustration. OpenAI's documentation lists the pattern in it among common mistakes: "Biased design: Creating eval datasets that don't faithfully reproduce production traffic patterns." [1] An eval, short for evaluation, is a repeatable test of an AI product's output, and its dataset is the list of cases the test runs. This page covers where those cases come from and how to store them. It is one step of the method in the guide to LLM evals, where LLM stands for large language model.
What a golden dataset is
"Golden dataset" is the name people search for. It means a set of test cases whose expected answers a person has confirmed. Google Cloud's documentation uses the same word for one field of a case, the reference, which it describes as the ground truth or "golden" answer "that you can compare the model's response against" [2]. Ground truth here means the answer accepted as correct.
Google marks the reference field as optional and says it is often required for the scores that count matching words [2]. A support answer can be correct in many different wordings, so one confirmed answer is a poor target. A written expectation works better: "states the 30 day return period and gives the returns link". A person still confirms that expectation. The post on how to test a feature you cannot fully specify covers features with no single correct output.
Where eval cases come from
OpenAI's guide lists six kinds of data to consider: "synthetic eval data, domain-specific eval data, purchased eval data, human-curated eval data, production data, and historical data" [1]. The same guide tells teams to log everything during development, so that the logs can later be searched for good cases [1]. Anthropic's engineering team suggests starting with 20 to 50 simple tasks drawn from real failures [4].
| Source of cases | What it is good for | The risk |
|---|---|---|
| Production logs: the stored record of real requests and responses | Shows what people ask, in their own words | Holds personal data, and most entries are easy requests |
| Support tickets and complaints | Failures that a user reported | Leaves out how common each failure is |
| Past requests the feature will replace, such as emails answered by staff | Available before release, with a person's answer attached | People may word a request to staff differently |
| Cases written by a subject expert | Costly situations that have never happened yet | Reflects what one expert expects to see |
| Cases written by a model | Fills gaps in coverage quickly | Cannot show how common a failure is [5] |
A team with no users yet starts from the third and fourth rows. A team with a live product starts from the first two. Eugene Yan and five co-authors, in a 2024 report on building with LLMs, give the same instruction for checks written in code: build them from "samples of inputs and outputs from production" [6].
What one case contains
A case needs four parts.
- Input. The request exactly as it arrived, typing mistakes included.
- Context. Everything else the model received: the documents a search step fetched, the customer's order record, the earlier turns of the conversation. Stored context lets the case run again in the same conditions.
- Expected behaviour. A written statement of what a correct response must do. Anthropic's test for a well-written one: "A good task is one where two domain experts would independently reach the same pass/fail verdict." [4] Yan and co-authors ask for expectations "based on at least three criteria" for each sample [6].
- Tags. Where the case came from, the date it was added, the failure type it tests, and whether a person or a model wrote it.
A reference answer is a fifth, optional part. Add one when the answer has a single correct form, such as a category label or an amount.
Include cases where the product must refuse. Anthropic's advice reads: "Test both the cases where a behavior should occur and where it shouldn't." [4] For the furniture shop, that means one case where the assistant must refuse a refund, next to one where it must grant it.
Cover common requests and difficult ones on purpose
Two aims compete. The first is to match real traffic. Google says example prompts should "represent the types of inputs that your models process in production" [2]. Anthropic's documentation says: "Design evals that mirror your real-world task distribution", and reminds the reader to include the rare cases [7]. Yan and co-authors give a concrete test: if typing mistakes are common in production inputs, the test data should contain them too [6].
The second aim is to contain enough failures to learn from. Yan, writing alone in November 2025, recommends "at least 50-100 failures out of 200+ total samples" for checking a grader, meaning the person or model that marks each output as pass or fail [8]. A random sample of traffic from a product that mostly works will contain few failures.
Keep both aims as two tagged groups inside one set. The first group is a random sample of traffic, and its pass rate estimates how the product performs for users. The second group is chosen by hand: the difficult requests, the failure types found during error analysis, the cases behind past complaints. Its pass rate shows whether a fix worked. Report the two rates separately, because a blended figure depends on how many difficult cases you chose to add.
Cases written by a model help coverage and mislead as the only source
Anthropic's documentation recommends model-written cases, often called synthetic data: "Get Claude to help you generate more from a baseline set of example test cases." [7] The sentence assumes a starting set that already exists.
Practitioners state two limits. Hamel Husain and Shreya Shankar write: "Synthetic data cannot tell you how common a failure is in production." [5] Yan describes model-written failures as "either too exaggerated or too subtle in ways that don't reflect what happens in production" [8].
A workable method follows from those limits.
- Start from real cases.
- Ask the model for variations along dimensions you name: a second language, a longer message, a missing order number.
- Have a person read every generated case before it enters the set.
- Tag each one as model-written, so its results can be removed from any figure that claims to describe real traffic.
Keep one part of the set that prompt writers never see
A prompt is the written instruction given to the model. When a team edits the prompt until every visible case passes, the pass rate on those cases stops predicting what happens on new requests. The remedy is a held-out set: a part of the dataset kept apart and used only for a final check.
Anthropic's documentation names a "held-out test set" in a sample target, and OpenAI's names "a held-out set of 1000 reference transcripts" [7][1]. As read on 30 September 2026, the two pages give those examples and no method for building such a set. The explicit splits come from practitioners, and both concern labelled examples used to check a model that grades outputs. Yan uses 75 percent of the samples for adjusting the grader's prompt and holds out 25 percent as the test set [8]. Husain and Shankar reserve 40 to 45 percent of their labelled examples "for one final test" [5].
Apply the same idea to the product's own prompt:
- Choose the reserved cases at random within each tag, so the reserved part has the same mix as the visible part.
- Store them in a location closed to prompt writers.
- Run them at release decisions only.
- Once someone studies a reserved case in detail to fix a failure, move it to the visible part and reserve a new one.
Remove personal data before a case is stored
Real requests carry names, email addresses, phone numbers and order numbers. An eval set is copied and shared with graders, so clean each case once, at the moment it enters the set.
- Replace each personal value with an invented value of the same shape: a made-up name for a name, a made-up order number with the same count of digits.
- Read the free-text fields and any attached documents. People paste personal details into the body of a message, and a rule that checks only named fields will miss them.
- Keep no table that maps invented values back to real ones.
- Record in the case's tags which fields were replaced.
- Limit access to the raw logs to the people who do the cleaning.
Which data protection law applies to your logs is a question for a lawyer. Products in regulated fields carry further duties, which the guide to software for regulated industries covers.
Version the set and change it on a schedule
Give the dataset a version number and raise it with every added, edited or retired case. Store the version next to every eval result. Compare two pass rates only when both come from the same version. The page on how many test cases an eval needs shows the comparison.
OpenAI's guide says to "grow the eval set over time" [1]. LangChain describes the route: problems found by online evaluations become offline test cases [3]. Online means measured on live traffic, and offline means run on the stored set. Husain and Shankar give a schedule from their own practice: review at least 100 fresh traces in each cycle, with cycles of 2 to 4 weeks, and add representative examples whenever monitoring shows a new failure pattern [5]. A trace is the full record of one request, from input to final response. The page on offline and online evals describes that loop.
Retire a case when the feature it tests is removed. Mark it retired instead of deleting it, so older results can still be reproduced.
How large the first version should be
Google's dataset page, as read on 30 September 2026, gives no number of examples [2], and the sizes in Anthropic's sample code are illustrations [7]. The figures that exist are named people's own working advice: 20 to 50 tasks to start from Anthropic's engineering team [4], and a set that "often grows to 100 or more examples" from Husain and Shankar, who add: "Coverage determines the final size." [5] The same two authors report spending 60 to 80 percent of development time on error analysis and evaluation in their own projects [5], which is a reason to put the dataset in the plan from the first week.
How Reveneau builds an eval dataset
At Reveneau we write the expected behaviour for each case before we write the prompt. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code exists, before it is released. For an AI feature, the dataset described on this page is part of that suite. We ask every client for the real requests the feature will handle: past tickets, emails, search queries. We clean personal data out of them before any case is stored, tag each case by source and failure type, and reserve a part that the people writing prompts do not open.
Because AI does the writing, a Reveneau build takes a small team, and that saving goes into the client's price. The time people spend on the build goes to reading cases and confirming expectations. Reveneau, as a company, takes responsibility for the whole project through production and after release, so the dataset keeps receiving new cases from live usage after the first release. To plan an eval dataset for your own product, see AI development at Reveneau.
Best for
- Teams with production logs, support tickets or past requests to draw cases from
- Features whose correct response can be described in a written expectation
- Products that will keep changing after release and need comparable results over time
Avoid if
- You have yet to read real outputs and name the failure types: do error analysis first
- The only cases available are model-written, with no real request to start from
Check before you decide
- Every case has an input, the stored context, an expected behaviour and tags
- Personal data was replaced before the case was stored
- Every eval result records the dataset version it ran on
- A reserved part exists that prompt writers have never opened
Common questions
What is a golden dataset for an LLM product?
A golden dataset is a set of test cases whose expected answers a person has confirmed. Each case holds a real request, the context the model received, and either a reference answer or a written statement of what a correct response must do. Google Cloud's documentation uses the word for the reference field, which it calls the ground truth or golden answer and marks as optional.
Where do LLM eval cases come from?
Eval cases come from the product's own usage first: production logs, support tickets and complaints, and past requests that staff answered before the feature existed. Cases written by a subject expert cover costly situations that have never happened, and cases written by a model fill gaps. OpenAI's guide lists six kinds of data to consider, including production data, historical data and synthetic data.
Can I use synthetic data for LLM evals?
Yes, as an addition to real cases. Synthetic data, meaning cases written by a model, fills gaps in coverage quickly, and Anthropic's documentation recommends generating more cases from a baseline set. Hamel Husain and Shreya Shankar state the limit: synthetic data cannot tell you how common a failure is in production. Have a person read each generated case and tag it as model-written.
How do I keep personal data out of an eval set?
Clean each case once, at the moment it enters the eval set. Replace names, email addresses, phone numbers and order numbers with invented values of the same shape, read the free-text fields where people paste personal details, keep no table linking invented values to real ones, and record which fields were replaced. Ask a lawyer which data protection law applies to your logs.
How often should an eval set change?
An eval set should change whenever live usage shows a failure type the set does not contain, and on a regular review cycle as well. Hamel Husain and Shreya Shankar describe cycles of 2 to 4 weeks in which they review at least 100 fresh traces and add representative examples. Give every change a new version number, and store that number with each result.
What should one eval case contain?
One eval case contains four parts: the input exactly as it arrived, the context the model received, a written statement of expected behaviour, and tags for source, date and failure type. Anthropic's test for the expected behaviour is that two domain experts would independently reach the same pass or fail verdict. A reference answer is optional and suits answers with one correct form.
Do I need a reference answer for every eval case?
No. A reference answer is needed only where the answer has one correct form, such as a label or an amount. Google Cloud's documentation marks the reference field as optional and says it is often required for scores that count matching words. For answers that can be worded many ways, a written expectation such as 'states the 30 day return period' is the better target.
What is a held-out set in an eval dataset?
A held-out set is a part of the eval dataset kept away from the people who write the prompt and used only for a final check before release. Edits made until every visible case passes can fit the prompt to those cases, and the held-out part shows whether the gain also appears on new requests. Eugene Yan holds out 25 percent of labelled samples when checking a grader.
What goes wrong when an eval dataset holds only easy cases?
An eval dataset of easy cases reports a high pass rate that says little about real use. OpenAI's guide lists this among common mistakes: building eval datasets that fail to reproduce production traffic patterns. A set of easy cases also holds too few failures to learn from. Eugene Yan recommends at least 50 to 100 failures out of 200 or more samples for checking a grader. Add difficult requests on purpose and tag them.
Can I build an eval dataset before the product has users?
Yes. Before release, an eval dataset can be built from past requests the feature will replace, such as emails and tickets answered by staff, and from cases a subject expert writes. OpenAI's guide names historical data and human-curated data among the sources to consider. Add model-written variations to widen coverage, then extend the set with cases from production logs once real usage exists.
How much work is building an eval dataset?
Building an eval dataset is a large share of the work on an AI feature. Hamel Husain and Shreya Shankar report spending 60 to 80 percent of development time on error analysis and evaluation in their own projects, which is their own account and comes with no dataset behind it. Plan for a person who knows the subject to read cases and confirm each expectation from the first week.
What should I do after the first version of the eval dataset exists?
After the first version of the eval dataset exists, run it on every change to the prompt or the model, store the dataset version with each result, and set a review cycle for adding cases from live usage. OpenAI's guide says to grow the eval set over time. The next decision is how each failure type will be scored: by a check in code, a person, or a model grader.
References
- [1] OpenAI, Evaluation best practices (API documentation, undated, read 30 September 2026): the six kinds of eval data to consider, the advice to log during development, the "Biased design" mistake, the held-out set of 1000 reference transcripts used in its example, and the instruction to grow the eval set over time.
- [2] Google Cloud, Prepare your evaluation dataset (read 30 September 2026): prompts should represent production inputs; the reference field is optional and holds the ground truth or "golden" answer; the page states no recommended number of examples.
- [3] LangChain, Evaluation concepts (LangSmith documentation, undated, read 30 September 2026): the statement that online evaluations find issues that become offline test cases.
- [4] Anthropic, Demystifying evals for AI agents (9 January 2026): 20 to 50 simple tasks drawn from real failures as a start; two domain experts should reach the same verdict; test cases where a behaviour should and should not occur.
- [5] Hamel Husain and Shreya Shankar, AI Evals: Everything You Need to Know (page dated 18 September 2026): synthetic data cannot show how common a failure is; 40 to 45 percent of labelled examples reserved for one final test; sets that grow to 100 or more examples; review cycles of 2 to 4 weeks with 100 fresh traces; 60 to 80 percent of development time in their own projects.
- [6] Eugene Yan, Bryan Bischof, Charles Frye, Hamel Husain, Jason Liu and Shreya Shankar, What We've Learned From A Year of Building with LLMs (8 June 2024): checks built from production samples with at least three criteria; held-out data should contain the typing mistakes that production inputs contain.
- [7] Anthropic, Define success criteria and build evaluations (Claude Platform documentation, undated, read 30 September 2026): evals should mirror the real task distribution; the held-out test set named in its example; sample sizes that are illustrations; generating more test cases from a baseline set.
- [8] Eugene Yan, Product Evals in Three Simple Steps (November 2025): at least 50 to 100 failures out of 200 or more samples; model-written failures are too exaggerated or too subtle; 75 percent for adjusting the grader and 25 percent held out.
Related reading
How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.
How to write an acceptance test a machine can run
Most acceptance criteria are written for a human reader who will add the missing details. A machine adds nothing. Here is how to write the sentence so an automated check can enforce it.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.