Start here

Error analysis: read the outputs before you choose a metric

Error analysis is the step that comes before any metric: read a sample of real outputs, write a short note on each failure, group the notes into failure types, and count them. Only then choose what to measure. Hamel Husain and Shreya Shankar recommend reading at least 100 traces and stopping when new ones show no new failure type. A 2024 study by Shankar and four co-authors, with nine practitioners, found that reviewers changed their criteria as they read more outputs, which is why criteria written before the reading miss failures. The reader should be a person who knows what a correct answer is.

Published September 30, 2026. Editorial.

Key takeaways

  • Error analysis has four steps: collect a random sample of real outputs, write a note on each failure, group the notes into failure types, and count each type before choosing any measure.
  • Hamel Husain and Shreya Shankar recommend reviewing at least 100 traces and stopping when new traces show no new failure type. These numbers are their own suggestions, with no dataset behind them.
  • A 2024 study by Shankar and four co-authors, with nine industry practitioners, found that people changed their grading criteria as they graded more outputs. Write the criteria after the first reading pass and relabel the early outputs.
  • A ready-made score measures a property somebody else chose. OpenAI's documentation lists relying solely on academic metrics such as BLEU among common mistakes in eval design.
  • The reader should be a person who knows what a correct answer is. Husain and Shankar recommend one domain expert with the final say on quality for most small and medium-sized companies.

Take a support assistant for an invented furniture shop. The team adds a ready-made "helpfulness" score from 1 to 5, given by a second model, and the average is 4.3. Customers keep complaining. Then one person reads 100 replies. In 11 of them the delivery date in the reply differs from the date in the order record. The helpfulness score had rated those 11 replies well, because each was polite and complete. Every figure in this example is invented to illustrate the method.

The reading that found the wrong dates has a name, error analysis, and it belongs at the start. Hamel Husain and Shreya Shankar, who teach a course on evals, put it this way: "Error analysis is the most important activity in evals. Error analysis helps you decide what evals to write in the first place." [1] This page gives the steps, the number of outputs to read, and the research finding that explains why the step cannot be skipped. It is the second step of the method in LLM evals: how to measure whether an AI product works.

Error analysis in four steps

An LLM is a large language model, an AI model that produces text. Error analysis for a product built on one has four steps:

  1. Collect a sample of real outputs.
  2. Write a short note on each output that failed.
  3. Group the notes into failure types.
  4. Count each failure type, then decide what to measure.

A failure type is a named group of failures that share one visible symptom, such as "the delivery date differs from the order record". The sources on this page call it a failure mode. Husain and Shankar describe the method as adapted from qualitative research, the branch of research that studies written material by reading and sorting it [1].

One caution about the sources. Husain, Shankar and Eugene Yan have written together, so their agreement is the view of one group and counts once. Their numbers are working guides from their own projects, with no dataset behind them. The research evidence on this page is one paper, covered below.

Step 1: collect a sample of real outputs

The unit to read is a trace: the full record of one request, holding what the user asked, what the model was given, and what it answered. You can only read traces that were saved. OpenAI's documentation tells teams to log everything during development, so that the logs can later supply eval cases [7]. Yan and five co-authors say the same about a live product: "we should log LLM inputs and outputs. By examining a sample of these logs daily, we can quickly identify and adapt to new patterns or failure modes." [5]

Three instructions for the sample:

  • Draw it at random. If you pick the outputs that look interesting, the counts in step 4 describe your choices, and how often users see each failure stays unknown.
  • Make each trace quick to read. Husain says to remove every obstacle from the process of looking at the data [3]. Put the request, the documents the model was given and the reply on one screen.
  • Before release, use what you have. Husain says a team can write likely user requests itself and generate synthetic data, which means test inputs written by a model [3]. Husain and Shankar state the limit: "Synthetic data cannot tell you how common a failure is in production." [1]

Where real cases come from, and how to keep personal data out of them, is covered in how to build an LLM eval dataset.

Steps 2 and 3: write a note on each failure, then group the notes

For each trace, first decide pass or fail. Husain starts by labelling examples as good or bad, and reports that finer ratings are harder to manage than two labels [3].

For each fail, write what you saw in plain words: "says delivery on 4 March, the order record says 11 March". Husain and Shankar call this open coding: free-text notes with no list of categories prepared in advance [1]. Leave the categories for later. A list written before the reading holds only the failures you expected, and the purpose of the reading is to find the others.

When the notes are written, sort them. Husain and Shankar call this axial coding and describe it like this: "Categorize the open-ended notes into a "failure taxonomy." In other words, group similar failures into distinct categories." [1] Give each group a name that states the symptom. Merge two groups when one fix would repair both. Split a group when its notes describe two different causes.

Step 4: count the failure types, then choose a check for each

This table continues the invented furniture shop example. The counts are an illustration.

Failure type Outputs out of 100 The check it leads to
Delivery date differs from the order record 11 Code check: compare the date in the reply with the date in the order record
Upset customer was kept with the assistant instead of being passed to a person 8 Human labels first, then a model grader compared with those labels
Reply is longer than the chat window displays 6 Code check: word count under the limit
Return policy quoted for the wrong country 5 Code check on the search step: the policy document fetched matches the customer's country
Order number missing from the reply 4 Code check: the reply contains the order number
Reply in a different language from the request 3 Code check: the language of the reply equals the language of the request
No failure found 63 None
Total 100

The six failure types add up to 37 outputs (11 + 8 + 6 + 5 + 4 + 3), and 37 plus 63 is 100.

Three things in this table could only come from reading. The largest failure type is one that no general quality score looks for. Five of the six checks are code, the cheapest of the three kinds of LLM eval, so the model grader the team started with was the wrong tool for most of the work. And the fourth row is a fault in the step that fetches documents, which is tested separately, as RAG evaluation explains.

Decide the order of work from two facts about each failure type: how often it occurs and how much harm one occurrence does. A count from 100 outputs is precise enough to rank the failure types. It is too small to publish as a rate, for the reasons in how many test cases an LLM eval needs.

How many outputs to read, and how often

No research paper read for this page tests how many outputs are enough. Every number below is one named person's or company's own suggestion.

Who The number What it refers to
Hamel Husain [4] 30 examples as a starting point A first batch, continued "until I do not see any new failure modes"
Husain and Shankar [1] At least 30 traces Notes written by a person before any AI tool is asked to suggest failures
Husain and Shankar [1] At least 100 traces A full first review
Husain and Shankar [1] 100 or more fresh traces every 2 to 4 weeks Each later review cycle
Anthropic engineering post, 9 January 2026 [8] 20 to 50 tasks The first set of test cases, "drawn from real failures"

The stopping rule matters more than any one number. Husain and Shankar write: "Continue until new traces stop revealing failure modes or changing existing ones. Qualitative researchers call this theoretical saturation." [1] In plain words, stop when more reading adds no new group to your list.

The last row of the table measures a different thing. Reading 100 traces is error analysis. The 20 to 50 tasks are the test cases that the reading produces.

On time, Husain and Shankar say that error analysis and evaluation took 60 to 80 percent of development time on their own projects [1]. That figure is their account of their own work.

Your criteria will change while you read

In 2024 Shreya Shankar and four co-authors published a study of how people grade model outputs. It was a qualitative study with nine industry practitioners [2]. Their main finding has a name, criteria drift: "users need criteria to grade outputs, but grading outputs helps users define criteria" [2]. Drift here means that the criteria change as the reviewer reads more.

The paper reports what the participants did: "Even when participants graded first, we observed that they still refined their criteria upon further grading, even going back to change previous grades." [2] It also states a stronger claim: "it is impossible to completely determine evaluation criteria prior to human judging of LLM outputs" [2].

The study has limits. It had nine participants, working on one supplied task in a single session, and it gives no rate for how often criteria changed. Yan and co-authors cite the same finding in their own guide [5].

The finding leads to three instructions:

  • Write the criteria after the first reading pass.
  • Go back and relabel the earliest outputs once the failure types have stopped changing.
  • Put a date on each version of the criteria, so that a pass rate from March can be compared with one from June.

A feature whose correct behaviour nobody can fully write down in advance still gets tested this way, as how to test a feature you cannot fully specify argues.

Why a ready-made metric is the wrong starting point

A ready-made metric measures a property that somebody else chose, for a different product. OpenAI's documentation lists this among common mistakes: "Overly generic metrics: Relying solely on academic metrics like perplexity or BLEU score." [7] Perplexity and BLEU are two scores from academic research. BLEU counts shared words between an output and a reference text. Husain makes a related point about scores on a scale: "Tracking a bunch of scores on a 1-5 scale is often a sign of a bad eval process" [4].

The furniture shop example shows the mechanism. A helpfulness score rewards a polite, complete reply. A reply with the wrong delivery date is polite and complete. The score stays high while the failure that customers complain about goes uncounted.

Automatic checks also need reading to stay correct. Anthropic's engineering post says: "You won't know if your graders are working well unless you read the transcripts and grades from many trials." [8] Yan writes that when a team has stopped reviewing outputs and customer feedback, automated evaluators will leave the product's problems in place [6].

Metrics come after the failure types. Once the table above exists, each row gets its own measure, and LLM eval metrics describes the choices.

Who should read the outputs

The reader has to know what a correct answer is. For the furniture shop that is the person who runs the support team, because they know the return policy and can tell when a reply misstates it. Husain warns that many developers try to act as the domain expert, the person who knows the subject the product answers questions on, or pick whoever is convenient, and that this goes badly [4].

Husain and Shankar recommend that most small and medium-sized companies appoint one domain expert who has the final say on quality [1]. One reader gives one consistent set of judgments. When a second reviewer joins, the two have to be compared, and human review and reviewer agreement gives the method.

An AI tool can help sort notes into groups. Husain and Shankar say to write notes on at least 30 traces yourself before asking a tool to suggest failures [1]. For a product that takes actions as well as writing text, the same reading applies to longer traces, as turning production traces into eval cases describes.

How Reveneau applies this

At Reveneau we read outputs before we choose a measure. For an AI product, we ask every client for two things at the start: a sample of real requests, or the requests the feature will replace, and a person on the client's side who knows what a correct answer is. We read the outputs with that person, write the failure types down, and turn each one into a written expectation with its own check.

That order follows from how we build. All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. The failure types found in error analysis become part of that specification, so the suite tests the failures users meet. Reveneau, as a company, takes responsibility for the whole project through production and after release, and the reading continues after release on fresh outputs. To start with a sample of your own requests, contact us.

Best for

  • A product with saved outputs that nobody has read in a structured way
  • A team about to choose its first metric or its first model grader
  • A product whose users complain while its quality score stays high

Avoid if

  • No person who knows the correct answers is available to read
  • The sample was chosen by hand, so the counts cannot be read as frequencies

Check before you decide

  • Each failure type has a name, a count, and one check assigned to it
  • A further batch of traces added no new failure type
  • The criteria carry a date and the earliest outputs were relabelled

Common questions

What is error analysis for an LLM product?

Error analysis for an LLM product is the practice of reading a sample of real outputs, writing a note on each failure, grouping the notes into failure types and counting them, before any metric is chosen. Hamel Husain and Shreya Shankar call it the most important activity in evals, because it decides which evals are worth writing in the first place.

How many outputs should I read before choosing a metric?

Read at least 100 outputs before choosing a metric, and stop when new ones show no new failure type. That is the recommendation of Hamel Husain and Shreya Shankar, who also suggest writing notes on the first 30 traces by hand. Both numbers are their own working guides from their projects, and no study on this page tests them.

What is a failure type in an LLM eval?

A failure type is a named group of failures that share one visible symptom, for example a delivery date in the reply that differs from the order record. The sources call it a failure mode. Husain and Shankar describe building the list by grouping free-text notes into what they call a failure taxonomy, and each failure type then gets its own check.

Why should a team avoid starting with a ready-made metric?

A ready-made metric measures a property that somebody else chose, so it can stay high while your product fails in a way the metric never looks at. OpenAI's documentation lists relying solely on academic metrics such as perplexity or BLEU among common mistakes. Choose the measure after error analysis has shown which failure types occur and how often.

Who should read the outputs during error analysis?

The outputs should be read by a person who knows what a correct answer is, such as the head of the support team for a support assistant. Husain and Shankar recommend that most small and medium-sized companies appoint one domain expert with the final say on quality. Hamel Husain warns that developers who act as the expert themselves get poor results.

How much time does error analysis take?

Error analysis takes a large share of the work. Hamel Husain and Shreya Shankar say that error analysis and evaluation together took 60 to 80 percent of development time on their own projects. That figure is their account of their own work, with no dataset behind it, so treat it as a planning guide and measure your own time on the first 100 traces.

Why do grading criteria change while a reviewer reads more outputs?

Grading criteria change because reading outputs teaches the reviewer what the product can get wrong. A 2024 study by Shreya Shankar and four co-authors, with nine industry practitioners, named this criteria drift. Participants who graded first still refined their criteria later and went back to change earlier grades. The study gives no rate, because it was qualitative.

Can an AI tool do the error analysis for me?

An AI tool can help group notes into failure types, and a person should do the first reading. Husain and Shankar recommend writing notes on at least 30 traces yourself before asking a tool to suggest failures. A tool sorts the failures it is shown, while a reader who knows the subject is the one who sees that an answer is wrong.

How do I do error analysis before the product has users?

Before the product has users, write the requests you expect and have a model generate more, then read the outputs those requests produce. Hamel Husain recommends this in a 2024 post. Husain and Shankar state the limit: synthetic data cannot tell you how common a failure is in production, so repeat the reading on real traces after release.

How often should error analysis be repeated after the first round?

Repeat error analysis on a regular cycle after the first round, because users keep sending requests that no earlier review covered. Husain and Shankar set a goal of 100 or more fresh traces in each review cycle and report cycles of 2 to 4 weeks. Yan and five co-authors go further and suggest reading a sample of the logs daily.

What happens when a team skips error analysis?

A team that skips error analysis measures what was easy to measure and misses the failures its users meet. Anthropic's engineering post of 9 January 2026 says a team will not know whether its graders work unless it reads the transcripts and grades from many trials. In the invented example on this page, a helpfulness score of 4.3 was reported while 11 of 100 replies gave a wrong delivery date.

What should I do after counting the failure types?

After counting the failure types, rank them by how often each occurs and how much harm one occurrence does, then assign each one a check. In the invented example on this page, five of six failure types led to a code check and one needed human labels followed by a model grader.

References