AI agent evals: how to test an agent that takes actions / Build and run the tests
Turning production traces into eval cases
Every production run of an agent leaves a trace, the record of each model call, tool call and result in that run. A trace becomes useful for testing when a person confirms that it shows a failure and writes down the outcome that should have happened. The method has six steps: record traces in a standard format, sample them, have a person review the failures, write each confirmed failure as a permanent case with its expected outcome, remove personal data before the case is stored, and re-run the whole set on every change. OpenTelemetry's naming conventions for generative AI are an open format for the first step, and their status reads Development as of 30 September 2026.
Published September 30, 2026. Editorial.
Key takeaways
- The first set can be small. Anthropic's January 2026 eval guide says 20 to 50 simple tasks drawn from real failures is a great start for an agent eval suite.
- A case is a starting state, an input and an expected outcome that two people who know the subject area would agree on. State the outcome as a fact about the world after the run, such as a refund record that exists.
- OpenTelemetry's generative AI conventions call message content sensitive and say recording tools should leave it off by default and let users switch it on. The conventions are now kept in their own GitHub repository with the status Development.
- No regulator in the sources for this page names AI traces. The UK regulator's general principles say personal data must be limited to what is necessary and kept no longer than needed, and applying them to traces is our inference for a lawyer to confirm.
- Remove personal data before a trace becomes a case, and replace each real value with an invented one of the same shape, because an eval set is a long-lived copy.
Take an invented furniture shop whose support agent tells a customer, "Your refund has been issued." Three days later the customer writes again: no money has arrived. Someone opens the record of that conversation and finds that the agent called the order lookup tool, made no call to the refund tool, and still reported success. The shop and the incident are an illustration.
That record is a trace. The shop can correct the agent's instructions and stop there, and the same fault can return with the next change. Or it can turn the trace into a test case that runs before every release. This page is the method for the second choice. It is part of the AI agent evals guide.
What a trace has to record
Anthropic's eval guide defines the record this way: "A transcript (also called a trace or trajectory) is the complete record of a trial, including outputs, tool calls, reasoning, intermediate results, and any other interactions" [1]. A trial is one attempt at one task. The page on why testing an agent differs from testing one model answer explains the term in full.
To be reusable as a test, a trace needs six things:
- The input that started the run.
- Every model call, in order.
- Every tool call, with its arguments and the result the tool returned.
- The final reply to the user.
- The versions in use: which model, which instructions, which tools.
- Time, cost and step count, which cost, time and step limits turns into checks.
Item 5 is easy to leave out. Without it, nobody can say whether a failure belongs to the current version or to one that was replaced last month.
A standard format: OpenTelemetry's conventions for generative AI
OpenTelemetry is an open standard, run as a Cloud Native Computing Foundation project, for recording what software did while it ran. Each recorded step is called a span. Its semantic conventions are agreed names for the things recorded, so that a trace written by one tool can be read by another.
One set of these conventions covers generative AI. As read on 30 September 2026, it covers model calls, agent runs, tool executions, errors and measurements of token use (a token is the unit of text a model reads and writes) [3]. The agent page defines named steps that include invoke agent, plan and execute tool, and the events page defines an event for recording an eval result [3].
Three facts about its condition matter before you depend on it:
- The status line of the conventions reads "Development" [3]. We could not open the page that defines that label, so we make no statement about what it guarantees.
- The conventions moved in 2026 to their own GitHub repository (a shared store of files). The page on the OpenTelemetry website that used to hold them is now titled "Moved: Generative AI semantic conventions" and says it "is no longer maintained in this repository" [3].
- In the new repository's introduction file, the heading "Schema URL" has one word under it: "TODO" [3].
Our position: use the conventions, because a shared format lets you change tracing tools and keep your old traces readable. Keep the mapping from your own field names to the convention's names in one file, so that a renamed field is a one-line change.
The six steps from a production run to a permanent case
| Step | Who does it | What is stored |
|---|---|---|
| 1. Record | The application, automatically | One trace per run, with message content only where you chose to record it |
| 2. Sample | A written rule, plus a person choosing extra runs | A review list, with the reason each trace was chosen |
| 3. Review | A person who knows the subject area | A label per trace (failure or acceptable) and the reason |
| 4. Write the case | The reviewer and an engineer | Starting state, input, expected outcome, link to the source trace |
| 5. Remove personal data | An engineer, checked by the reviewer | The case, with invented values in place of real ones |
| 6. Re-run | The test system, on every change | Pass or fail per case, per version of the agent |
Which traces to sample
Nobody can read every trace, so a rule picks them. The vendors' documentation starts from runs that went badly. LangChain's LangSmith documentation says to "filter the most interesting traces, such as traces that were tagged with poor user feedback, and add them to a dataset", and describes automation rules that add a trace when it has, for example, a low feedback score [5]. (A dataset here is a stored set of test cases.) Langfuse's documentation describes the same practice: "A common workflow is to select production traces where the application did not perform as expected" [6].
Add a second rule: a random sample of runs nobody complained about. Users report the failures they can see, and an agent that wrote a wrong delivery date into a record would get no complaint until the delivery failed. A fixed number of random traces each week, read by a person, finds failures that no user reported. Choosing what to look for in those traces is the subject of error analysis before metrics.
The set can start small. Anthropic's guide says: "20-50 simple tasks drawn from real failures is a great start" [1].
How to write one confirmed failure as a case
A reviewer who knows the subject area reads the trace and decides whether it shows a failure. If it does, four things are written down.
The starting state. What the order and the customer record looked like before the run. The case runs in a test environment rebuilt to this state each time, as test environments for agents explains.
The input. The user's message or messages, with personal data replaced.
The expected outcome. State it as a fact about the world after the run: a refund record exists for this order, for this amount. A correct agent can word its reply in many ways, so the agent's production wording stays out of the pass condition. Anthropic's test for a good task applies: "A good task is one where two domain experts would independently reach the same pass/fail verdict" [1]. Langfuse's documentation gives the job to a reviewer: "you let an expert add the expected output" [6]. The choice between checking the result and checking the steps is covered in outcome evals vs trajectory evals.
A link back to the source trace, so a later reader can see where the case came from.
Then run the new case several times against the version of the agent that produced the failure. It should fail at least once. A case that passes every time on that version is testing something else and needs rewriting before it is kept.
What four tools' own documentation says
Each row below is a vendor describing its own product, as read on 30 September 2026. We have made no comparison of how well these features work.
| Tool | What its documentation says | Detail to know |
|---|---|---|
| LangSmith (LangChain) [5] | "A common pattern for constructing datasets is to convert notable traces from your application into dataset examples." | Reviewers can change the inputs, outputs and reference outputs (the expected answers) before a trace is added |
| Langfuse [6] | An "Add to dataset" control on any recorded step of a production trace | A dataset item keeps the identifier of its source trace |
| Braintrust [7] | "Build datasets from production logs, user feedback, manual curation" | A record's input "can hold a reference to a logged trace" in place of a copied value |
| Arize Phoenix [8] | Any recorded step, or group of steps, can be added to a dataset from the interface | The reviewer can "make any changes you might need to make before saving the example" |
The Braintrust detail raises a design question. A case that refers to a stored trace depends on that trace still existing. Delete a source trace in a test and see what happens to the case.
Personal data: remove it before the trace becomes a case
A production trace holds what the customer typed. The OpenTelemetry conventions say so directly: "Model instructions, user messages, and model outputs are considered sensitive and are often large in size" [4]. They state the rule for tools that record traces: such tools "SHOULD NOT capture them by default" and "SHOULD provide an option for users to opt in" [4], meaning to switch the recording on by choice. The fields that hold messages have the requirement level "Opt-In" [4]. This text has the same Development status as the other conventions.
Two vendors document ways to limit recording at the source. LangSmith's documentation describes hiding inputs and outputs completely, or rule-based masking, which replaces the sensitive parts with a placeholder [9]. Langfuse's documentation says to "redact sensitive information before trace data leaves your application" [10]. Both pages describe features and make no claim about how complete the masking is. Test it with traces that contain invented personal data, and count what remains unmasked.
What the law says, and the limit of what we can tell you
We found no regulator that speaks about AI traces by name. What follows rests on OpenTelemetry's wording above and on general data protection guidance from the Information Commissioner's Office (ICO), the UK regulator. It covers UK law only, and a lawyer should confirm what applies to you.
The ICO's guidance on principle (c), data minimisation, says personal data must be "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed", and explains the last part as "you do not hold more than you need for that purpose" [11]. The page shows a notice that the guidance is under review because of the Data (Use and Access) Act [11]. The guidance makes no mention of traces. Our inference is that a trace holding a customer's name or message is personal data, so the general rule applies to it.
The purpose of an eval case is to test behaviour, and the behaviour rarely depends on the real name. Replace each real value with an invented one of the same shape: a different name, a made-up order number in the same format.
How long to keep traces
The ICO's guidance on principle (e), storage limitation, states: "You must not keep personal data for longer than you need it" [12]. It adds a duty to write the periods down: "You need a policy setting standard retention periods wherever possible" [12]. The guidance gives no number of days, and this page gives none either.
Applied to this method, again as our inference, which a lawyer should confirm: raw traces and eval cases are two stores with two purposes, so they can have two periods. Raw traces serve sampling and review, and that purpose ends once a trace has been read and labelled. Cases serve testing for as long as the agent exists, which is the reason step 5 removes personal data from them. The ICO's minimisation guidance also says people "have right to get you to delete any data that is not necessary for your purpose" [11], so keep a list of every place a trace is copied to.
Re-run the set on every change
A case is useful when it runs on every change to the instructions, the model or the tools. OpenAI's documentation states the move in one sentence: "Once you know what 'good' looks like, move from individual traces to repeatable datasets and eval runs" [2]. Braintrust's documentation describes datasets as "versioned collections of test cases" [7]. When a case is corrected, record which version of the set each result came from.
Checks on the final state are written in code. Checks that need judgment, such as whether the reply told the customer the truth, need a grader. Grading agent traces with Jev describes one way to do that, and the post never let the model grade its own work argues for keeping the grader separate from the model being graded. Repeat each case, as agent reliability across repeated runs describes.
A set built from production describes failures that already happened, and a feature that has not been released has no traces. New features need cases written from the specification before the code.
How Reveneau applies this
At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. Cases taken from production traces are the second source for that suite: the specification supplies the cases that can be foreseen, and production supplies the ones nobody foresaw. For every agent we build, we record traces with versions attached, we ask every client for a reviewer who knows the subject area, and we add each confirmed failure to the suite as a permanent case with personal data replaced.
The judged checks in the suite are graded by Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team and that saving goes into the client's price.
Reveneau, as a company, takes responsibility for the whole project through production and after release. Under this method, a failure found in production becomes a check that every later change must pass. Our AI development service describes how to start.
Best for
- An agent already in production with recorded runs
- Teams where a person who knows the subject area can review sampled traces each week
- Failures that users reported and that must be checked on every later change
Avoid if
- Traces do not record which model, instructions and tools were in use
- Message content is recorded with no decision about personal data
- The only cases you have describe features that are already released
Check before you decide
- Each new case fails at least once on the version that produced the failure
- No real name, address or order number remains in any stored case
- A written retention period exists for raw traces and another for cases
Common questions
What should a production trace record for an agent run?
A production trace should record the input that started the run, every model call in order, every tool call with its arguments and result, the final reply, the versions of the model, instructions and tools in use, and the time, cost and step count. Anthropic's eval guide defines the trace as the complete record of a trial, including outputs, tool calls, reasoning and intermediate results.
How do I turn a production failure into a test case?
Turn a production failure into a test case by having a person who knows the subject area confirm the failure, then writing down the starting state, the input with personal data replaced, the expected outcome and a link to the source trace. Run the new case several times against the version that failed, and keep it only when it fails at least once. Langfuse's documentation describes an expert adding the expected output.
How long should I keep production traces?
Keep production traces only as long as their purpose lasts, and write the period down. The UK Information Commissioner's Office says personal data must not be kept longer than you need it and asks for a policy setting standard retention periods. Its guidance gives no number of days and does not name AI traces, so applying it here is our inference, and a lawyer should confirm what applies to you.
How do I handle personal data in traces?
Handle personal data in traces by deciding whether to record message content at all, limiting it before it leaves your application, and replacing every real value with an invented one of the same shape before a trace becomes a case. OpenTelemetry's conventions say recording tools should not capture message content by default and should offer a setting to switch it on. LangSmith and Langfuse both document masking features.
What is OpenTelemetry for AI?
OpenTelemetry is an open standard for recording what software did while it ran, and its semantic conventions for generative AI are agreed names for model calls, agent runs and tool executions, so that a trace written by one tool can be read by another. As read on 30 September 2026, these conventions are kept in their own GitHub repository and their status line reads Development.
Are OpenTelemetry's generative AI conventions finished?
The status line of OpenTelemetry's generative AI conventions read Development on 30 September 2026, and the heading Schema URL in the repository's introduction file had only the word TODO under it. We could not open the page that defines the Development label, so we state no guarantee. Keep the mapping from your own field names to the convention's names in one file.
Which production traces should I review first?
Review first the production traces from runs that went badly: LangSmith's documentation suggests filtering traces tagged with poor user feedback, and Langfuse's documentation suggests traces where the application did not perform as expected. Then add a random sample of runs nobody complained about, because users report only the failures they can see, and a wrong value written into a record gets no complaint at first.
How many cases do I need before a trace-based eval set is useful?
A trace-based eval set is useful from a small number of cases. Anthropic's eval guide of January 2026 says 20 to 50 simple tasks drawn from real failures is a great start. Each case costs a reviewer's time to confirm and an engineer's time to write, so begin with the failures that users reported and grow the set every week from sampled traces.
What is the difference between a trace and an eval case?
A trace is the record of one run that already happened, and an eval case is a test that will run again on every change. The case keeps the starting state, the input and the expected outcome, and drops what the agent said in production. A trace usually holds personal data and can have a shorter retention period, while a case is kept as long as the agent exists.
What goes wrong when a case expects the agent's exact words?
A case that expects the agent's exact words fails a correct agent that phrases its reply differently, so it reports a failure where the agent was correct. State the expected outcome as a fact about the world after the run, such as a refund record that exists for the order. Anthropic's test for a good task is that two domain experts would independently reach the same pass or fail verdict.
Does this method work for an agent that has not been released?
The method on this page needs production traces, and an agent that has not been released has none, so its cases must be written from the specification before the code. The method starts once real runs exist, and a set built from production describes failures that already happened. Both sources are needed: the specification for what can be foreseen, and production traces for what nobody foresaw.
Which tools can turn a trace into a dataset case?
LangSmith, Langfuse, Braintrust and Arize Phoenix each document a feature for turning a trace into a dataset case, by each vendor's own documentation as read on 30 September 2026. LangSmith and Phoenix describe editing the example before saving it, Langfuse keeps the identifier of the source trace, and Braintrust lets a record hold a reference to a logged trace. We have made no comparison of how well they work.
References
- [1] Anthropic, Demystifying evals for AI agents (9 January 2026): definition of a transcript, also called a trace; 20 to 50 simple tasks drawn from real failures as a start; a good task is one where two domain experts reach the same verdict.
- [2] OpenAI, API documentation, Evaluate agent workflows (read 30 September 2026): the instruction to move from individual traces to repeatable datasets and eval runs.
- [3] OpenTelemetry, Semantic conventions for generative AI systems, README and linked files in the same repository (read 30 September 2026): status Development; what the conventions cover; the agent steps and the eval result event; the move from the OpenTelemetry website; the Schema URL heading with TODO under it.
- [4] OpenTelemetry, Semantic conventions for generative client AI spans (read 30 September 2026): message content is considered sensitive; instrumentations SHOULD NOT capture it by default and SHOULD provide an opt-in; the message fields carry the requirement level Opt-In.
- [5] LangChain, LangSmith documentation, Create and manage datasets in the UI (read 30 September 2026): converting notable traces into dataset examples; filtering traces tagged with poor user feedback; automation rules; reviewers can modify a trace before it is added.
- [6] Langfuse, documentation, Datasets (read 30 September 2026): selecting production traces where the application did not perform as expected; an expert adds the expected output; a dataset item keeps its source trace identifier; the Add to dataset control.
- [7] Braintrust, documentation, Datasets (read 30 September 2026): datasets as versioned collections of test cases built from production logs, user feedback and manual curation; a record's input can hold a reference to a logged trace.
- [8] Arize AI, Phoenix documentation, Creating Datasets (read 30 September 2026): adding a span or group of spans to a dataset from the interface; making changes before saving the example.
- [9] LangChain, LangSmith documentation, Prevent logging of sensitive data in traces (read 30 September 2026): hiding inputs and outputs completely; rule-based masking.
- [10] Langfuse, documentation, Masking (read 30 September 2026): redacting sensitive information before trace data leaves the application.
- [11] Information Commissioner's Office (UK), Principle (c): Data minimisation (read 30 September 2026): personal data must be adequate, relevant and limited to what is necessary; the right to have unnecessary data deleted; the notice that the guidance is under review.
- [12] Information Commissioner's Office (UK), Principle (e): Storage limitation (read 30 September 2026): personal data must not be kept longer than needed; periodic review; a policy setting standard retention periods.
Related reading
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.
More in Build and run the tests
Test environments for agents: sandboxes, fake tools and simulated users
An agent eval must never run against production, because an agent under test takes real actions: it issues refunds, sends email and deletes records. A test environment has three building blocks. A sandbox is a closed copy of the system where nothing the agent does reaches real data, and it is reset to the same starting state before every run. Fake tools return recorded or scripted results in place of services you cannot copy. A simulated user is a language model that plays the customer in a conversation that takes several turns. Each block distorts the result in a known way, and this page explains how to measure and limit each distortion.
Security evals for AI agents: prompt injection and unsafe actions
An agent reads text it did not write: web pages, emails, documents and the replies of its tools. Some of that text can be written by an attacker to look like an instruction, which is called prompt injection. A security eval plants such text in a test environment and measures how often the agent does what the attacker wanted. Published tests show that the measured rate depends on who writes the attack. In a January 2025 test by the Center for AI Standards and Innovation at NIST, the US government's standards body, the strongest earlier attack succeeded 11 percent of the time against one agent, and the strongest new attack 81 percent. A passing result lowers the measured rate, and attacks nobody has tried remain untested.