Start here

Why testing an agent is different from testing one model answer

An agent eval is a repeatable test of an AI agent, a language model that works in a loop: it picks a step, calls a tool such as a search or a database, reads the result and repeats until the task is done. Testing an agent differs from testing one model answer in four ways. The agent takes many steps, so a small error rate at each step multiplies. It changes real systems. The same task can be completed correctly by different routes. Each run takes longer and costs more. An agent eval therefore checks the final state of the system, the recorded steps, the pass rate over repeated runs and the cost, where a single-answer eval checks one piece of text.

Published September 30, 2026. Editorial.

Key takeaways

  • As an illustration, if every step succeeds 95 percent of the time and the steps are independent, a 10-step task succeeds 59.87 percent of the time and a 20-step task 35.85 percent.
  • An agent's final message and the final state of the system are two separate things. Anthropic's example is a booking agent that says the flight is booked, while the test looks for the reservation in the database.
  • Two correct runs of the same task can use different steps, so a test that demands one fixed sequence of tool calls can fail a correct agent. Anthropic and Google give opposite advice on this point.
  • Agent runs cost money per task. The paper AI Agents That Matter reports a limit of USD 4 per task for one coding agent, and says a single run of the whole benchmark could cost over USD 8,000.
  • A product that makes one model call per message and takes no actions needs LLM evals only. Agent evals start when the model chooses its own steps and calls tools.

Take a refund assistant for an invented furniture shop. A customer writes that a chair arrived broken. The assistant looks up the order, reads the returns policy, issues the refund in the payment system and writes a reply. That is four actions and one message. If you test it the way you test a single model answer, you read the reply, and the reply can say "Your refund has been issued" while the payment system holds no refund at all.

An AI agent is a large language model (LLM) working in a loop: it picks a step, calls a tool, reads the result and repeats until the task is done. Testing it differs from testing one answer in four ways, and each one adds a check that a single-answer test never needed. This page is the starting point of the AI agent evals guide. The business overview is how to evaluate an AI agent before you trust it.

What an agent eval is

An eval is a repeatable test with a recorded result. An LLM eval gives a model one input and grades the one output, a method covered in the LLM evals guide. An agent eval gives the agent a task and a place to carry it out, lets it run to the end, and then grades what happened. Anthropic's engineering guide to agent evals, published on 9 January 2026, supplies the vocabulary this guide uses [1]. One attempt at a task is a trial. The complete record of a trial is a transcript, which the guide says is also called a trace or a trajectory, and which includes "outputs, tool calls, reasoning, intermediate results, and any other interactions". The outcome is "the final state in the environment at the end of the trial".

A tool is any function the agent can ask the surrounding software to run for it: a search, a database lookup, a payment request. A tool call is one such request. OpenAI's documentation defines a trace in the same way, as a record that "captures the end-to-end record of model calls, tool calls, guardrails, and handoffs for one run" [2]. A guardrail, in that sentence, is a rule in the software that blocks or checks an action, and a handoff is one agent passing the task to another. The unit under test is therefore a whole run: the task, every step, and the state of the system afterwards.

What you test for one answer and what you test for an agent

Question One model answer An agent
What is graded The text of the answer The final state of the system, plus the recorded steps
What can go wrong A wrong, unsupported or off-topic answer A wrong action, a wrong tool, wrong tool inputs, or a harmful step before a correct result
How many correct results exist Often one answer or a small set Many valid routes to the same end state
How often each case runs Once, or a few times Several times, because the same task can pass and then fail
What a run costs One model call Many model calls and tool calls
Where the test runs Anywhere: text goes in and text comes out In a closed copy of the system, reset before each run

Many steps multiply a small error rate

Anthropic's guide gives the reason in a few words: agents "use tools across many turns, modifying state in the environment and adapting as they go", so "mistakes can propagate and compound" [1]. In plain words, an early mistake is carried into later steps and grows.

The table below is an illustration, and it rests on two assumptions. Each step succeeds 95 percent of the time. The steps are independent, meaning that a failure at one step leaves the chance of failure at every other step unchanged. The chance that every step succeeds is then 0.95 multiplied by itself once for each step.

Steps in the task Working Chance that every step succeeds
1 0.95 95.00 percent
5 0.95 to the power of 5 = 0.7737809375 77.38 percent
10 0.95 to the power of 10 = 0.5987369392 59.87 percent
20 0.95 to the power of 20 = 0.3584859224 35.85 percent

A step that looks reliable when tested alone produces a 20-step task that fails 64.15 percent of the time.

Real agents do not meet the independence assumption, and the difference can raise or lower the pass rate. An agent that reads the result of each tool call can notice an error and repair it, which is why Anthropic's building guide says an agent must gain "ground truth" from the environment at each step, meaning real results such as the output of a tool call [3]. An early wrong step can also make every later step wrong. The reliable way to learn a task's pass rate is to run the whole task many times, as described in agent reliability across repeated runs.

The agent changes real systems

A model answer is text that a person reads before deciding what to do. An agent acts: it writes to a database, sends an email, moves money. Two consequences follow for testing.

First, the final message is weak evidence. Anthropic's example is a flight-booking agent that ends its transcript with "Your flight has been booked", while the outcome is whether a reservation exists in the environment's database [1]. An agent eval reads the database. The method is in outcome evals vs trajectory evals.

Second, a run can do harm during its steps, whatever its final result. HAL is a public ranking of AI agents run by a Princeton University team. The 31 authors of the paper about it, submitted to the research archive arXiv on 13 October 2025, point out that accuracy scores "assign the same score (zero) to an agent that abstains from answering, and another one that leaks a user's credit card information online in the process of solving a task" [4].

This is why agent tests run in a sandbox: a closed copy of the system where nothing the agent does reaches real data or real customers. Anthropic recommends "extensive testing in sandboxed environments, along with the appropriate guardrails" [3]. Building one is covered in test environments for agents, and tests against text planted by an attacker are in security evals for AI agents.

Different routes can both be correct

Ask an agent to refund an order and it may look up the customer first or the order first, check the policy before or after, and search the help pages at some point. All of those runs can end with the right refund.

A study titled Towards a Science of AI Agent Reliability, submitted to arXiv on 18 February 2026, tested 15 models on two benchmarks (public sets of tasks with a fixed scoring method) and ran each task 5 times. It found that agents "reliably select similar action types across runs but vary in execution order" [5]. A test that demands one exact sequence of steps will therefore mark some correct runs as failures.

The vendors disagree on what to do about it. Anthropic's guide calls checking a specific sequence of tool calls "too rigid" and says that "agents regularly find valid approaches that eval designers didn't anticipate" [1]. Google's Agent Development Kit documentation compares the steps the agent took "against the list of steps we expect the agent to have taken", and its default pass mark for that comparison is 1.0, "requiring a 100% match in the tool usage trajectory" [6]. Each is a company describing its own tools, with no measurement attached. Our position, argued on the outcome and trajectory page linked above, is to grade the end state first and to check steps only for rules that must always hold.

A run takes longer and costs more

One model answer costs one model call. An agent run costs one call per step, plus the tools. The cost then multiplies again, because each task has to be run several times: Anthropic's guide says that "because model outputs vary between runs, we run multiple trials to produce more consistent results" [1].

The paper AI Agents That Matter, submitted to arXiv on 1 July 2024 by five authors at Princeton University, gives figures. The authors of one coding agent, SWE-Agent, set a cost limit of USD 4 per task, and running it on the whole benchmark "could cost over USD 8,000 for a single evaluation run" [7]. On HumanEval, a coding benchmark of 164 problems on which each agent was run five times, the paper found two agent designs that cost over 50 percent more than the authors' own simple method and one that cost over 50 times more. It reports "no significant accuracy difference" between that simple method and the best-performing design [7]. Those results come from one benchmark and from models of the GPT-3.5 and GPT-4 period.

The paper concludes that "useful agent evaluations must control for cost" [7]. For an agent, cost is a test result: a run that reaches the right end state but uses more steps or money than its limit can still be a failed run. Setting those limits is covered in cost, time and step limits as agent evals.

What stays the same as in LLM evals

You still write the expected result before you run the test. You still need a set of cases drawn from real use, and Anthropic's guide says that "20-50 simple tasks drawn from real failures is a great start" [1]. You still choose a grader for each check, and the guide names three kinds for agents: code-based, model-based and human [1]. A code-based grader is a program that compares the result with the expected one. A model-based grader is a second model that reads the result and gives a verdict.

The guide's test for a well-written task also still applies: "A good task is one where two domain experts would independently reach the same pass/fail verdict" [1].

What changes is where each grader looks. In an agent eval, a code-based grader reads the end state and the list of tool calls, and a model-based or human grader reads the trace. Checks on a single step, such as whether the right tool was chosen with the right inputs, are covered in tool-call evals. Our reasons for keeping the grading model separate from the model that did the work are in never let the model grade its own work.

When a chatbot needs agent evals

Anthropic separates two kinds of system. Workflows "are systems where LLMs and tools are orchestrated through predefined code paths", and agents suit "open-ended problems where it's difficult or impossible to predict the required number of steps" [3].

A chatbot that makes one model call per message and uses no tools is covered by LLM evals. So is a fixed workflow in which your code decides every step and the model fills in text: test each model call as a single answer, and test the code as ordinary software.

Agent evals start at the moment the model decides what happens next. A chatbot that can look up an order, change a booking or open a support ticket is an agent for testing purposes. Ask one question of your product: could a wrong run leave a system in a different state? If it could, test the state.

How Reveneau applies this

All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. When the product is an agent, we write that suite as agent evals. Each task has an expected end state, written before the prompt. The task runs in a closed copy of the system, it runs more than once, and each run has a limit on steps and cost.

The checks that need judgment, such as whether a trace followed the plan it was given, are graded with Jev, TypeSafe AI's decision model, which returns a probability for a written question. On our own suite the run is ten times faster than with our previous language-model grader. The method for traces is in grading agent traces with Jev.

Reveneau, as a company, takes responsibility for the whole project through production and after release. If you are putting an agent into production and want it tested this way, see AI development at Reveneau or contact us.

Common questions

What is an agent eval?

An agent eval is a repeatable test of a whole agent run: the agent is given a task and a closed copy of the system, it runs to the end, and the test grades what happened. Anthropic's January 2026 guide calls one attempt a trial, the full record a transcript or trace, and the final state of the environment the outcome. The grade can come from code, from a second model or from a person.

How is an agent eval different from an LLM eval?

An agent eval grades a whole run, where an LLM eval grades one piece of text. Four things change: the agent takes many steps, it changes real systems, different routes can be correct, and each run costs more. Anthropic's guide defines the outcome as the final state in the environment at the end of the trial, so an agent eval checks that state, the recorded steps, repeated runs and cost.

Why do errors add up in an agent?

Errors add up in an agent because a task succeeds only when every step succeeds, so the chances multiply. As an illustration, if each step succeeds 95 percent of the time and the steps are independent, a 10-step task succeeds 59.87 percent of the time and a 20-step task 35.85 percent. Anthropic's guide says that mistakes can propagate and compound, meaning they carry into later steps and grow.

What is a trace in an agent eval?

A trace is the recorded sequence of model calls, tool calls and results for one run of an agent. OpenAI's documentation describes it as the end-to-end record of one run, covering model calls, tool calls, the safety rules that were applied and any passing of the task between agents. Anthropic's guide treats trace, transcript and trajectory as the same thing. Graders read the trace to check the steps, and engineers read the trace to find where a failed run went wrong.

Do I need agent evals for a simple chatbot?

A simple chatbot that makes one model call per message and uses no tools needs LLM evals only. Agent evals become necessary when the model chooses its own next step or calls a tool that reads or changes a system, such as an order lookup or a booking change. Anthropic makes the same distinction between workflows, which follow code paths fixed in advance, and agents, which handle tasks whose number of steps is unknown beforehand.

Why does an agent eval cost more to run than an LLM eval?

An agent eval costs more because one run makes many model calls and tool calls, and each task is run several times. The paper AI Agents That Matter, from July 2024, reports that the authors of one coding agent set a limit of USD 4 per task and that a single run of the whole benchmark could cost over USD 8,000. Plan the eval budget per task and per repeat.

What goes wrong if I test only the agent's final message?

Testing only the final message lets a run pass when the agent said the work was done and the system shows otherwise. Anthropic's example is a booking agent that ends with the words Your flight has been booked, while the real outcome is whether a reservation exists in the database. The final message also leaves out harmful steps taken during the run, which only the recorded steps reveal.

Can I reuse my LLM evals when I test an agent?

Most of an LLM eval method still applies to agents: expected results written first, cases drawn from real use, and a chosen grader for each check. Anthropic's guide names the same three grader types for agents, code-based, model-based and human, and keeps the same test for a well-written task, which is that two domain experts would independently reach the same pass or fail verdict. The new work is checking the end state, the steps, repeat runs and cost.

Does the 95 percent per step illustration describe real agents?

The 95 percent illustration is arithmetic under two stated assumptions, a fixed success rate per step and independent steps, and real agents do not meet either one. An agent that reads each tool result can notice and repair an error, which Anthropic's building guide describes as checking real results from the environment at each step. An early wrong step can also make every later step wrong. Measure whole runs on your own tasks.

Is a workflow with fixed steps an agent?

A workflow with fixed steps and an agent are two separate types of system, by Anthropic's definition. In a workflow, language models and tools are orchestrated through predefined code paths, so your code decides each step. In an agent, the model decides. For a fixed workflow, test each model call as a single answer and test the code as ordinary software. Add agent evals when the model starts choosing the steps.

What should I build first when I start testing an agent?

Start with a small set of tasks that each have an expected end state you can check in code, and a closed copy of the system to run them in. Anthropic's guide suggests 20 to 50 simple tasks drawn from real failures. Run each task more than once, record the trace, and add limits on steps and cost. Checks on the order of steps can come later, for the rules that must always hold.

References