Start here

Outcome evals vs trajectory evals: grade the result and the steps

An outcome eval checks the final state of the system after an agent has run: the record in the database, the file, the reply that was sent. A trajectory eval checks the sequence of steps the agent took against a sequence written in advance. Grade the outcome first, because it matches what the user wanted and a record or a file can be checked in code. Grade the steps only for rules that must always hold, such as never calling the delete tool or always confirming before a payment. A strict step-by-step match fails correct agents that found another valid route, and an outcome-only grade hides harmful steps, so a production suite uses both, each for its own job.

Published September 30, 2026. Editorial.

Key takeaways

  • An outcome eval reads the system after the run. The tau-bench benchmark grades this way: it compares the database at the end of a conversation with a goal state written by a person.
  • Google Cloud's documentation defines exact, in-order and any-order matches plus precision and recall, and every one of them needs a reference list of tool calls written in advance.
  • An exact match fails any run with an extra or reordered step. A February 2026 study of 15 models found that agents choose similar types of action across runs and vary the order.
  • An outcome score gives the same zero to an agent that declines the task and to one that leaks a card number, say the authors of the paper on HAL, a public ranking of AI agents. The recorded steps of the run separate the two.
  • Grade the end state first. Add step checks for rules that must always hold, written as narrow checks on the trace: a forbidden tool is absent, a confirmation comes before a payment.

HAL is a public ranking of AI agents run by a Princeton University team. In October 2025 the 31 authors of the paper about it, published on the research archive arXiv, reported 21,730 agent runs across 9 models and 9 benchmarks (public sets of tasks with a fixed scoring method), and they released the logs [1]. Reading those logs, they found agents that searched a public website for the benchmark itself "instead of solving a task", and agents "misusing credit cards in flight booking tasks" [1]. A pass or fail on the final result records no such behaviour.

That is the argument for checking the steps. The argument against: an agent that completes a refund correctly by a route nobody wrote down is marked as failed by a test that expected one fixed route. This page, part of the AI agent evals guide, explains both checks and gives our position. The vocabulary comes from why testing an agent is different from testing one model answer.

What an outcome eval checks

An outcome eval reads the system after the agent has finished. Anthropic's guide to agent evals defines the outcome as "the final state in the environment at the end of the trial", where a trial is one attempt at a task. Its example is a flight-booking agent: the agent may say the flight is booked, and the outcome is whether a reservation exists in the environment's database [2]. Google Cloud's documentation calls this final response evaluation [3].

Outcomes come in three forms, and each has its own check:

  • A record, such as a row in a database, a ticket or a calendar entry. Query the system and compare the fields.
  • A file, such as a report, a code change or a spreadsheet. Open it and test its content.
  • A reply to a person. Check that the reply says what was done and that what it says matches the record.

Code can decide the first two. The third often needs judgment, and the three kinds of LLM eval covers how text from a large language model (LLM) is graded.

How to check the final state

Two public benchmarks show the pattern. The tau-bench benchmark, published in June 2024, has an agent serve simulated customers of a retail shop and an airline, and its grader "compares the database state at the end of a conversation with the annotated goal state" [4], which is the correct end state as the benchmark's authors wrote it down. OSWorld, a benchmark of 369 computer tasks published in April 2024, gives each task "a detailed initial state setup configuration and a custom execution-based evaluation script" [5], meaning a program that inspects the system after the run. Both follow the same four steps:

  1. Put the test system in a known starting state.
  2. Write down the goal state before the run: which records must exist, and with which values.
  3. Run the agent.
  4. Compare the whole end state with the goal state in code.

Comparing the whole state, as tau-bench does, also fails a run that made the right refund and then changed a second order nobody asked it to touch. A check that looks only for the expected record would pass that run.

The closed test system that these checks run in is described in test environments for agents.

What a trajectory eval checks

Google's Agent Development Kit documentation defines a trajectory as "the sequence of steps taken to reach the solution" [6]. Google Cloud's definition is narrower: trajectory evaluation looks at "the path (sequence of tool calls) the agent took to reach the final response" [3]. A tool call is one request from the agent to the surrounding software, such as a lookup or a payment.

A trajectory eval needs a reference trajectory, which is the list of tool calls a person expects for the task, written in advance. Google Cloud states that a reference trajectory "is required for all metrics except trajectory_single_tool_use" [3]. The match type decides how strictly the agent's list is compared with the reference.

The match types, with one example each

Take one invented task for the furniture shop: refund a broken chair. The reference trajectory has three tool calls in this order: look_up_order, check_policy, issue_refund. The table applies Google Cloud's definitions [3] to invented runs.

Match type The run scores 1 when Example
Exact match It has "the exact same tool calls in the exact same order" as the reference look_up_order, check_policy, issue_refund scores 1. Add one search and the score is 0.
In-order match It has every reference call in the same order, and it "may also have extra tool calls" look_up_order, search_help, check_policy, issue_refund scores 1.
Any-order match It has every reference call, in any order, with extra calls allowed check_policy, look_up_order, issue_refund scores 1. The in-order match gives this run 0.

Two more measures return a fraction. For precision, Google Cloud's instruction is to count the actions in the run that also appear in the reference and divide by the total number of actions in the run. For recall, count the reference actions that appear in the run and divide by the number of actions in the reference [3]. Two invented runs show the difference:

Run Precision Recall
look_up_order, search_help, check_policy, issue_refund 3 of 4 calls are in the reference: 0.75 3 of 3 reference calls are in the run: 1.00
look_up_order, issue_refund 2 of 2 calls are in the reference: 1.00 2 of 3 reference calls are in the run: 0.67

Low precision means the agent did work outside the reference. Low recall means the agent left out a required call: the second run skipped the policy check.

Names differ by product. Google's Agent Development Kit calls the three sequence matches EXACT, IN_ORDER and ANY_ORDER, uses EXACT by default, scores each call to the agent 1.0 on a match and 0.0 otherwise, and reports the average [7]. LangChain's documentation for its agentevals package lists four modes: strict, unordered, subset and superset [8]. Each vendor is describing its own product, so read the definition in the tool you use before you compare two scores.

When a strict match fails a correct agent

In the furniture example, an agent that searches the help pages before issuing the correct refund leaves the database in the goal state and still scores 0 on an exact match, because of one extra search.

Anthropic's guide describes checking that an agent followed a specific sequence of tool calls in the right order as "too rigid", because "agents regularly find valid approaches that eval designers didn't anticipate", and it advises grading what the agent produced in preference to the path it took [2]. One study supports the premise. Towards a Science of AI Agent Reliability, submitted to arXiv on 18 February 2026, ran 15 models on two benchmarks, 5 times per task, and found that agents "reliably select similar action types across runs but vary in execution order" [9].

Google's products default to the strict check. The Agent Development Kit's tool trajectory score defaults to 1.0, "requiring a 100% match in the tool usage trajectory" [6], and its documentation advises EXACT "when you need to enforce a specific tool execution path" [7]. That advice fits tasks where the order is itself the requirement. On any other task, an exact match measures how closely the agent copied the person who wrote the reference.

What an outcome-only grade hides

A position paper by eleven authors, submitted to arXiv on 8 May 2026 under the title Log analysis is necessary for credible evaluation of AI agents, opens with the problem: "Agent benchmarks typically report only final outcomes: pass or fail" [10]. One of the three threats it names is that "capability scores may conceal dangerous or catastrophic actions taken by the agent" [10]. We read the paper's abstract only.

The HAL authors put that threat in one sentence: accuracy scores "assign the same score (zero) to an agent that abstains from answering, and another one that leaks a user's credit card information online in the process of solving a task" [1].

An outcome tells you that a run failed, and the trace, the full record of the run, tells you where. OpenAI's documentation makes this argument for grading traces, which it says "provide more data to better understand why an agent succeeds or fails" [11]. Turning failed traces into permanent test cases is covered in turning production traces into eval cases.

What each check catches and what it misses

Check What it catches What it misses
End state compared in code Wrong, missing or extra changes to records and files Harmful steps that left no record, and the cause of a failure
Reply graded by a person or a model A reply that misstates what was done Whether the action happened
Exact match Any skipped, added or reordered step Whether a different route was also correct
In-order match A skipped step, or required steps in the wrong order Wasted extra steps
Precision and recall Wasted steps and missing steps, as fractions The order of steps, and whether the task was completed
A narrow rule on the trace A forbidden tool call, or a missing confirmation Anything the rule did not name

Whether one call used the right tool with the right inputs is a separate check, covered in tool-call evals.

Grade the outcome first, then the rules that must hold

Grade the outcome first. It matches what the user asked for, it stays valid when the model or the prompt changes because it does not depend on the route, and code can decide it when the outcome is a record.

Then add trajectory checks and keep them narrow. A rule that must always hold is true in every acceptable run, whatever the route. Three examples:

  • The agent never calls the delete tool. Check that the tool is absent from the trace. Google Cloud's single tool use metric checks whether a named tool is used anywhere in the run [3], so for this rule the pass is that the tool is absent.
  • The agent confirms with the user before it pays. Write a two-call reference, confirm_with_user then issue_payment, and use an in-order match, which allows any other calls around them.
  • The agent stays within a step limit. Count the calls in the trace, as described in cost, time and step limits as agent evals.

The OWASP Top 10 for LLM Applications 2025, written by a community security project, lists "Require human approval for high-risk actions" among its defences against prompt injection, which is text that changes a model's behaviour in ways nobody intended, including text planted by an attacker [12]. The second rule above checks that approval on every run. The attack tests are in security evals for AI agents, and the business guide covers how to set an agent's permissions.

Use a full reference trajectory only where the order is the requirement, such as a regulated process with fixed steps, and prefer in-order to exact there, so that an extra lookup still passes.

Who grades each check

Anthropic's guide says agent evals "typically combine three types of graders: code-based, model-based, and human" [2]. End-state comparison and trace rules are code. Whether a reply explained the refund accurately needs a person or a model, and LangChain's documentation says that a model used as judge of a trajectory "requires an LLM call and is less deterministic" [8], meaning the same trace can receive different grades on different runs. Keep that judge separate from the model that did the work, for the reasons in never let the model grade its own work. Grading traces with a decision model is covered in grading agent traces with Jev, and how many times to repeat each check is in agent reliability across repeated runs.

How Reveneau applies this

All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an agent, we write the goal state of each task before the prompt, and the outcome check compares the end state of a closed test system with that goal in code. Step checks are limited to the rules the specification says must always hold, so a correct run by an unexpected route still passes.

The checks that need judgment, such as whether a reply describes the action accurately, are graded with Jev, TypeSafe AI's decision model, which returns a probability for a written question. On our own suite the run is ten times faster than with our previous language-model grader.

Reveneau, as a company, takes responsibility for the whole project through production and after release. If you want an agent built against outcome checks that exist before the first prompt, see AI development at Reveneau or contact us.

Best for

  • Outcome checks: any task whose result is a record, a file or a state that code can read
  • Narrow trace rules: actions that must never happen, and approvals that must come first
  • In-order match: a regulated process where the sequence of steps is itself the requirement

Avoid if

  • Exact match as the default pass mark: it fails correct runs that add or reorder one step
  • Outcome-only grading for an agent that can move money or read private data
  • A model as judge for a check that code can decide

Check before you decide

  • The goal state was written before the agent was run
  • The outcome check compares the whole end state, so unwanted extra changes fail the run
  • Each match type means what you expect in the tool you use, because names differ by vendor

Common questions

What is an outcome eval for an AI agent?

An outcome eval is a test that reads the system after the agent has finished and checks that the required result exists. Anthropic's guide defines the outcome as the final state in the environment at the end of the trial, and gives the example of a booking agent whose outcome is whether a reservation exists in the database. The agent's closing message is left out of the check, because the message can be wrong.

What is a trajectory eval?

A trajectory eval is a test that compares the sequence of steps an agent took with a reference sequence written in advance. Google Cloud's documentation describes the trajectory as the path, meaning the sequence of tool calls, that the agent took to reach the final response. The comparison can be strict or loose: an exact match, an in-order match, an any-order match, or a fraction such as precision or recall.

Should I grade the steps or the result?

Grade the result first, and grade the steps only for rules that must always hold. The result is what the user asked for and a record can be checked in code. Anthropic's guide calls checking a fixed sequence of tool calls too rigid, because agents regularly find valid approaches the eval designer did not anticipate. Keep step checks for rules such as never calling the delete tool or always confirming before a payment.

What is the difference between an exact match and an in-order match?

An exact match passes only when the agent made the same tool calls as the reference in the same order with nothing added, while an in-order match also passes when extra calls appear between them. Google Cloud's documentation says an in-order run contains all the tool calls from the reference trajectory in the same order and may also have extra tool calls. One added search fails an exact match and passes an in-order match.

How do I check the final state after an agent run?

Check the final state in four steps: put the test system in a known starting state, write the goal state before the run, run the agent, then compare the whole end state with the goal in code. The tau-bench benchmark works this way, comparing the database state at the end of a conversation with an annotated goal state. Comparing the whole state also fails runs that made unwanted extra changes.

What do precision and recall mean for an agent's steps?

Precision is the share of the agent's tool calls that appear in the reference, and recall is the share of the reference calls that appear in the agent's run. In an invented example with three reference calls, a run that adds one search has precision 0.75 and recall 1.00, by Google Cloud's definitions. Low precision means wasted work. Low recall means a required step was skipped.

Why would a strict step match fail an agent that did the task correctly?

A strict step match fails a correct agent because it compares the route with one reference route, and several routes can be valid. A study submitted to arXiv in February 2026 tested 15 models on two benchmarks and found that agents reliably select similar action types across runs but vary in execution order. An agent that adds one harmless lookup, or swaps two steps, scores 0 on an exact match.

What does an outcome-only score hide?

An outcome-only score hides what the agent did before it reached the result. The authors of the paper on HAL, a public ranking of AI agents run by a Princeton University team, reported 21,730 agent runs in October 2025. They note that accuracy assigns the same zero to an agent that abstains from answering and to one that leaks a user's credit card information. The same score also hides shortcuts, such as an agent searching online for the benchmark in place of solving the task.

Do Google, Anthropic and LangChain use the same names for trajectory checks?

Google and LangChain use different names for trajectory checks, and Anthropic gives different advice. Google's Agent Development Kit has EXACT, IN_ORDER and ANY_ORDER and uses EXACT by default. LangChain's agentevals package has strict, unordered, subset and superset modes. Anthropic's guide advises grading what the agent produced in preference to its path. Read each tool's own definition before comparing scores from two products.

Which step rules should an agent eval always include?

An agent eval should include a step rule for every action that must never happen and every approval that must come first. Examples are that the delete tool is never called and that the user confirms before a payment. The OWASP Top 10 for LLM Applications 2025 lists requiring human approval for high-risk actions as a defence, and an in-order check on the trace shows the approval came before the action.

Does a trajectory eval cost more to run than an outcome eval?

A trajectory eval decided by code needs no model call, because a program compares two lists of tool calls, so it adds no grading cost to an outcome eval. The cost rises when a model grades the trace. LangChain's documentation says that a model used as judge of a trajectory requires an LLM call and is less deterministic, so use a model only for questions that code cannot decide.

References