Build and run the tests

Test environments for agents: sandboxes, fake tools and simulated users

An agent eval must never run against production, because an agent under test takes real actions: it issues refunds, sends email and deletes records. A test environment has three building blocks. A sandbox is a closed copy of the system where nothing the agent does reaches real data, and it is reset to the same starting state before every run. Fake tools return recorded or scripted results in place of services you cannot copy. A simulated user is a language model that plays the customer in a conversation that takes several turns. Each block distorts the result in a known way, and this page explains how to measure and limit each distortion.

Published September 30, 2026. Editorial.

Key takeaways

  • Reset the environment before every run. Anthropic's January 2026 eval guide says leftover files and stored data cause failures unrelated to the agent, and reports that in some internal evals Claude gained an unfair advantage by reading the file history of earlier runs.
  • Use the real tool code against a test copy of the database wherever a tool changes data, so a script can check the final state. Use scripted fake tools for outside services such as payments and email, and script their error replies too.
  • A simulated user makes mistakes of its own. The tau2-bench paper (June 2025) recorded user simulator error rates of 40 percent in its retail part, 47 percent in airline and 16 percent in telecom, so read a sample of simulated turns before trusting a score.
  • Check which isolation your tool gives you. A sandbox is a closed copy of the system. Inspect, the UK AI Security Institute's eval tool, lists eight sandbox types, and its own documentation describes the one named local as the local file system with no sandbox.
  • WebArena's 14.41 percent (2023) and OSWorld's 12.24 percent (2024) are the original papers' results for the models of those years. Both figures are historical baselines, and this page states no current score.

Take an invented furniture shop with a support agent that can issue refunds. The team writes 40 test tasks and runs each one 5 times, which is 200 runs (40 x 5 = 200). If those runs go to the shop's real systems, every refund task that passes moves real money. The shop and its numbers are an illustration. The problem applies to any agent that acts: the test itself takes the actions.

So an agent eval needs a place to run that is separate from production (the systems real customers use). Anthropic's guide to building agents recommends "extensive testing in sandboxed environments, along with the appropriate guardrails" [2]. A sandbox is a closed copy of the system where nothing the agent does reaches real data, and guardrails are limits placed on what the agent may do. This page describes the three building blocks of a test environment and what each one distorts. It belongs to the AI agent evals guide.

What a test environment has to do

A test environment has three jobs. It keeps every action inside: a refund issued during a test changes a test record and nothing else. It starts every run from the same state. And it lets a script (a short program) read the final state. Anthropic's eval guide separates what the agent says from what happened: "A flight-booking agent might say 'Your flight has been booked' at the end of the transcript, but the outcome is whether a reservation exists in the environment's SQL database" [1]. (A SQL database is a store of records that programs query in a language called SQL.) The page on outcome evals and trajectory evals explains how that check is written.

The sandbox: a closed copy that actions cannot leave

There are two common ways to build a sandbox. A container is a self-contained copy of a program and its files that runs apart from the rest of the machine. A virtual machine is a complete simulated computer running inside a real one. The OSWorld benchmark, which tests agents that operate a desktop computer, runs every task in a virtual machine and gives the reason: "Virtual machine offers a safe isolated environment and prevents the agent resulting in irreversible damaging effect on the real host machine" [4].

Inspect is an eval tool published by the UK AI Security Institute, with its code open for anyone to read. As read on 30 September 2026, its documentation lists eight sandbox types: two built in and six available as separate packages [3]. One of the two built-in types is named local, and the documentation's own description of it is "Local file system (no sandbox)" [3]. Read the description of every setting before the first run, because a setting listed among the sandbox types can give no isolation at all.

Check network access next. An agent that can reach the internet from inside the sandbox can still call a real service. The example configuration in the Inspect documentation sets the network mode to none, which gives the container no network [3].

Reset the state before every run

Anthropic's eval guide states the rule: "Each trial should be 'isolated' by starting from a clean environment. Unnecessary shared state between runs (leftover files, cached data, resource exhaustion) can cause correlated failures due to infrastructure flakiness rather than agent performance" [1]. A trial is one attempt at one task. In plain words: if run 7 leaves a file behind and run 8 fails because of it, that failure says nothing about the agent.

Shared state can also help the agent. By Anthropic's own account, "in some internal evals we observed Claude gaining an unfair advantage on some tasks by examining the git history from previous trials" [1]. Git history is the saved record of earlier changes to a set of files.

Three practices make a reset dependable:

  • Keep the starting data in one file under version control (the system that records every change to a file), and build the environment from that file each time. OSWorld gives each task its own starting state configuration and its own checking script [4].
  • Destroy the environment after the run and build a new one, instead of deleting the rows a run added.
  • Fix the date and time the tools report, and any identifier the system would generate at random.

You can prove that a reset works. Run the same task twice with a scripted agent that takes fixed steps, and compare the two final states. They must be identical. A related post explains why unreliable tests are worse when agents write them.

Fake tools or real ones

A tool is a function the agent can call, such as an order lookup. In a test, each tool can be one of three things.

The real tool code, connected to a test copy of the data. The agent's call runs production's code against a database inside the sandbox. tau-bench, a customer service benchmark published in June 2024, is built in a similar way: its tools act on a database. Its retail part has 500 users, 50 products and 1,000 orders in a database, with 7 tools that write to it and 8 that only read [6]. The benchmark grades a run by a rule that needs a real database: it "compares the database state at the end of a conversation with the annotated goal state" [6].

A recorded fake. The tool returns a reply that was saved from an earlier real call. The reply has the exact shape production gives, and it is fixed: an agent that sends different arguments (the values passed to a tool) still receives the old reply.

A scripted fake. The tool returns whatever the test author wrote, including error replies: a timeout (no reply within the allowed time), a refusal, an empty result.

Our position: use real tool code against test data for every tool that changes your own records, because that is what makes the final state checkable. Use scripted fakes for outside services you cannot copy, such as payments and email, and write their error replies as well as their successful ones. Keep recorded fakes for read-only lookups where the exact shape of the reply matters. The checks to run on each individual call are in tool-call evals.

Simulated users for tasks that take several turns

Many agent tasks are conversations, and a fixed list of user messages cannot cover them. Google's Agent Development Kit documentation gives the example: "if the agent needs the user to supply two values to perform a task, it may ask for those values one at a time or both at once" [8]. That documentation describes a scenario for its own simulator in three parts: a starting prompt, a conversation plan and a user persona, which is a description of the user's traits [8].

A simulated user is a language model that plays the customer. The tau-bench authors used one: "We use a language model (gpt-4-0613) to simulate a human user interacting with the agent" [6]. Two details of its setup are worth copying. The simulated user could not see the agent's tool calls, as a real customer could not [6]. And the two models ran at different values of temperature, the setting that controls how much a model's output varies: 0.0 for the agent and 1.0 for the user [6].

The simulated user makes errors too

The tau2-bench paper (June 2025), whose authors include one of the tau-bench authors, checked the simulator itself. It recorded a user simulator error rate of 40 percent in its retail part and 47 percent in its airline part, with 12 and 13 percent being critical errors that prevented the task from being completed. In its newer telecom part the rate was 16 percent, with 6 percent critical [7]. The first tau-bench paper found a related fault: of 40 failed gpt-4o retail runs it examined, 4 were traced to faults in the instruction given to the simulated user and 36 to the agent [6].

A simulated user also costs money. In tau-bench the agent cost USD 0.38 per task and the user simulation USD 0.23 per task [6]. The simulated user was therefore 0.23 / (0.38 + 0.23) = 37.7 percent of the spending on model use, at the 2024 prices of those models.

Read a sample of simulated user turns from every test run. Label each failed run as an agent fault or a simulator fault, and report the two counts separately. Run each task more than once, because the simulated user is a second source of variation. Agent reliability across repeated runs explains how many repeats to run.

What each building block gives you and what it distorts

Building block What it gives you What it distorts
Sandbox, reset before every run Actions that stay inside the test, and runs that can be compared Test data is smaller and more regular than production data
Real tool code on test data A final state a script can check The delays and failures of the real service are missing
Recorded fake tool Replies with the exact shape production gives The reply is fixed, so new arguments receive an old answer
Scripted fake tool Error replies whenever the test needs them The replies are the author's belief about how the service behaves
Simulated user Conversations that take several turns The simulator makes its own errors: 16 to 47 percent in tau2-bench [7]

Every row has a distortion, so a score from a test environment is a score for that environment.

Two public environments and what their figures mean

Two research projects show what a full environment looks like. WebArena is a set of working websites of four kinds, including a shop and a discussion forum [5]. OSWorld is a set of desktop computer tasks run inside virtual machines [4].

Fact WebArena [5] OSWorld [4]
First published July 2023 April 2024
Tasks 812, written from 241 templates 369
Best model in the paper 14.41 percent (a GPT-4 agent) 12.24 percent
People in the paper 78.24 percent Over 72.36 percent

Read those percentages with their dates. Each one is the original paper's result for the models of 2023 or 2024, and this page states no current score for either benchmark. WebArena's human figure came from five computer science graduate students working on one task from each of 170 templates [5]. OSWorld limited every run to 15 steps and 30 minutes [4]; see cost, time and step limits. What these benchmarks measure is covered in agent benchmarks explained.

Your own environment can be far smaller. It needs the same properties at the size of your own agent: its tools, a starting data file, and a checking script per task.

The environment has faults of its own

A test environment is software, and it contains mistakes. The maintainers of tau2-bench say so in the project's own README (its introduction file). A notice dated July 2026 for version 1.0.1 lists "Task Quality (75+ fixes)" and states that "results produced with tau2-bench < 1.0.1 are not comparable with >= 1.0.1" [9].

Treat your environment the same way. Give it a version number. When you repair a task or a checking script, run the previous version of the agent again so that both results come from the same environment.

What a test environment cannot show

A sandbox contains what its authors thought to put in it, and real customers write messages nobody scripted. Anthropic's guide lists the other methods a team needs next to evals: "production monitoring, user feedback, A/B testing, manual transcript review, and systematic human evaluation" [1]. (An A/B test gives two versions to two groups of users and compares the results.)

When production shows a failure the environment never produced, the failure becomes a new test task: see turning production traces into eval cases and offline evals vs online evals. The same sandbox is where attack tests run, which security evals for AI agents covers.

How Reveneau builds test environments

At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an agent, that order means the test environment comes first: we write the starting data file and the checking script for each task before we write the agent's instructions. We reset the environment before every run, and we run the real tool code against test data wherever a tool changes records.

The judged checks in our suite are graded by Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. A faster run leaves time to repeat every task.

We use AI instead of hiring more engineers, so a build takes a small team and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To have an agent built together with its test environment, see our AI development service or contact us.

Best for

  • Any agent with a tool that changes records, sends messages or moves money
  • Tasks that take several turns of conversation with a user
  • Teams that need to compare two versions of an agent on the same starting state

Avoid if

  • The setting you chose gives no isolation, such as the Inspect type named local
  • You plan to delete test rows after each run instead of rebuilding the environment
  • Nobody will read a sample of the simulated user's turns

Check before you decide

  • Two runs of a fixed-step scripted agent end in identical final states
  • The sandbox has no network route to a real service
  • Each failed run is labelled as an agent fault, a simulator fault or an environment fault

Common questions

What is a sandbox for agent testing?

A sandbox for agent testing is a closed copy of the system in which the agent can read, write and delete while nothing it does reaches real data or real people. It is usually built as a container or a virtual machine. The OSWorld paper gives the reason for its virtual machines: a safe isolated environment that prevents irreversible damage to the real host machine.

How do I test an AI agent without touching production?

Run the agent inside a sandbox (a closed copy of the system) with the network switched off, connect its tools to a test copy of the data, and replace outside services such as payments and email with scripted fake tools. Rebuild the environment from one starting data file before every run. The example configuration in the Inspect documentation from the UK AI Security Institute sets the network mode to none, which gives the container no network.

What is a simulated user in an agent eval?

A simulated user in an agent eval is a language model that plays the customer, so that a task needing several turns of conversation can be tested without a person typing. The tau-bench benchmark used the model gpt-4-0613 for this and did not let the simulated user see the agent's tool calls. Google's Agent Development Kit documentation describes a scenario for its simulator as a starting prompt, a conversation plan and a user persona.

Should I use fake tools or real tools when testing an agent?

Use the real tool code against test data for every tool that changes your own records, because a script can then check the final state of the database, which is how tau-bench grades a run. Use scripted fake tools for outside services you cannot copy, and write their error replies as well. Keep recorded fake tools for read-only lookups where the exact shape of the reply matters.

How do I reset state between agent test runs?

Reset state between agent test runs by keeping the starting data in one file, building the environment from that file before each run, and destroying the environment afterwards. Fix the date, the time and any randomly generated identifier. Anthropic's eval guide says each trial should start from a clean environment, because leftover files and stored data cause failures that are unrelated to the agent.

How much does a simulated user add to the cost of an agent eval?

In the tau-bench paper of June 2024, the agent cost USD 0.38 per task and the user simulation USD 0.23 per task, so the simulated user was 37.7 percent of the spending on model use (0.23 divided by 0.61). Those are 2024 prices for the models in that paper. The share in your own eval depends on which model plays the user and how long the conversations run.

How often does a simulated user make mistakes?

The tau2-bench paper of June 2025 measured the simulated user itself and recorded an error rate of 40 percent in its retail part, 47 percent in airline and 16 percent in telecom, with 12, 13 and 6 percent being critical errors that prevented the task from being completed. Read a sample of simulated turns from every test run and count simulator faults separately from agent faults.

What goes wrong when agent test runs share state?

Two things go wrong when agent test runs share state. A file or record left by one run can make the next run fail for a reason unrelated to the agent. Shared state can also help the agent unfairly: Anthropic reports that in some internal evals its model Claude gained an advantage on some tasks by examining the git history left by previous trials.

What is the difference between a recorded fake tool and a scripted fake tool?

A recorded fake tool returns a reply saved from an earlier real call, so the reply has the exact shape production gives and never changes. A scripted fake tool returns what the test author wrote, which lets you test error replies such as a timeout or a refusal. The recorded kind gives an old answer to new arguments, and the scripted kind reflects the author's belief about the service.

Do I need a test environment as large as WebArena or OSWorld?

No. WebArena has 812 tasks across four kinds of website and OSWorld has 369 desktop computer tasks, because both are research benchmarks meant to compare many agents. Your own agent needs the same properties at its own size: its tools connected to test data, one starting data file, and a checking script for each task. Start small, and make the reset correct before adding tasks.

Does a test environment replace monitoring an agent in production?

No. A test environment contains only what its authors put in it, and real users, real data and real outside services behave in ways it does not reproduce. Anthropic's eval guide lists production monitoring, user feedback, A/B testing, manual transcript review and systematic human evaluation as the other methods a team needs beside evals. Failures found in production should become new test tasks in the environment.

What should I do after the agent test environment is built?

After the agent test environment is built, prove the reset works by running a scripted agent with fixed steps twice and comparing the two final states. Then give the environment a version number and write that number next to every result. The maintainers of tau2-bench state that results from before version 1.0.1 are not comparable with later ones, in a notice that lists 75+ task quality fixes.

References

More in Build and run the tests

Turning production traces into eval cases

Every production run of an agent leaves a trace, the record of each model call, tool call and result in that run. A trace becomes useful for testing when a person confirms that it shows a failure and writes down the outcome that should have happened. The method has six steps: record traces in a standard format, sample them, have a person review the failures, write each confirmed failure as a permanent case with its expected outcome, remove personal data before the case is stored, and re-run the whole set on every change. OpenTelemetry's naming conventions for generative AI are an open format for the first step, and their status reads Development as of 30 September 2026.

Security evals for AI agents: prompt injection and unsafe actions

An agent reads text it did not write: web pages, emails, documents and the replies of its tools. Some of that text can be written by an attacker to look like an instruction, which is called prompt injection. A security eval plants such text in a test environment and measures how often the agent does what the attacker wanted. Published tests show that the measured rate depends on who writes the attack. In a January 2025 test by the Center for AI Standards and Innovation at NIST, the US government's standards body, the strongest earlier attack succeeded 11 percent of the time against one agent, and the strongest new attack 81 percent. A passing result lowers the measured rate, and attacks nobody has tried remain untested.