AI agent evals: how to test an agent that takes actions
An AI agent is a language model working in a loop: it picks a step, calls a tool such as a search or a database, reads the result and repeats until the task is done. Because it takes many steps and acts on real systems, a test of its final answer covers a small part of what can go wrong. This guide is the method. Grade the end state and the steps, test each tool call, repeat every case because one pass proves little, set limits on cost, time and steps, run the tests in a closed copy of the system, test for attacks, and turn production runs into new test cases.
Published September 30, 2026. Editorial.
Key takeaways
- As an illustration, if each step succeeds 95 percent of the time and the steps are independent, a 10-step task succeeds 59.87 percent of the time. Test whole runs, because step scores multiply.
- Grade the end state of the system first, in code. Check the steps only for rules that must always hold, such as a confirmation before a payment.
- In the tau-bench paper of June 2024, gpt-4o completed 61.2 percent of retail tasks on one run, and the measure of 8 passes out of 8 fell below 25 percent. Repeat every case.
- A vendor's strict mode covers the format of a tool call. In tau-bench, gpt-4o usually chose the right type of call and filled in one or more arguments incorrectly, so test the values.
- Give every task a limit on cost, time and steps. No primary source we read recommends a number, so set limits from measured runs of your own tasks.
- Attack results depend on the attacker: on the AgentDojo test environment, NIST's Center for AI Standards and Innovation raised the success rate from 11 percent to 81 percent with new attacks against one model.
- Turn confirmed production failures into permanent test cases, and remove personal data first. OpenTelemetry's conventions say trace recording should leave out message content by default.
In June 2024 the authors of the tau-bench benchmark gave the best model they tested, gpt-4o, a set of 115 customer-service tasks for a simulated retail shop. On a single run it completed 61.2 percent of them. The authors then measured a stricter result: the same task completed 8 times out of 8. That measure, averaged over the tasks, was below 25 percent [1].
That one experiment shows most of the points in this guide. The agent acted on a database, so the test read the database instead of the agent's reply. The agent's typical error was a wrong input to a tool. One run proved little, so every task was repeated. Each task had a limit of 30 actions. The customer was played by a second model, and the follow-up paper, tau2-bench, measured how often its simulated customer made errors. The sections below take those points one at a time, and each links to a page with the full method.
What this guide covers and who it is for
This guide is for engineers and product owners who are putting an AI agent into production. An agent is a large language model (LLM) working in a loop: it picks a step, calls a tool such as a search, a database or a payment service, reads the result and repeats until the task is done. An eval is a repeatable test with a recorded result.
This is the method reference. The business overview is how to evaluate an AI agent before you trust it. How to read a vendor's score on a public agent benchmark, a published set of tasks with a fixed scoring method, is covered in agent benchmarks explained.
Waiting for a better model is a weak plan for reliability. A study of 15 models on two benchmarks, submitted to the research archive arXiv on 18 February 2026, found that "recent capability gains have only yielded small improvements in reliability" [2]. The table lists what an agent eval suite tests.
| What to test | The check | How it is decided |
|---|---|---|
| End state | The system after the run matches a goal state written in advance | Code |
| Steps | Rules on the record of the run: forbidden tools absent, approvals first | Code |
| One tool call | Right tool, valid inputs, correct values, errors handled | Mostly code |
| Repeat runs | The share of tasks that pass on every run | Code, over several runs |
| Cost, time and steps | Each stays within a limit set per task | Code |
| Attacks | Planted text leaves the agent's actions unchanged | Code, plus people writing new attacks |
| Replies and other checks that need judgment | A written question about the reply or the run | A person or a second model |
Testing an agent differs from testing one answer in four ways
A test of one model answer grades a piece of text. A test of an agent grades a whole run, and four things change. The agent takes many steps, so a small error rate at each step multiplies. It changes real systems, so its final message and the final state can disagree. Different routes can be correct. Every run costs more.
The first difference is the least obvious, so here is the arithmetic. It is an illustration with two stated assumptions: each step succeeds 95 percent of the time, and the steps are independent. The chance that all 10 steps of a task succeed is 0.95 to the power of 10, which is 0.5987, or 59.87 percent. For 20 steps it is 0.95 to the power of 20, which is 0.3585, or 35.85 percent. Anthropic's guide to agent evals, published on 9 January 2026, says the same in words: agents "use tools across many turns", so "mistakes can propagate and compound" [3].
A chatbot that makes one model call per message and uses no tools needs the methods in the LLM evals guide and none of the checks here. The full comparison, with a table of what to test in each case, is in why testing an agent is different from testing one model answer.
Grade the result first, and the steps for rules that must hold
An outcome eval reads the system after the run: the record in the database, the file, the reply. Anthropic's example is a flight-booking agent that says "Your flight has been booked", while the outcome is whether a reservation exists in the database [3]. A trajectory eval compares the agent's sequence of tool calls with a reference sequence written in advance. Google Cloud's documentation defines the comparisons: an exact match needs "the exact same tool calls in the exact same order", an in-order match allows extra calls, an any-order match ignores the order, and precision and recall return fractions [4].
Grade the outcome first, in code where you can. A strict step match fails a correct agent that added one lookup, and Anthropic calls that type of check "too rigid" [3]. Keep step checks for rules that hold in every acceptable run: the delete tool is never called, and a confirmation comes before a payment.
Those rules matter because a pass or fail on the outcome hides what happened during the run. HAL is a public ranking of AI agents run by a Princeton University team. The 31 authors of the paper about it, submitted to arXiv on 13 October 2025, note that accuracy scores "assign the same score (zero) to an agent that abstains from answering, and another one that leaks a user's credit card information online in the process of solving a task" [5]. The match types, with a worked example of each, are in outcome evals vs trajectory evals.
Test each tool call: the tool, the inputs and the error handling
A tool call is one request from the agent to the software around it: run this search, read this record, make this payment. It is the smallest unit you can test, and five checks cover it: whether a tool was needed at all, whether the right tool was chosen, whether the arguments (the inputs passed to the tool) fit the format the tool expects, whether the values are correct, and whether the agent dealt with an error the tool returned.
Format is the part a vendor feature can cover. OpenAI's documentation says that its strict mode "will ensure function calls reliably adhere to the function schema" [6], where the schema is the written description of the inputs a tool accepts. That is the vendor's statement about its own feature, and it concerns format. Correct values are a separate matter. In the tau-bench paper, the gpt-4o agent "usually makes the right type of tool call(s) but fills in one or more arguments incorrectly", and it made 0.46 tool calls per retail task with a user, product or order identifier that did not exist [1].
Most of these checks are a program comparing the call the agent made with the call you expected, so they are cheap to run on every change. Each check, with an example failure, is in tool-call evals.
Repeat every case, because one passing run proves little
Run the same agent on the same task twice and it can pass once and fail once. The tau-bench authors defined a measure for this, pass^k: "the chance that all k i.i.d. task trials are successful, averaged across tasks" [1]. A trial is one attempt, and i.i.d. means the attempts are independent and made under the same conditions. The figures at the top of this page are that measure for gpt-4o on retail tasks: 61.2 percent for one run, and below 25 percent for 8 runs. The models date from mid 2024. The paper gives no exact value for the second figure: its abstract and its results section print "<25%", and its introduction gives 25 percent and marks it as approximate.
An illustration shows why the fall is so large. If a task passes 60 percent of single runs and the runs are independent, the chance of 8 passes in a row is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent. A real task set mixes easier and harder tasks, so its own figure has to be measured.
A released agent handles the same type of request many times, so the number that matters for release is the share of tasks that pass on every run. A single-run pass rate overstates it. How many repeats to run, and what to report, are in agent reliability across repeated runs.
Put limits on cost, time and steps, and fail runs that exceed them
A run that reaches the right end state can still be a failed run if it used more steps or money than the limit set for the task. Measure four things per task: the cost of model usage, the time from start to finish, the number of steps and the number of tool calls. Give each a limit, and fail the run when it goes over. A step limit also stops a loop, which is an agent repeating the same step without progress.
The paper AI Agents That Matter, submitted to arXiv on 1 July 2024, argues that "useful agent evaluations must control for cost", because accuracy "can be improved by scientifically meaningless methods such as retrying" [7].
We found no primary source that recommends a step or cost budget as a number. The sources show the limits researchers used in their own tests: at most 30 agent actions per task in tau-bench [1], and USD 4 per task for the coding agent SWE-Agent, as reported in AI Agents That Matter [7]. Set your own limits from measured runs of your own tasks. The method, along with the research group METR's measurements of how long a task current models can complete, is in cost, time and step limits as agent evals.
Run the tests in a closed copy of the system
An agent eval must never touch production. Three building blocks make that possible. A sandbox is a closed copy of the system where nothing the agent does reaches real data, and it is reset before every run. Fake tools return recorded or scripted results in place of the real service. A simulated user is a second model that plays the customer in tasks with several turns.
Each one distorts the result in its own way. Start with the reset. Anthropic's guide says each trial should start "from a clean environment", because leftover files and stored data from one run can cause failures in the next that are unrelated to the agent [3]. By Anthropic's own account, in some internal evals its model gained "an unfair advantage on some tasks by examining the git history from previous trials", meaning the record of earlier changes to the files [3].
The simulated user is a source of error too. The tau2-bench paper of June 2025 checked its own user simulator and recorded an error rate of 40 percent on the retail tasks and 47 percent on the airline tasks, with 12 and 13 percent being errors that prevented the task from being completed [8]. A failed run in a simulated conversation has to be read before it is counted against the agent. How to build each block, and what each one distorts, is in test environments for agents.
Turn production runs into test cases
Every production run leaves a trace: the recorded sequence of model calls, tool calls and results. A test set stays current through one repeated routine. Record traces, sample them, have a person review the failures, turn each confirmed failure into a permanent case with its expected outcome, remove personal data, and re-run the set on every change. Langfuse's documentation describes the same workflow for its own product: "select production traces where the application did not perform as expected. Then you let an expert add the expected output to test new versions of your application on the same data" [9].
Two cautions apply. The shared format for recording traces, OpenTelemetry's conventions for generative AI, carried the status "Development" when we read it on 30 September 2026, so check the status again before you depend on a field name [10]. And traces hold what customers typed. The same conventions call model instructions, user messages and model outputs "sensitive". They say that the software which records traces "SHOULD NOT capture them by default" and should let a team switch the capture on by choice [10].
Decide what a trace stores and how long you keep it. A lawyer should confirm which data protection rules apply to you. The routine, step by step, is in turning production traces into eval cases, and grading traces with a decision model is in grading agent traces with Jev.
Test the agent against attacks
An agent reads text from web pages, emails and documents, and some of that text can be written by an attacker to look like an instruction. This is prompt injection. OWASP, a community security project, describes the indirect form as occurring "when an LLM accepts input from external sources, such as websites or files" [11].
A security eval asks four things: whether planted text changes the agent's actions, whether the agent can be made to send data out, whether it takes an irreversible action without confirmation, and whether it stays inside its permissions. Published results show how much the answer depends on who wrote the attack. On AgentDojo, a test environment with 97 tasks and 629 security test cases published in June 2024, the paper's own attacks succeeded "against the best performing agents in less than 25% of cases" [12]. In January 2025, the Center for AI Standards and Innovation at NIST, the US government's standards body, wrote new attacks for the same test environment against one model, and the success rate rose "from 11% for the strongest baseline attack to 81% for the strongest new attack" [13]. The baseline attack is the strongest one that was already in the test set.
Passing these tests lowers the measured rate for the attacks you tried, and a stronger attack may still succeed. Defences outside the model, such as narrow permissions and a confirmation step before a payment, are still needed. The tests are in security evals for AI agents, and a way to review what each tool allows is in how to audit an AI agent's tool permissions.
The order to build an agent eval suite in
Build in this order:
- A closed test system with a reset.
- A set of tasks, each with a goal state. Anthropic's guide says that "20-50 simple tasks drawn from real failures is a great start" [3].
- Outcome checks in code, run several times per task.
- Limits on steps, time and cost.
- The narrow step rules and the attack cases.
- The link to production, so that every confirmed failure becomes a case.
Once the task set exists, the same guide notes that response time, model usage, cost per task and error rates can be tracked on it from one change to the next [3].
Evals are one method among several. Anthropic's guide lists the others a team needs: "production monitoring, user feedback, A/B testing, manual transcript review, and systematic human evaluation" [3]. A/B testing means giving two versions to separate groups of users and comparing the results. An eval suite tells you whether a change is safe to release, and those other methods tell you what happened after release.
How Reveneau tests agents
All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an agent, that suite follows the method in this guide. We write the goal state of every task before the prompt. We run each task in a closed copy of the system, several times, with limits on steps, time and cost. We write step rules only for what the specification says must always hold, and we add attack cases for every tool that reads outside text.
The checks that need judgment are graded with Jev, TypeSafe AI's decision model, which returns a probability for a written question. On our own suite the run is ten times faster than with our previous language-model grader.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To have an agent built and tested this way, see AI development at Reveneau or contact us.
Explore the guide
Start here
Why testing an agent is different from testing one model answer
An agent eval is a repeatable test of an AI agent, a language model that works in a loop: it picks a step, calls a tool such as a search or a database, reads the result and repeats until the task is done. Testing an agent differs from testing one model answer in four ways. The agent takes many steps, so a small error rate at each step multiplies. It changes real systems. The same task can be completed correctly by different routes. Each run takes longer and costs more. An agent eval therefore checks the final state of the system, the recorded steps, the pass rate over repeated runs and the cost, where a single-answer eval checks one piece of text.
Outcome evals vs trajectory evals: grade the result and the steps
An outcome eval checks the final state of the system after an agent has run: the record in the database, the file, the reply that was sent. A trajectory eval checks the sequence of steps the agent took against a sequence written in advance. Grade the outcome first, because it matches what the user wanted and a record or a file can be checked in code. Grade the steps only for rules that must always hold, such as never calling the delete tool or always confirming before a payment. A strict step-by-step match fails correct agents that found another valid route, and an outcome-only grade hides harmful steps, so a production suite uses both, each for its own job.
What to measure
Tool-call evals: did the agent pick the right tool and the right arguments
A tool call is one request an AI agent sends to another program: the name of a tool and the values to pass to it. A tool-call eval grades that single step with five checks. Was a tool needed, was the right tool named, do the arguments fit the tool's written definition, are the argument values correct, and did the agent respond sensibly when the tool returned an error. Four of the five can be decided by code, which makes them cheap to run on every change. The Berkeley Function Calling Leaderboard checks the name and the arguments this way, and the tau-bench paper found in 2024 that its best agent usually chose the right tool and filled in one or more arguments incorrectly.
Agent reliability: why one passing run proves little
An agent that passes a task once can fail the same task on the next run. The tau-bench paper of June 2024 measured this with pass^k, the chance that all k runs of a task pass: its best agent passed 61.2 percent of single retail runs and had a pass^8 below 25 percent. One passing run is therefore weak evidence. Run every case several times, at least 5 before a release decision, and report two numbers: the pass rate across all runs and the share of tasks that passed every run. A user of the released product gets one run, so the second number describes what users experience.
Cost, time and step limits as agent evals
A run that returns the right result can still be a failed run if it took too many steps or cost too much. An agent eval should record model usage cost, elapsed time, steps and tool calls for every run, and fail any run that exceeds a written limit on one of them. The 2024 paper AI Agents That Matter found coding agents of similar accuracy whose cost differed by almost two orders of magnitude (two orders is a factor of one hundred), so accuracy reported without cost cannot be compared. No source gives a recommended limit as a number: set each one from your own passing runs. METR's time horizon figures, quoted here with their dates, show how long a task models complete.
Build and run the tests
Test environments for agents: sandboxes, fake tools and simulated users
An agent eval must never run against production, because an agent under test takes real actions: it issues refunds, sends email and deletes records. A test environment has three building blocks. A sandbox is a closed copy of the system where nothing the agent does reaches real data, and it is reset to the same starting state before every run. Fake tools return recorded or scripted results in place of services you cannot copy. A simulated user is a language model that plays the customer in a conversation that takes several turns. Each block distorts the result in a known way, and this page explains how to measure and limit each distortion.
Turning production traces into eval cases
Every production run of an agent leaves a trace, the record of each model call, tool call and result in that run. A trace becomes useful for testing when a person confirms that it shows a failure and writes down the outcome that should have happened. The method has six steps: record traces in a standard format, sample them, have a person review the failures, write each confirmed failure as a permanent case with its expected outcome, remove personal data before the case is stored, and re-run the whole set on every change. OpenTelemetry's naming conventions for generative AI are an open format for the first step, and their status reads Development as of 30 September 2026.
Security evals for AI agents: prompt injection and unsafe actions
An agent reads text it did not write: web pages, emails, documents and the replies of its tools. Some of that text can be written by an attacker to look like an instruction, which is called prompt injection. A security eval plants such text in a test environment and measures how often the agent does what the attacker wanted. Published tests show that the measured rate depends on who writes the attack. In a January 2025 test by the Center for AI Standards and Innovation at NIST, the US government's standards body, the strongest earlier attack succeeded 11 percent of the time against one agent, and the strongest new attack 81 percent. A passing result lowers the measured rate, and attacks nobody has tried remain untested.
Common questions
What does this guide to AI agent evals cover?
The AI agent evals guide covers the method for testing an AI agent that takes actions: grading the end state and the steps, testing each tool call, repeating runs, limiting cost, time and steps, building a closed test system, testing for attacks, and turning production runs into test cases. It opens with the tau-bench result of June 2024, in which gpt-4o completed 61.2 percent of retail tasks on one run.
Where should a team start with agent evals?
A team should start with a closed test system that can be reset, and a small set of tasks that each have a goal state written in advance. Anthropic's guide to agent evals says that 20 to 50 simple tasks drawn from real failures is a great start. Add outcome checks in code, run each task several times, then add limits, step rules and attack cases.
Will a newer model make agent evals unnecessary?
Agent evals stay necessary with a newer model, on the evidence available. A study of 15 models on two benchmarks, submitted to arXiv on 18 February 2026, found that recent capability gains have only yielded small improvements in reliability. A model that solves harder tasks can still pass and fail the same task on different runs, so each new model has to be run against your own task set.
Are evals enough to know that an agent is ready for release?
Evals are one of several methods and are best used together with the others. Anthropic's guide lists the others a team needs: production monitoring, user feedback, A/B testing, manual transcript review and systematic human evaluation. An eval suite tells you whether a change is safe to release, and the other methods tell you what happened after release. Confirmed production failures then go back into the suite as new cases.
Which agent checks can code decide and which need a person?
Code can decide most agent checks: the end state against a goal state, rules on the recorded steps, the format and values of a tool call, repeat-run pass rates, and limits on cost, time and steps. A person or a second model is needed for checks that need judgment, such as whether a reply explained the action accurately. Anthropic's guide names the three grader types as code-based, model-based and human.
How is this guide different from the overview of agent evaluation in the agents in production guide?
The AI agent evals guide is the method reference for people building the tests, and the overview in the agents in production guide is written for a reader deciding what to fund. Each section here states a check, gives the source for its figures, and links to a full page. For example, the section on repeat runs gives the tau-bench definition of pass^k and its 61.2 percent single-run figure.
What does an agent eval suite measure besides pass and fail?
An agent eval suite also measures cost, time, steps and tool calls per task, and how those change from one version to the next. Anthropic's guide notes that response time, model usage, cost per task and error rates can be tracked on a fixed set of tasks. The paper AI Agents That Matter argues that useful agent evaluations must control for cost, because accuracy alone can be raised by retrying.
How far does an agent's pass rate fall when the same task is repeated?
An agent's pass rate can fall by more than half when every repeat must pass. In the tau-bench paper of June 2024, gpt-4o completed 61.2 percent of retail tasks on a single run, and the measure of 8 passes out of 8 was below 25 percent. As an illustration, a task that passes 60 percent of independent single runs passes 8 in a row 1.68 percent of the time.
Is there a standard budget for an agent's steps or cost?
A standard budget for an agent's steps or cost is missing from every primary source we read. The sources show limits that researchers chose for their own tests: at most 30 agent actions per task in tau-bench, and USD 4 per task for the coding agent SWE-Agent, as reported in AI Agents That Matter. Measure runs of your own tasks and set each limit from those measurements.
Does passing a security eval mean an agent is safe?
Passing a security eval lowers the measured rate for the attacks that were tried, and a stronger attack may still succeed. On the AgentDojo test environment, the paper's own attacks succeeded in less than 25 percent of cases, and in January 2025 NIST's Center for AI Standards and Innovation raised the rate from 11 percent to 81 percent with new attacks against one model. Keep defences outside the model, such as narrow permissions.
Can a simulated user be trusted in agent tests?
A simulated user is useful for tasks with several turns, and its own mistakes have to be allowed for. The tau2-bench paper of June 2025 recorded an error rate of 40 percent for its user simulator on the retail tasks and 47 percent on the airline tasks, with 12 and 13 percent being errors that prevented task completion. Read a failed simulated conversation before counting it against the agent.
How often should an agent eval suite run?
An agent eval suite should run on every change to the prompt, the tools, the model or the code, and it should grow whenever production shows a new failure. Langfuse's documentation describes the workflow: select production traces where the application did not perform as expected, then have an expert add the expected output so that new versions are tested on the same data.
References
- [1] Yao, Shinn, Razavi and Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains, arXiv 2406.12045 (17 June 2024): gpt-4o scored 61.2 on retail tasks (115 tasks) for one run and pass^8 below 25% (mid-2024 models); definition of pass^k; at most 30 agent actions per task; the gpt-4o agent usually makes the right type of tool call and fills in arguments incorrectly; 0.46 tool calls with non-existent IDs per retail task.
- [2] Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan, Towards a Science of AI Agent Reliability, arXiv 2602.16666 (18 February 2026): 15 models on two benchmarks; recent capability gains have only yielded small improvements in reliability.
- [3] Anthropic, Demystifying evals for AI agents (9 January 2026): mistakes can propagate and compound; the flight-booking example; checking a fixed sequence of tool calls is too rigid; each trial starts from a clean environment; Anthropic's own account of a model examining the git history from earlier trials; 20-50 simple tasks as a start; what can be tracked on a static bank of tasks; the other methods a team needs beside evals.
- [4] Google Cloud, Evaluate Gen AI agents (page marked Preview, read 30 September 2026): definitions of exact, in-order and any-order match, precision and recall.
- [5] Kapoor, Stroebl, Kirgis and 28 co-authors, Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation, arXiv 2510.11977 (13 October 2025): accuracy assigns the same score (zero) to an agent that abstains and to one that leaks a user's credit card information.
- [6] OpenAI API documentation, Function calling (no date on page, read 30 September 2026): OpenAI's statement that strict mode ensures function calls adhere to the function schema.
- [7] Kapoor, Stroebl, Siegel, Nadgir and Narayanan, AI Agents That Matter, arXiv 2407.01502 (1 July 2024): useful agent evaluations must control for cost; accuracy can be improved by retrying; SWE-Agent cost limit of USD 4 per task.
- [8] Barres, Dong, Ray, Si and Narasimhan, tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment, arXiv 2506.07982 (9 June 2025): user simulator error rates of 40% (retail) and 47% (airline), with 12% and 13% critical errors.
- [9] Langfuse documentation, Datasets (no date on page, read 30 September 2026): the workflow of selecting production traces that did not perform as expected and having an expert add the expected output.
- [10] OpenTelemetry, Semantic conventions for generative client AI spans (main branch, read 30 September 2026): status Development as read on 30 September 2026; model instructions, user messages and model outputs are sensitive and SHOULD NOT be captured by default.
- [11] OWASP Gen AI Security Project, LLM01:2025 Prompt Injection (2025 edition): definition of indirect prompt injection.
- [12] Debenedetti, Zhang, Balunovic, Beurer-Kellner, Fischer and Tramer, AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents, arXiv 2406.13352 (19 June 2024): 97 tasks and 629 security test cases; the paper's attacks succeed against the best performing agents in less than 25% of cases.
- [13] NIST, Center for AI Standards and Innovation, Technical Blog: Strengthening AI Agent Hijacking Evaluations (17 January 2025): new attacks raised the success rate on AgentDojo from 11% to 81% against one model.
Related reading
An AI agent that passes a test once can fail it on the next run
The same agent on the same task passes on some runs and fails on others, and the tau-bench paper of 2024 measured by how much. Run each test case several times and report the share of tasks that passed every run.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
How to audit an AI agent's tool permissions
An agent that can call more tools than its task needs is a permanent risk, and most teams find that out by reading an incident report instead of an audit.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
What a decision model changes about agent safety
A safety check that answers in under half a second and costs a fraction of a cent can run on every message and every tool call. Here is what that changes for an agent in front of real users, and the one thing it does not change.