AI agent evals: how to test an agent that takes actions / What to measure
Agent reliability: why one passing run proves little
An agent that passes a task once can fail the same task on the next run. The tau-bench paper of June 2024 measured this with pass^k, the chance that all k runs of a task pass: its best agent passed 61.2 percent of single retail runs and had a pass^8 below 25 percent. One passing run is therefore weak evidence. Run every case several times, at least 5 before a release decision, and report two numbers: the pass rate across all runs and the share of tasks that passed every run. A user of the released product gets one run, so the second number describes what users experience.
Published September 30, 2026. Editorial.
Key takeaways
- Pass^k is the chance that an agent passes all k runs of the same task, averaged across tasks. The tau-bench paper introduced the measure in June 2024.
- In that paper the best agent, built on gpt-4o, passed 61.2 percent of single retail runs, and its pass^8 was below 25 percent. The models are from mid 2024.
- As an illustration: if a task passes 60 percent of independent single runs, the chance of eight passes in a row is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent.
- Report two numbers: the pass rate across all runs, and the share of tasks that passed every run, with the number of runs stated beside it.
- Three runs are too few to trust a pass: a task with an 80 percent pass rate passes three independent runs out of three with a probability of 0.512.
The tau-bench paper, submitted in June 2024, tested agents that act as customer service assistants for a simulated shop and a simulated airline. Its best agent used gpt-4o and passed 61.2 percent of single runs on the 115 retail tasks [1]. The authors also measured how often the agent passes all eight of eight runs of the same task. For that agent on retail the figure was below 25 percent [1]. The paper gives no exact value: its abstract and its results section print "<25%", and its introduction gives 25 percent and marks it as approximate.
An agent in production meets the same request many times, from many customers, and each customer gets one run. A test set that runs every case once reports the first number and gives no information about the second. This page explains the measure the tau-bench authors introduced, the arithmetic behind it, how many repeats to run, and which two numbers to report.
What pass^k measures
The tau-bench authors define pass^k (read "pass hat k") as "the chance that all k i.i.d. task trials are successful, averaged across tasks" [1]. A trial is one attempt at a task. The letters i.i.d. stand for independent and identically distributed: each attempt starts fresh, under the same conditions, with no knowledge of the others. So pass^1 is the ordinary single-run pass rate, and pass^8 is the chance of eight passes in eight runs.
The paper contrasts it with a second measure, pass@k: "the chance that at least one out of k i.i.d. task trials is successful" [1]. The two move in opposite directions as k grows. Anthropic's guide to agent evals says that at k=1 the two are "identical (both equal the per-trial success rate)", and that by k=10 "pass@k approaches 100% while pass^k falls to 0%" [2].
Which one fits depends on how the agent is used. Where a person looks at several attempts and keeps the best, such as several drafts of a piece of code, pass@k describes what that person gets. Where each user receives a single run and the agent acts on it, pass^k describes what the user gets. A released agent that changes records belongs to the second case.
The arithmetic of eight passes in a row
Here is an illustration with invented numbers. Suppose an agent passes a task on 60 percent of single runs, every run is independent, and the rate is the same on every run. The chance of eight passes in a row is 0.6 multiplied by itself eight times:
0.6^8 = 0.01679616, which is 1.68 percent.
Anthropic's guide gives a smaller example of the same sum: at a 75 percent single-run rate, three passes in a row have a probability of 0.75^3, which is 0.421875, or 42 percent as the guide rounds it [2]. That figure is arithmetic, and nobody measured it.
The table shows the same calculation for three single-run rates. Every cell is the rate raised to the power of the number of runs, under the same two assumptions.
| Single-run pass rate | 3 runs in a row | 5 runs in a row | 8 runs in a row |
|---|---|---|---|
| 90 percent | 72.9 percent | 59.049 percent | 43.046721 percent |
| 80 percent | 51.2 percent | 32.768 percent | 16.777216 percent |
| 60 percent | 21.6 percent | 7.776 percent | 1.679616 percent |
A 90 percent agent sounds dependable. Under these assumptions it completes eight runs of one task without a failure less than half the time.
What the tau-bench paper reported
The paper's main results are single-run rates. Each task was run at least 3 times, each run was limited to 30 agent actions, and the agent model's temperature was 0.0 [1]. Temperature is the setting that controls how much randomness a model uses when it chooses its words.
| Agent model | Retail pass^1 | Airline pass^1 |
|---|---|---|
| gpt-4o | 61.2 | 35.2 |
| gpt-4-turbo | 57.7 | 32.4 |
| claude-3-opus | 44.2 | 34.7 |
| gpt-3.5-turbo | 20.0 | 10.8 |
These are the paper's figures for models available in mid 2024, on 115 retail tasks and 50 airline tasks [1]. They record that test and describe no current model. For the eight-run measure the paper gives one statement: for the gpt-4o agent, "pass^8 drops to <25%" [1].
The paper's figure, which its introduction gives as 25 percent and marks as approximate, is higher than the 1.68 percent in the illustration, and the reason affects how you read your own results. The illustration assumed every task has the same pass rate. Real tasks differ. Take three invented tasks with single-run rates of 100, 60 and 20 percent. Their average is 60 percent. Their eight-run rates are 1, 0.01679616 and 0.00000256, and the average of those three is 0.33893291, or 33.9 percent. Two agents with the same 60 percent average can therefore have eight-run results of 1.68 percent and 33.9 percent, depending on whether the failures are spread across all tasks or concentrated in a few.
METR, a research group that measures how long a task AI agents can complete, describes this pattern in a later agent. Its time horizons page, last updated on 8 May 2026, says that on tasks that take a human expert 90 minutes to 3 hours, a GPT-5 agent succeeds every time on a third of the tasks, fails every time on another third, and "sometimes succeeds and sometimes fails on the remaining third of tasks" [3]. METR's own wording for each of the first two shares is "around one-third", so the split is approximate.
The measure also changes with the setting. In the tau2-bench paper of June 2025, gpt-4.1 had a pass^1 of 74 percent on retail, 56 percent on airline and 34 percent on a new telecom setting. For a second model, claude-3.7-sonnet, the single-run rate on telecom was 49 percent, which the authors say matches its airline rate, and they report that as k increases "the pass^k scores decline more rapidly for telecom compared to airline" [4].
Why the same agent passes and fails on the same input
A flaky agent test is one that passes on some runs and fails on others with the same input. Four causes produce that result, and they call for different responses.
The model's output varies. Anthropic's guide gives the reason for repeating runs in one line: "Because model outputs vary between runs, we run multiple trials to produce more consistent results" [2]. Setting the temperature to zero reduces the variation, and some of it remains. A February 2026 paper by Stephan Rabanser and five co-authors ran each task 5 times at temperature zero on 15 models across two benchmarks (public sets of tasks with a fixed scoring method), GAIA and tau-bench. It found that "outcome consistency remains low across all models", which the authors explain as "agents that can solve a task often fail to do so consistently" [5]. The same paper reports that agents choose similar types of action from run to run and change the order of them [5].
The simulated user varies. Many agent tests use a second model to play the customer. The tau2-bench authors checked that simulated user and recorded an error rate of 40 percent in retail and 47 percent in airline, with 12 and 13 percent being critical errors that prevent the task from being completed. In the telecom setting the rate was 16 percent, with 6 percent critical [4]. Some failed runs are caused by the test itself.
State is left over from an earlier run. Anthropic's guide says each trial should start from a clean environment, because leftover files or stored data can cause failures that come from the test setup and say nothing about the agent [2]. Resetting between runs is covered in test environments for agents.
The task is faulty. Of the 40 failed gpt-4o retail runs the tau-bench authors examined, they traced 4 to faults in the instruction given to the simulated user [1]. The maintainers of tau2-bench list "Task Quality (75+ fixes)" in their repository notes, and state that results from versions before 1.0.1 "are not comparable with >= 1.0.1" [6].
The first cause is the agent's own behaviour, and it is the thing you are measuring. The other three are faults in the test, and they should be fixed before the number is trusted. Reading the recorded steps of each mixed task shows which cause applies.
How many times to run each case
No source used for this page gives a rule for the number of repeats. The published tests show a range of practice:
- tau-bench ran at least 3 trials per task for its main results [1].
- The 2026 reliability paper ran each task 5 times [5].
- The 2024 paper "AI Agents That Matter" ran each agent five times on the 164 problems of the HumanEval coding benchmark [7].
- METR launches 6 independent runs for each task [3].
The arithmetic gives a reason to prefer the upper end. Under the independence assumption, a task the agent passes 80 percent of the time passes 3 runs out of 3 with a probability of 0.8^3 = 0.512. Three repeats show that task as fully passing more often than they show its failure. At 5 repeats the probability is 0.32768, and at 8 it is 0.16777216.
Our position is a judgment, since no source measured it: run every case at least 5 times before a release decision, and 8 or more times for cases where one failure is costly, such as a payment or a deletion. On small day-to-day changes, fewer repeats on a subset of cases reduce the cost.
Repeats multiply cost directly: 5 repeats cost 5 times one pass. The tau-bench authors reported that one retail task cost 0.38 US dollars for the agent and 0.23 for the simulated user in 2024 [1], and cost, time and step limits works through the sum. How many separate cases the set needs is a different question, answered in how many test cases an LLM eval needs.
What to report
Report two numbers and state k beside the second.
- The pass rate across all runs. Count every run of every task and divide the passes by the total. This is pass^1. It tells you how often a single attempt works.
- The share of tasks that passed every run. This is pass^k for the k you ran. It tells you how many tasks a user can rely on.
Add a third item for the engineers: the list of mixed tasks, the ones that passed on some runs and failed on others. The next fixes are in that list.
Leave best-of-k figures out of the report. The authors of "AI Agents That Matter" point out that accuracy "can be improved by scientifically meaningless methods such as retrying" [7]. A best-of-k figure is pass@k, and it rises with every extra attempt even when the agent is unchanged.
The released product needs the second number because of what a failure means there. In a test, a failed run is a line in a report. In production it is a customer whose refund went to the wrong order. The reliability paper's opening claim is that "compressing agent behavior into a single success metric obscures critical operational flaws" [5].
What reliability is enough for production
The threshold differs from agent to agent. It depends on two facts about the task: what one failure costs, and whether a person checks the result before it takes effect. An agent that drafts a reply for a staff member to approve can be released at a lower all-runs rate than an agent that issues refunds unattended.
Higher reliability also shortens the task an agent can be given. METR's paper on task length reports that the length of task models complete at 80 percent reliability is 4 to 6 times shorter than the length they complete at 50 percent [8], and METR's page says that measuring a 99 percent figure "would require many more tasks" [3].
Set the threshold for each class of task in writing before the runs, so that nobody changes the rule after seeing the result. Using evals to decide whether an AI feature is ready to release covers that decision, and how to make an AI agent reliable enough to release is the business overview. Failures inside a single step are the subject of tool-call evals, and the full method is in the AI agent evals guide. A related post asks can you trust an AI agent with real work yet?.
How Reveneau applies this
At Reveneau, all code is written by AI, and every change must pass a large eval suite written from the specification before the code. For an agent, we write each case's pass condition and its required all-runs rate from the specification before the prompt exists, and we run every case several times before a release decision. We report the single-run pass rate and the share of tasks that passed every run, with the number of runs stated. Repeats multiply the number of judged checks, so the speed of grading affects how many repeats are practical. We grade those checks with Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To discuss reliability targets for an agent you are planning, see AI development at Reveneau or contact us.
Best for
- Agents that act without a person checking each result
- Release decisions, where each user will get a single run
- Comparing two versions of an agent whose single-run rates are close
Avoid if
- The test environment keeps data from one run to the next, which makes the repeats depend on each other
- The report would show the best of several runs as the result
Check before you decide
- The number of runs is stated next to every all-runs figure
- Each run starts from a reset environment
- Mixed tasks are listed and their recorded steps were read
Common questions
What is pass^k in agent evaluation?
Pass^k is the chance that an agent passes all k runs of the same task, averaged across tasks. The tau-bench paper introduced it in June 2024 and defines it as the chance that all k independent trials of a task are successful. At k equal to 1 it is the ordinary pass rate. At higher k it shows how often an agent completes a task every time, which is what a user who gets one run depends on.
What is the difference between pass^k and pass@k?
Pass^k is the chance that every one of k runs passes, and pass@k is the chance that at least one of k runs passes. Anthropic's guide to agent evals notes that the two are identical at k equal to 1 and move apart as k grows: by 10 runs pass@k approaches 100 percent while pass^k falls to 0. Use pass@k when a person keeps the best attempt, and pass^k when each user gets one.
How many times should I run each agent test?
Run each agent test at least 5 times before a release decision, and 8 or more times where one failure is costly. That is Reveneau's position, and no source gives a measured rule. Published tests use at least 3 runs per task (tau-bench), 5 (the 2026 reliability paper) and 6 (METR). With 3 runs, a task that passes 80 percent of the time shows three passes with a probability of 0.512.
Why does my agent pass a test sometimes and fail it other times?
An agent passes and fails the same test because its model's output varies between runs, even when the model's randomness setting, called temperature, is at zero. A 2026 paper that ran each task 5 times at temperature zero on 15 models found that agents that can solve a task often fail to do so consistently. Three other causes are in the test: a simulated user that makes errors, data left over from an earlier run, and a faulty task.
What is a flaky agent test?
A flaky agent test is a test that passes on some runs and fails on others with the same input. The cause can be the agent, whose output varies from run to run, or the test, through a simulated user's errors, leftover data or a faulty task. The tau2-bench authors measured simulated user error rates of 40 percent in retail and 47 percent in airline conversations, so read the recorded steps before deciding the agent is at fault.
What reliability is enough to release an agent to production?
The reliability needed to release an agent depends on what one failure costs and whether a person checks the result before it takes effect. Set a required all-runs pass rate for each class of task in writing before the runs. METR's paper reports that the task length models complete at 80 percent reliability is 4 to 6 times shorter than at 50 percent, so a higher target means giving the agent shorter tasks.
How do I calculate the chance of eight passes in a row?
Multiply the single-run pass rate by itself eight times, assuming the runs are independent and the rate is the same on each. For a 60 percent rate the result is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent. That calculation is an illustration. Results across a set of tasks are higher when tasks differ in difficulty, because a task that passes every time counts fully at any number of runs.
What did the tau-bench paper find about repeated runs?
The tau-bench paper found that its best agent in June 2024, built on gpt-4o, passed 61.2 percent of single runs on 115 retail tasks, and that its pass^8, the chance of eight passes in eight runs averaged across tasks, was below 25 percent. The paper states the eight-run figure only as below 25 percent. The models are from mid 2024, so the figures describe that test and no agent available today.
Which numbers should an agent eval report show?
An agent eval report should show two numbers: the pass rate across all runs, and the share of tasks that passed every run, with the number of runs stated. A list of the tasks that passed on some runs and failed on others belongs beside them. Leave out best-of-several figures: the authors of AI Agents That Matter note that accuracy can be raised by retrying, which leaves the agent unchanged.
Does repeating every test make agent evals too expensive?
Repeating tests multiplies the cost of the run by the number of repeats, so 5 repeats cost 5 times one pass. Control it by running the full repeats before a release decision and fewer repeats on a subset of cases for small daily changes. The tau-bench authors reported 0.38 US dollars per retail task for the agent and 0.23 for the simulated user in 2024, which shows the scale for one benchmark.
Do I need repeated runs if my agent gave the same result three times?
Yes, keep the repeats. Three identical passes are weak evidence: a task with an 80 percent pass rate produces three passes in a row with a probability of 0.512, assuming independent runs. The 2026 reliability paper found low consistency of outcomes across all 15 models it tested with the randomness setting at zero, so an agent that agreed with itself on a few runs can still vary on later ones.
What should I do with a task that passes on some runs and fails on others?
Read the recorded steps of the failing runs of that task first, then decide whether the cause is the agent or the test. The tau-bench authors traced 4 of 40 examined failures to faults in the instructions given to the simulated user. Fix the test faults, treat the remaining mixed tasks as the work list for the agent, and re-run the full set with the same number of repeats.
References
- [1] Yao, Shinn, Razavi and Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, submitted 17 June 2024): definitions of pass^k and pass@k; Table 2 pass^1 figures for four models on 115 retail and 50 airline tasks; pass^8 below 25 percent for gpt-4o on retail; at least 3 trials per task, 30 actions, agent temperature 0.0; 4 of 40 examined failures traced to the user instruction; 0.38 and 0.23 US dollars per retail task.
- [2] Anthropic, Demystifying evals for AI agents (9 January 2026): why trials are repeated, how pass@k and pass^k move apart by k=10, the worked figure of 0.75 cubed, and the rule that each trial starts from a clean environment.
- [3] METR, Task-Completion Time Horizons of Frontier AI Models, live page and data file (last updated 8 May 2026, read 30 September 2026): a GPT-5 agent on tasks of 90 minutes to 3 hours succeeds every time on a third, fails every time on a third and is mixed on the rest; 6 independent runs per task; a 99 percent horizon would require many more tasks.
- [4] Barres, Dong, Ray, Si and Narasimhan, tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment (arXiv, submitted 9 June 2025): gpt-4.1 pass^1 of 74, 56 and 34 percent on retail, airline and telecom; claude-3.7-sonnet at 49 percent on telecom, with pass^k declining faster on telecom than on airline; simulated user error rates of 40, 47 and 16 percent with 12, 13 and 6 percent critical.
- [5] Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan, Towards a Science of AI Agent Reliability (arXiv, submitted 18 February 2026): each task run 5 times at temperature zero on 15 models across GAIA and tau-bench; outcome consistency low across all models; order of actions varies between runs.
- [6] Sierra Research, tau2-bench repository README (notice for version 1.0.1 dated July 2026, read 30 September 2026): more than 75 task quality fixes, and results before version 1.0.1 are not comparable with later ones.
- [7] Kapoor, Stroebl, Siegel, Nadgir and Narayanan, AI Agents That Matter (arXiv, submitted 1 July 2024): each agent run five times on the 164 HumanEval problems; accuracy can be improved by retrying.
- [8] Kwa and 25 co-authors (METR), Measuring AI Ability to Complete Long Software Tasks (arXiv, first version 18 March 2025, fourth version 10 July 2026): models' 80 percent time horizons are 4 to 6 times shorter than their 50 percent horizons.
Related reading
An AI agent that passes a test once can fail it on the next run
The same agent on the same task passes on some runs and fails on others, and the tau-bench paper of 2024 measured by how much. Run each test case several times and report the share of tasks that passed every run.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
What founders get wrong about AI agents
An impressive agent demo and a reliable agent are two different things. Most of the work, and most of the risk, is in the final step before production, which nobody shows in the demo.
How autonomous are AI coding agents, really?
Engineers at the leading AI labs now say a model writes one hundred percent of their code. Read the quotes closely and a person is still involved in every one of them. Here is what the 2026 evidence supports, and what it does not.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
More in What to measure
Tool-call evals: did the agent pick the right tool and the right arguments
A tool call is one request an AI agent sends to another program: the name of a tool and the values to pass to it. A tool-call eval grades that single step with five checks. Was a tool needed, was the right tool named, do the arguments fit the tool's written definition, are the argument values correct, and did the agent respond sensibly when the tool returned an error. Four of the five can be decided by code, which makes them cheap to run on every change. The Berkeley Function Calling Leaderboard checks the name and the arguments this way, and the tau-bench paper found in 2024 that its best agent usually chose the right tool and filled in one or more arguments incorrectly.
Cost, time and step limits as agent evals
A run that returns the right result can still be a failed run if it took too many steps or cost too much. An agent eval should record model usage cost, elapsed time, steps and tool calls for every run, and fail any run that exceeds a written limit on one of them. The 2024 paper AI Agents That Matter found coding agents of similar accuracy whose cost differed by almost two orders of magnitude (two orders is a factor of one hundred), so accuracy reported without cost cannot be compared. No source gives a recommended limit as a number: set each one from your own passing runs. METR's time horizon figures, quoted here with their dates, show how long a task models complete.