An AI agent that passes a test once can fail it on the next run

In June 2024 four researchers published a test called tau-bench. It gave AI agents 115 customer service tasks for a simulated shop. An AI agent is a language model working in a loop: it picks a step, calls a tool such as a database lookup, reads the result and repeats until the task is done. The best agent in the paper used the gpt-4o model and completed 61.2 percent of the retail tasks on a single run.
The authors then measured the same agent on the same tasks a second way: how often does it complete a task on all eight of eight runs? Averaged across the tasks, the answer was below 25 percent. The paper prints that figure only as "<25%" and gives no exact value.
Both figures belong to 2024: the models are from mid 2024, and the numbers describe no current model. The gap between the two measures is the reason for this post's position. One passing run of an agent test is weak evidence. If you are about to release an agent, run each test case several times and report the share of tasks that passed every run.
The same agent on the same task gives different results
An eval is a repeatable test with a recorded result. In ordinary software the same input gives the same output, so a test that passes once is counted as done.
A language model varies. Anthropic's guide to agent evals, published on 9 January 2026, gives this as the reason for repeating: "Because model outputs vary between runs, we run multiple trials to produce more consistent results." An agent makes many model calls in one task, and a different choice at an early step can change the steps after it. Our guide page on why agent evals differ from LLM evals covers that multiplication.
One setting reduces the variation. Temperature is the setting that controls how much randomness a model uses when it chooses its words and at zero the model varies less. Some variation remains. A paper submitted in February 2026 by Stephan Rabanser and five co-authors ran each task 5 times at temperature zero, on 15 models, across two benchmarks, which are public sets of tasks with a fixed scoring method. The authors report that "outcome consistency remains low across all models", and explain it this way: "agents that can solve a task often fail to do so consistently." The tau-bench agent also ran at a temperature of 0.0.
METR, a research group that measures how long a task AI agents can complete, describes what the variation looks like across a set of tasks. Its time horizons page, last updated on 8 May 2026, says that a GPT-5 agent succeeds every time on a third of one group of tasks of similar length, fails every time on another third, and "sometimes succeeds and sometimes fails on the remaining third of tasks". METR gives those shares as approximate. A single run sorts that last third into passes and failures by chance.
What eight passes in a row require
The tau-bench authors named their stricter measure pass^k, read "pass hat k". They define it as "the chance that all k i.i.d. task trials are successful, averaged across tasks". A trial is one attempt at a task. The letters i.i.d. mean that the attempts are independent and run under the same conditions. With k set to 1, the measure is the ordinary single-run pass rate. With k set to 8, it is the chance of eight passes in eight runs.
Arithmetic shows why the second number is lower than the first. The following is an illustration that assumes that every run is independent and that the pass rate is the same on every run. We computed each figure with python3.
- At a 60 percent single-run rate, eight passes in a row have a chance of 0.6 to the power of 8, which is 0.01679616, or 1.68 percent.
- At a 90 percent single-run rate, the chance is 0.9 to the power of 8, which is 0.43046721, or 43.05 percent.
Under these two assumptions, an agent that passes 90 percent of single runs completes eight runs of a task with no failure less than half the time.
The tau-bench figure, below 25 percent, is higher than the 1.68 percent in the illustration, and the reason matters when you read your own results. The illustration gave every task the same pass rate, and real tasks differ. Take three invented tasks with single-run rates of 100, 60 and 20 percent. Their average is 60 percent. Their eight-run rates are 1, 0.01679616 and 0.00000256, and the average of those is 0.33893291, or 33.9 percent. So two agents with the same 60 percent single-run rate can have eight-run rates of 1.68 percent and 33.9 percent. The first agent fails some of the time on every task. The second fails on a few tasks and completes the others every time. Repeated runs show which agent you have.
The paper also defines the opposite measure, pass@k: "the chance that at least one out of k i.i.d. task trials is successful". It rises as k grows: Anthropic's guide says that by k=10, "pass@k approaches 100% while pass^k falls to 0%". A released agent that changes records gives each customer one run and acts on it, so pass^k describes what the customer gets.
What a one-run report leaves out
Here is a worked example. Every number in it is invented, and the arithmetic was computed with python3.
A team builds an agent that handles returns for an invented bicycle shop and writes 20 test cases. It runs each case 5 times, which is 100 runs. Twelve cases pass on all 5 runs. Five cases pass on 3 runs of 5, and we call these mixed cases. Three cases fail on all 5.
The passes add up to 12 x 5 + 5 x 3 + 3 x 0 = 75. The pass rate across all runs is 75 of 100, or 75 percent. The share of cases that passed every run is 12 of 20, or 60 percent.
Now suppose the same team had run each case once. The 12 cases that always pass are passes, and the 3 that always fail are failures. Each of the 5 mixed cases passes with a chance of 3 in 5. On average the single run shows 12 + 5 x 0.6 = 15 passes of 20, which is the same 75 percent. On a given day it can show any count from 12 to 17. If the runs are independent, all 5 mixed cases pass together with a chance of 0.6 to the power of 5, which is 0.07776, and on those days the report reads 17 of 20, or 85 percent.
Whatever the count, the one-run report places each mixed case in the pass column or the fail column and gives no sign that it is mixed. Those 5 cases are 25 percent of the set. They are also the cases where one customer's return is processed correctly and the next customer's is processed wrongly.
How many runs, and which numbers to report
No source we read gives a rule for the number of repeats. The published tests show a range of practice. The tau-bench authors ran at least 3 trials per task for their main results. The 2026 reliability paper ran each task 5 times. The 2024 paper AI Agents That Matter ran each agent five times on the 164 problems of the HumanEval coding benchmark. METR launches 6 independent runs for each task.
The arithmetic favours the upper end of that range. Under the same independence assumption, a task that the agent passes 80 percent of the time passes 3 runs of 3 with a chance of 0.8 to the power of 3, which is 0.512. Three repeats show that task as fully passing more often than they show its failure. At 5 repeats the chance is 0.32768, and at 8 it is 0.16777216.
Our recommendation is Reveneau's judgment, and no source measured it: run every case at least 5 times before a release decision, and 8 or more times for cases where one failure is costly, such as a payment or a deletion. Repeats multiply the cost of a test run directly, so 5 repeats cost 5 times one pass. For small daily changes, fewer repeats on a subset of cases reduce that cost.
Then report two numbers. The first is the pass rate across all runs. The second is the share of tasks that passed every run, with the number of runs written beside it. Add the list of mixed tasks for the engineers, because the next fixes are in that list. Leave best-of-several figures out of a release report. The authors of AI Agents That Matter point out that accuracy "can be improved by scientifically meaningless methods such as retrying".
Set the required rate for each class of task in writing before the runs, so that nobody changes the rule after seeing the result. The full method is on our guide page on agent reliability across repeated runs.
Read the failed runs before changing the agent
A mixed result has more than one possible cause, and some of the causes are in the test.
Many agent tests use a second model to play the customer. The authors of tau2-bench, a follow-up paper submitted in June 2025, checked that simulated customer and recorded an error rate of 40 percent in retail conversations and 47 percent in airline conversations. The task itself can be faulty: of 40 failed gpt-4o retail runs that the tau-bench authors examined, they traced 4 to faults in the instruction given to the simulated customer. Data left over from an earlier run is a third cause. Anthropic's guide says each trial should start "from a clean environment", and our page on test environments for agent evals explains how to reset between runs.
So read the recorded steps of each mixed task first. Fix the faults in the test, run the set again with the same number of repeats, and treat the mixed tasks that remain as the work list for the agent. The AI agent evals guide covers the other checks an agent needs, and two earlier posts cover related questions: can you trust an AI agent with real work yet? and why flaky tests are worse when agents write them.
Where Reveneau fits
At Reveneau, all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. When we build an agent, we run every case in that suite several times before a release decision, and we report the pass rate across all runs and the share of tasks that passed every run, with the number of runs stated. Repeats multiply the number of judged checks, so the speed of grading affects how many repeats are practical. We grade those checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. Reveneau, as a company, takes responsibility for the whole project through production and after release. To set reliability targets for an agent you plan to release, see AI development at Reveneau.
Sources
- tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains: Yao, Shinn, Razavi and Narasimhan, arXiv, submitted 17 June 2024. The definitions of pass^k and pass@k, the 61.2 percent single-run rate for gpt-4o on 115 retail tasks, the eight-run figure of below 25 percent, the run settings, and the 4 of 40 failures traced to the user instruction.
- Demystifying evals for AI agents: Anthropic, 9 January 2026. Why trials are repeated, how pass@k and pass^k move apart by 10 runs, and the rule that each trial starts from a clean environment.
- Towards a Science of AI Agent Reliability: Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan, arXiv, submitted 18 February 2026. Each task run 5 times at temperature zero on 15 models across two benchmarks, with low outcome consistency across all models.
- Task-Completion Time Horizons of Frontier AI Models: METR, last updated 8 May 2026, read 30 September 2026. Six independent runs per task, and the GPT-5 agent that always succeeds on a third of tasks, always fails on a third and is mixed on the rest.
- tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment: Barres, Dong, Ray, Si and Narasimhan, arXiv, submitted 9 June 2025. Simulated user error rates of 40 percent in retail and 47 percent in airline.
- AI Agents That Matter: Kapoor, Stroebl, Siegel, Nadgir and Narayanan, arXiv, submitted 1 July 2024. Each agent run five times on the 164 HumanEval problems, and the point that retrying raises accuracy.


