Engineering

An AI agent that passes a test once can fail it on the next run

Editorial · Reveneau · October 3, 2026

An AI agent that passes a test once can fail it on the next run

In June 2024 four researchers published a test called tau-bench. It gave AI agents 115 customer service tasks for a simulated shop. An AI agent is a language model working in a loop: it picks a step, calls a tool such as a database lookup, reads the result and repeats until the task is done. The best agent in the paper used the gpt-4o model and completed 61.2 percent of the retail tasks on a single run.

The authors then measured the same agent on the same tasks a second way: how often does it complete a task on all eight of eight runs? Averaged across the tasks, the answer was below 25 percent. The paper prints that figure only as "<25%" and gives no exact value.

Both figures belong to 2024: the models are from mid 2024, and the numbers describe no current model. The gap between the two measures is the reason for this post's position. One passing run of an agent test is weak evidence. If you are about to release an agent, run each test case several times and report the share of tasks that passed every run.

The same agent on the same task gives different results

An eval is a repeatable test with a recorded result. In ordinary software the same input gives the same output, so a test that passes once is counted as done.

A language model varies. Anthropic's guide to agent evals, published on 9 January 2026, gives this as the reason for repeating: "Because model outputs vary between runs, we run multiple trials to produce more consistent results." An agent makes many model calls in one task, and a different choice at an early step can change the steps after it. Our guide page on why agent evals differ from LLM evals covers that multiplication.

One setting reduces the variation. Temperature is the setting that controls how much randomness a model uses when it chooses its words and at zero the model varies less. Some variation remains. A paper submitted in February 2026 by Stephan Rabanser and five co-authors ran each task 5 times at temperature zero, on 15 models, across two benchmarks, which are public sets of tasks with a fixed scoring method. The authors report that "outcome consistency remains low across all models", and explain it this way: "agents that can solve a task often fail to do so consistently." The tau-bench agent also ran at a temperature of 0.0.

METR, a research group that measures how long a task AI agents can complete, describes what the variation looks like across a set of tasks. Its time horizons page, last updated on 8 May 2026, says that a GPT-5 agent succeeds every time on a third of one group of tasks of similar length, fails every time on another third, and "sometimes succeeds and sometimes fails on the remaining third of tasks". METR gives those shares as approximate. A single run sorts that last third into passes and failures by chance.

What eight passes in a row require

The tau-bench authors named their stricter measure pass^k, read "pass hat k". They define it as "the chance that all k i.i.d. task trials are successful, averaged across tasks". A trial is one attempt at a task. The letters i.i.d. mean that the attempts are independent and run under the same conditions. With k set to 1, the measure is the ordinary single-run pass rate. With k set to 8, it is the chance of eight passes in eight runs.

Arithmetic shows why the second number is lower than the first. The following is an illustration that assumes that every run is independent and that the pass rate is the same on every run. We computed each figure with python3.

  • At a 60 percent single-run rate, eight passes in a row have a chance of 0.6 to the power of 8, which is 0.01679616, or 1.68 percent.
  • At a 90 percent single-run rate, the chance is 0.9 to the power of 8, which is 0.43046721, or 43.05 percent.

Under these two assumptions, an agent that passes 90 percent of single runs completes eight runs of a task with no failure less than half the time.

The tau-bench figure, below 25 percent, is higher than the 1.68 percent in the illustration, and the reason matters when you read your own results. The illustration gave every task the same pass rate, and real tasks differ. Take three invented tasks with single-run rates of 100, 60 and 20 percent. Their average is 60 percent. Their eight-run rates are 1, 0.01679616 and 0.00000256, and the average of those is 0.33893291, or 33.9 percent. So two agents with the same 60 percent single-run rate can have eight-run rates of 1.68 percent and 33.9 percent. The first agent fails some of the time on every task. The second fails on a few tasks and completes the others every time. Repeated runs show which agent you have.

The paper also defines the opposite measure, pass@k: "the chance that at least one out of k i.i.d. task trials is successful". It rises as k grows: Anthropic's guide says that by k=10, "pass@k approaches 100% while pass^k falls to 0%". A released agent that changes records gives each customer one run and acts on it, so pass^k describes what the customer gets.

What a one-run report leaves out

Here is a worked example. Every number in it is invented, and the arithmetic was computed with python3.

A team builds an agent that handles returns for an invented bicycle shop and writes 20 test cases. It runs each case 5 times, which is 100 runs. Twelve cases pass on all 5 runs. Five cases pass on 3 runs of 5, and we call these mixed cases. Three cases fail on all 5.

The passes add up to 12 x 5 + 5 x 3 + 3 x 0 = 75. The pass rate across all runs is 75 of 100, or 75 percent. The share of cases that passed every run is 12 of 20, or 60 percent.

Now suppose the same team had run each case once. The 12 cases that always pass are passes, and the 3 that always fail are failures. Each of the 5 mixed cases passes with a chance of 3 in 5. On average the single run shows 12 + 5 x 0.6 = 15 passes of 20, which is the same 75 percent. On a given day it can show any count from 12 to 17. If the runs are independent, all 5 mixed cases pass together with a chance of 0.6 to the power of 5, which is 0.07776, and on those days the report reads 17 of 20, or 85 percent.

Whatever the count, the one-run report places each mixed case in the pass column or the fail column and gives no sign that it is mixed. Those 5 cases are 25 percent of the set. They are also the cases where one customer's return is processed correctly and the next customer's is processed wrongly.

How many runs, and which numbers to report

No source we read gives a rule for the number of repeats. The published tests show a range of practice. The tau-bench authors ran at least 3 trials per task for their main results. The 2026 reliability paper ran each task 5 times. The 2024 paper AI Agents That Matter ran each agent five times on the 164 problems of the HumanEval coding benchmark. METR launches 6 independent runs for each task.

The arithmetic favours the upper end of that range. Under the same independence assumption, a task that the agent passes 80 percent of the time passes 3 runs of 3 with a chance of 0.8 to the power of 3, which is 0.512. Three repeats show that task as fully passing more often than they show its failure. At 5 repeats the chance is 0.32768, and at 8 it is 0.16777216.

Our recommendation is Reveneau's judgment, and no source measured it: run every case at least 5 times before a release decision, and 8 or more times for cases where one failure is costly, such as a payment or a deletion. Repeats multiply the cost of a test run directly, so 5 repeats cost 5 times one pass. For small daily changes, fewer repeats on a subset of cases reduce that cost.

Then report two numbers. The first is the pass rate across all runs. The second is the share of tasks that passed every run, with the number of runs written beside it. Add the list of mixed tasks for the engineers, because the next fixes are in that list. Leave best-of-several figures out of a release report. The authors of AI Agents That Matter point out that accuracy "can be improved by scientifically meaningless methods such as retrying".

Set the required rate for each class of task in writing before the runs, so that nobody changes the rule after seeing the result. The full method is on our guide page on agent reliability across repeated runs.

Read the failed runs before changing the agent

A mixed result has more than one possible cause, and some of the causes are in the test.

Many agent tests use a second model to play the customer. The authors of tau2-bench, a follow-up paper submitted in June 2025, checked that simulated customer and recorded an error rate of 40 percent in retail conversations and 47 percent in airline conversations. The task itself can be faulty: of 40 failed gpt-4o retail runs that the tau-bench authors examined, they traced 4 to faults in the instruction given to the simulated customer. Data left over from an earlier run is a third cause. Anthropic's guide says each trial should start "from a clean environment", and our page on test environments for agent evals explains how to reset between runs.

So read the recorded steps of each mixed task first. Fix the faults in the test, run the set again with the same number of repeats, and treat the mixed tasks that remain as the work list for the agent. The AI agent evals guide covers the other checks an agent needs, and two earlier posts cover related questions: can you trust an AI agent with real work yet? and why flaky tests are worse when agents write them.

Where Reveneau fits

At Reveneau, all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. When we build an agent, we run every case in that suite several times before a release decision, and we report the pass rate across all runs and the share of tasks that passed every run, with the number of runs stated. Repeats multiply the number of judged checks, so the speed of grading affects how many repeats are practical. We grade those checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. Reveneau, as a company, takes responsibility for the whole project through production and after release. To set reliability targets for an agent you plan to release, see AI development at Reveneau.

Sources

  • tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains: Yao, Shinn, Razavi and Narasimhan, arXiv, submitted 17 June 2024. The definitions of pass^k and pass@k, the 61.2 percent single-run rate for gpt-4o on 115 retail tasks, the eight-run figure of below 25 percent, the run settings, and the 4 of 40 failures traced to the user instruction.
  • Demystifying evals for AI agents: Anthropic, 9 January 2026. Why trials are repeated, how pass@k and pass^k move apart by 10 runs, and the rule that each trial starts from a clean environment.
  • Towards a Science of AI Agent Reliability: Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan, arXiv, submitted 18 February 2026. Each task run 5 times at temperature zero on 15 models across two benchmarks, with low outcome consistency across all models.
  • Task-Completion Time Horizons of Frontier AI Models: METR, last updated 8 May 2026, read 30 September 2026. Six independent runs per task, and the GPT-5 agent that always succeeds on a third of tasks, always fails on a third and is mixed on the rest.
  • tau2-Bench: Evaluating Conversational Agents in a Dual-Control Environment: Barres, Dong, Ray, Si and Narasimhan, arXiv, submitted 9 June 2025. Simulated user error rates of 40 percent in retail and 47 percent in airline.
  • AI Agents That Matter: Kapoor, Stroebl, Siegel, Nadgir and Narayanan, arXiv, submitted 1 July 2024. Each agent run five times on the 164 HumanEval problems, and the point that retrying raises accuracy.

Common questions

Why is one passing run of an AI agent test weak evidence?

One passing run is weak evidence because the same agent on the same task passes on some runs and fails on others. In the tau-bench paper of June 2024, the best agent, built on gpt-4o, completed 61.2 percent of 115 retail tasks on a single run, and its rate for completing a task on all eight of eight runs was below 25 percent. Those models are from mid 2024, and the figures describe no current model.

What did the tau-bench paper measure about repeated runs of an agent?

The tau-bench paper measured how often an agent completes the same task on every one of several runs, a measure its authors call pass^k. For the gpt-4o agent on the retail tasks, the single-run rate was 61.2 percent and the eight-run figure is printed only as below 25 percent, with no exact value. The main results used at least 3 runs per task, with the agent model's randomness setting at 0.0.

How many times should a team run each agent test case before a release?

Run each agent test case at least 5 times before a release decision, and 8 or more times where one failure is costly, such as a payment. That is Reveneau's judgment, because no source we read gives a rule for the number of repeats. Published tests show the range of practice: tau-bench ran at least 3 runs per task, a 2026 reliability paper ran 5, and METR launches 6.

How do you work out the chance that an agent passes eight runs in a row?

Multiply the single-run pass rate by itself eight times. For a rate of 60 percent that is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent. This is an illustration that assumes independent runs with the same rate on each. A real set of tasks gives a higher figure when its failures sit in a few tasks, because a task that passes every time counts fully at any number of runs.

Which numbers belong in an agent test report?

An agent test report needs two numbers: the pass rate across all runs, and the share of tasks that passed every run, with the number of runs written beside the second. Add the list of tasks that passed on some runs and failed on others. In an invented set of 20 cases run 5 times each, 75 passes of 100 runs and 12 cases with no failure give 75 percent and 60 percent.

Does setting the temperature to zero make an agent give the same result on every run?

No, an agent still varies between runs when temperature, the setting that controls how much randomness a model uses, is at zero. A February 2026 paper by Stephan Rabanser and five co-authors ran each task 5 times at temperature zero on 15 models across two benchmarks. It reports that outcome consistency remains low across all models, and that agents that can solve a task often fail to do so consistently.

Why should a best-of-several result stay out of an agent release report?

A best-of-several result rises with every extra attempt while the agent stays the same, so it says little about a product where each user gets one run. Anthropic's guide to agent evals notes that by 10 runs the chance of at least one pass approaches 100 percent while the chance of passing every run falls to 0. The authors of AI Agents That Matter make the same point about accuracy raised by retrying.

What should a team do with a task that passes on some runs and fails on others?

Read the recorded steps of the failed runs of that task before changing the agent, because the cause can be in the test. The tau-bench authors examined 40 failed gpt-4o retail runs and traced 4 of them to faults in the instruction given to the simulated customer. Fix faults in the test first, then treat the tasks that still pass and fail as the work list for the agent.

Does repeating every case make agent testing too expensive?

Repeating every case multiplies the cost of the test run by the number of repeats, so 5 repeats cost 5 times one pass. If that cost is too high for daily work, run the full repeats before a release decision and fewer repeats on a subset of cases for small changes. For scale, the tau-bench authors reported 0.38 US dollars per retail task for the agent and 0.23 for the simulated customer in 2024.