What to measure

Agent reliability: why one passing run proves little

An agent that passes a task once can fail the same task on the next run. The tau-bench paper of June 2024 measured this with pass^k, the chance that all k runs of a task pass: its best agent passed 61.2 percent of single retail runs and had a pass^8 below 25 percent. One passing run is therefore weak evidence. Run every case several times, at least 5 before a release decision, and report two numbers: the pass rate across all runs and the share of tasks that passed every run. A user of the released product gets one run, so the second number describes what users experience.

Published September 30, 2026. Editorial.

Key takeaways

  • Pass^k is the chance that an agent passes all k runs of the same task, averaged across tasks. The tau-bench paper introduced the measure in June 2024.
  • In that paper the best agent, built on gpt-4o, passed 61.2 percent of single retail runs, and its pass^8 was below 25 percent. The models are from mid 2024.
  • As an illustration: if a task passes 60 percent of independent single runs, the chance of eight passes in a row is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent.
  • Report two numbers: the pass rate across all runs, and the share of tasks that passed every run, with the number of runs stated beside it.
  • Three runs are too few to trust a pass: a task with an 80 percent pass rate passes three independent runs out of three with a probability of 0.512.

The tau-bench paper, submitted in June 2024, tested agents that act as customer service assistants for a simulated shop and a simulated airline. Its best agent used gpt-4o and passed 61.2 percent of single runs on the 115 retail tasks [1]. The authors also measured how often the agent passes all eight of eight runs of the same task. For that agent on retail the figure was below 25 percent [1]. The paper gives no exact value: its abstract and its results section print "<25%", and its introduction gives 25 percent and marks it as approximate.

An agent in production meets the same request many times, from many customers, and each customer gets one run. A test set that runs every case once reports the first number and gives no information about the second. This page explains the measure the tau-bench authors introduced, the arithmetic behind it, how many repeats to run, and which two numbers to report.

What pass^k measures

The tau-bench authors define pass^k (read "pass hat k") as "the chance that all k i.i.d. task trials are successful, averaged across tasks" [1]. A trial is one attempt at a task. The letters i.i.d. stand for independent and identically distributed: each attempt starts fresh, under the same conditions, with no knowledge of the others. So pass^1 is the ordinary single-run pass rate, and pass^8 is the chance of eight passes in eight runs.

The paper contrasts it with a second measure, pass@k: "the chance that at least one out of k i.i.d. task trials is successful" [1]. The two move in opposite directions as k grows. Anthropic's guide to agent evals says that at k=1 the two are "identical (both equal the per-trial success rate)", and that by k=10 "pass@k approaches 100% while pass^k falls to 0%" [2].

Which one fits depends on how the agent is used. Where a person looks at several attempts and keeps the best, such as several drafts of a piece of code, pass@k describes what that person gets. Where each user receives a single run and the agent acts on it, pass^k describes what the user gets. A released agent that changes records belongs to the second case.

The arithmetic of eight passes in a row

Here is an illustration with invented numbers. Suppose an agent passes a task on 60 percent of single runs, every run is independent, and the rate is the same on every run. The chance of eight passes in a row is 0.6 multiplied by itself eight times:

0.6^8 = 0.01679616, which is 1.68 percent.

Anthropic's guide gives a smaller example of the same sum: at a 75 percent single-run rate, three passes in a row have a probability of 0.75^3, which is 0.421875, or 42 percent as the guide rounds it [2]. That figure is arithmetic, and nobody measured it.

The table shows the same calculation for three single-run rates. Every cell is the rate raised to the power of the number of runs, under the same two assumptions.

Single-run pass rate 3 runs in a row 5 runs in a row 8 runs in a row
90 percent 72.9 percent 59.049 percent 43.046721 percent
80 percent 51.2 percent 32.768 percent 16.777216 percent
60 percent 21.6 percent 7.776 percent 1.679616 percent

A 90 percent agent sounds dependable. Under these assumptions it completes eight runs of one task without a failure less than half the time.

What the tau-bench paper reported

The paper's main results are single-run rates. Each task was run at least 3 times, each run was limited to 30 agent actions, and the agent model's temperature was 0.0 [1]. Temperature is the setting that controls how much randomness a model uses when it chooses its words.

Agent model Retail pass^1 Airline pass^1
gpt-4o 61.2 35.2
gpt-4-turbo 57.7 32.4
claude-3-opus 44.2 34.7
gpt-3.5-turbo 20.0 10.8

These are the paper's figures for models available in mid 2024, on 115 retail tasks and 50 airline tasks [1]. They record that test and describe no current model. For the eight-run measure the paper gives one statement: for the gpt-4o agent, "pass^8 drops to <25%" [1].

The paper's figure, which its introduction gives as 25 percent and marks as approximate, is higher than the 1.68 percent in the illustration, and the reason affects how you read your own results. The illustration assumed every task has the same pass rate. Real tasks differ. Take three invented tasks with single-run rates of 100, 60 and 20 percent. Their average is 60 percent. Their eight-run rates are 1, 0.01679616 and 0.00000256, and the average of those three is 0.33893291, or 33.9 percent. Two agents with the same 60 percent average can therefore have eight-run results of 1.68 percent and 33.9 percent, depending on whether the failures are spread across all tasks or concentrated in a few.

METR, a research group that measures how long a task AI agents can complete, describes this pattern in a later agent. Its time horizons page, last updated on 8 May 2026, says that on tasks that take a human expert 90 minutes to 3 hours, a GPT-5 agent succeeds every time on a third of the tasks, fails every time on another third, and "sometimes succeeds and sometimes fails on the remaining third of tasks" [3]. METR's own wording for each of the first two shares is "around one-third", so the split is approximate.

The measure also changes with the setting. In the tau2-bench paper of June 2025, gpt-4.1 had a pass^1 of 74 percent on retail, 56 percent on airline and 34 percent on a new telecom setting. For a second model, claude-3.7-sonnet, the single-run rate on telecom was 49 percent, which the authors say matches its airline rate, and they report that as k increases "the pass^k scores decline more rapidly for telecom compared to airline" [4].

Why the same agent passes and fails on the same input

A flaky agent test is one that passes on some runs and fails on others with the same input. Four causes produce that result, and they call for different responses.

The model's output varies. Anthropic's guide gives the reason for repeating runs in one line: "Because model outputs vary between runs, we run multiple trials to produce more consistent results" [2]. Setting the temperature to zero reduces the variation, and some of it remains. A February 2026 paper by Stephan Rabanser and five co-authors ran each task 5 times at temperature zero on 15 models across two benchmarks (public sets of tasks with a fixed scoring method), GAIA and tau-bench. It found that "outcome consistency remains low across all models", which the authors explain as "agents that can solve a task often fail to do so consistently" [5]. The same paper reports that agents choose similar types of action from run to run and change the order of them [5].

The simulated user varies. Many agent tests use a second model to play the customer. The tau2-bench authors checked that simulated user and recorded an error rate of 40 percent in retail and 47 percent in airline, with 12 and 13 percent being critical errors that prevent the task from being completed. In the telecom setting the rate was 16 percent, with 6 percent critical [4]. Some failed runs are caused by the test itself.

State is left over from an earlier run. Anthropic's guide says each trial should start from a clean environment, because leftover files or stored data can cause failures that come from the test setup and say nothing about the agent [2]. Resetting between runs is covered in test environments for agents.

The task is faulty. Of the 40 failed gpt-4o retail runs the tau-bench authors examined, they traced 4 to faults in the instruction given to the simulated user [1]. The maintainers of tau2-bench list "Task Quality (75+ fixes)" in their repository notes, and state that results from versions before 1.0.1 "are not comparable with >= 1.0.1" [6].

The first cause is the agent's own behaviour, and it is the thing you are measuring. The other three are faults in the test, and they should be fixed before the number is trusted. Reading the recorded steps of each mixed task shows which cause applies.

How many times to run each case

No source used for this page gives a rule for the number of repeats. The published tests show a range of practice:

  • tau-bench ran at least 3 trials per task for its main results [1].
  • The 2026 reliability paper ran each task 5 times [5].
  • The 2024 paper "AI Agents That Matter" ran each agent five times on the 164 problems of the HumanEval coding benchmark [7].
  • METR launches 6 independent runs for each task [3].

The arithmetic gives a reason to prefer the upper end. Under the independence assumption, a task the agent passes 80 percent of the time passes 3 runs out of 3 with a probability of 0.8^3 = 0.512. Three repeats show that task as fully passing more often than they show its failure. At 5 repeats the probability is 0.32768, and at 8 it is 0.16777216.

Our position is a judgment, since no source measured it: run every case at least 5 times before a release decision, and 8 or more times for cases where one failure is costly, such as a payment or a deletion. On small day-to-day changes, fewer repeats on a subset of cases reduce the cost.

Repeats multiply cost directly: 5 repeats cost 5 times one pass. The tau-bench authors reported that one retail task cost 0.38 US dollars for the agent and 0.23 for the simulated user in 2024 [1], and cost, time and step limits works through the sum. How many separate cases the set needs is a different question, answered in how many test cases an LLM eval needs.

What to report

Report two numbers and state k beside the second.

  1. The pass rate across all runs. Count every run of every task and divide the passes by the total. This is pass^1. It tells you how often a single attempt works.
  2. The share of tasks that passed every run. This is pass^k for the k you ran. It tells you how many tasks a user can rely on.

Add a third item for the engineers: the list of mixed tasks, the ones that passed on some runs and failed on others. The next fixes are in that list.

Leave best-of-k figures out of the report. The authors of "AI Agents That Matter" point out that accuracy "can be improved by scientifically meaningless methods such as retrying" [7]. A best-of-k figure is pass@k, and it rises with every extra attempt even when the agent is unchanged.

The released product needs the second number because of what a failure means there. In a test, a failed run is a line in a report. In production it is a customer whose refund went to the wrong order. The reliability paper's opening claim is that "compressing agent behavior into a single success metric obscures critical operational flaws" [5].

What reliability is enough for production

The threshold differs from agent to agent. It depends on two facts about the task: what one failure costs, and whether a person checks the result before it takes effect. An agent that drafts a reply for a staff member to approve can be released at a lower all-runs rate than an agent that issues refunds unattended.

Higher reliability also shortens the task an agent can be given. METR's paper on task length reports that the length of task models complete at 80 percent reliability is 4 to 6 times shorter than the length they complete at 50 percent [8], and METR's page says that measuring a 99 percent figure "would require many more tasks" [3].

Set the threshold for each class of task in writing before the runs, so that nobody changes the rule after seeing the result. Using evals to decide whether an AI feature is ready to release covers that decision, and how to make an AI agent reliable enough to release is the business overview. Failures inside a single step are the subject of tool-call evals, and the full method is in the AI agent evals guide. A related post asks can you trust an AI agent with real work yet?.

How Reveneau applies this

At Reveneau, all code is written by AI, and every change must pass a large eval suite written from the specification before the code. For an agent, we write each case's pass condition and its required all-runs rate from the specification before the prompt exists, and we run every case several times before a release decision. We report the single-run pass rate and the share of tasks that passed every run, with the number of runs stated. Repeats multiply the number of judged checks, so the speed of grading affects how many repeats are practical. We grade those checks with Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader.

We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To discuss reliability targets for an agent you are planning, see AI development at Reveneau or contact us.

Best for

  • Agents that act without a person checking each result
  • Release decisions, where each user will get a single run
  • Comparing two versions of an agent whose single-run rates are close

Avoid if

  • The test environment keeps data from one run to the next, which makes the repeats depend on each other
  • The report would show the best of several runs as the result

Check before you decide

  • The number of runs is stated next to every all-runs figure
  • Each run starts from a reset environment
  • Mixed tasks are listed and their recorded steps were read

Common questions

What is pass^k in agent evaluation?

Pass^k is the chance that an agent passes all k runs of the same task, averaged across tasks. The tau-bench paper introduced it in June 2024 and defines it as the chance that all k independent trials of a task are successful. At k equal to 1 it is the ordinary pass rate. At higher k it shows how often an agent completes a task every time, which is what a user who gets one run depends on.

What is the difference between pass^k and pass@k?

Pass^k is the chance that every one of k runs passes, and pass@k is the chance that at least one of k runs passes. Anthropic's guide to agent evals notes that the two are identical at k equal to 1 and move apart as k grows: by 10 runs pass@k approaches 100 percent while pass^k falls to 0. Use pass@k when a person keeps the best attempt, and pass^k when each user gets one.

How many times should I run each agent test?

Run each agent test at least 5 times before a release decision, and 8 or more times where one failure is costly. That is Reveneau's position, and no source gives a measured rule. Published tests use at least 3 runs per task (tau-bench), 5 (the 2026 reliability paper) and 6 (METR). With 3 runs, a task that passes 80 percent of the time shows three passes with a probability of 0.512.

Why does my agent pass a test sometimes and fail it other times?

An agent passes and fails the same test because its model's output varies between runs, even when the model's randomness setting, called temperature, is at zero. A 2026 paper that ran each task 5 times at temperature zero on 15 models found that agents that can solve a task often fail to do so consistently. Three other causes are in the test: a simulated user that makes errors, data left over from an earlier run, and a faulty task.

What is a flaky agent test?

A flaky agent test is a test that passes on some runs and fails on others with the same input. The cause can be the agent, whose output varies from run to run, or the test, through a simulated user's errors, leftover data or a faulty task. The tau2-bench authors measured simulated user error rates of 40 percent in retail and 47 percent in airline conversations, so read the recorded steps before deciding the agent is at fault.

What reliability is enough to release an agent to production?

The reliability needed to release an agent depends on what one failure costs and whether a person checks the result before it takes effect. Set a required all-runs pass rate for each class of task in writing before the runs. METR's paper reports that the task length models complete at 80 percent reliability is 4 to 6 times shorter than at 50 percent, so a higher target means giving the agent shorter tasks.

How do I calculate the chance of eight passes in a row?

Multiply the single-run pass rate by itself eight times, assuming the runs are independent and the rate is the same on each. For a 60 percent rate the result is 0.6 to the power of 8, which is 0.01679616, or 1.68 percent. That calculation is an illustration. Results across a set of tasks are higher when tasks differ in difficulty, because a task that passes every time counts fully at any number of runs.

What did the tau-bench paper find about repeated runs?

The tau-bench paper found that its best agent in June 2024, built on gpt-4o, passed 61.2 percent of single runs on 115 retail tasks, and that its pass^8, the chance of eight passes in eight runs averaged across tasks, was below 25 percent. The paper states the eight-run figure only as below 25 percent. The models are from mid 2024, so the figures describe that test and no agent available today.

Which numbers should an agent eval report show?

An agent eval report should show two numbers: the pass rate across all runs, and the share of tasks that passed every run, with the number of runs stated. A list of the tasks that passed on some runs and failed on others belongs beside them. Leave out best-of-several figures: the authors of AI Agents That Matter note that accuracy can be raised by retrying, which leaves the agent unchanged.

Does repeating every test make agent evals too expensive?

Repeating tests multiplies the cost of the run by the number of repeats, so 5 repeats cost 5 times one pass. Control it by running the full repeats before a release decision and fewer repeats on a subset of cases for small daily changes. The tau-bench authors reported 0.38 US dollars per retail task for the agent and 0.23 for the simulated user in 2024, which shows the scale for one benchmark.

Do I need repeated runs if my agent gave the same result three times?

Yes, keep the repeats. Three identical passes are weak evidence: a task with an 80 percent pass rate produces three passes in a row with a probability of 0.512, assuming independent runs. The 2026 reliability paper found low consistency of outcomes across all 15 models it tested with the randomness setting at zero, so an agent that agreed with itself on a few runs can still vary on later ones.

What should I do with a task that passes on some runs and fails on others?

Read the recorded steps of the failing runs of that task first, then decide whether the cause is the agent or the test. The tau-bench authors traced 4 of 40 examined failures to faults in the instructions given to the simulated user. Fix the test faults, treat the remaining mixed tasks as the work list for the agent, and re-run the full set with the same number of repeats.

References