What to measure

Cost, time and step limits as agent evals

A run that returns the right result can still be a failed run if it took too many steps or cost too much. An agent eval should record model usage cost, elapsed time, steps and tool calls for every run, and fail any run that exceeds a written limit on one of them. The 2024 paper AI Agents That Matter found coding agents of similar accuracy whose cost differed by almost two orders of magnitude (two orders is a factor of one hundred), so accuracy reported without cost cannot be compared. No source gives a recommended limit as a number: set each one from your own passing runs. METR's time horizon figures, quoted here with their dates, show how long a task models complete.

Published September 30, 2026. Editorial.

Key takeaways

  • Record five things for every agent run: model usage cost, elapsed time, number of steps, number of tool calls, and whether the task passed. Put a limit on each of the first four.
  • The 2024 paper AI Agents That Matter found that, on a coding benchmark of 164 problems, cost differed by almost two orders of magnitude (two orders is a factor of one hundred) between agents of similar accuracy.
  • A step limit ends a loop, which is an agent repeating the same step. Anthropic's guide to building agents names a maximum number of iterations as a common stopping condition.
  • No source recommends a step or cost limit for a product as a number. Set each limit from the largest value among your own passing runs, plus a margin you write down.
  • METR's paper reports that models' 80 percent time horizons are 4 to 6 times shorter than their 50 percent horizons, so a higher reliability target means giving an agent shorter tasks.

HAL is a leaderboard (a public ranking table) of AI agents run by a team at Princeton University. As read on 30 September 2026, it listed two 2025 runs on a benchmark (a public set of tasks with a fixed scoring method) named AssistantBench. Both rows used the same agent software, Browser-Use. With the o3 model at its medium setting (April 2025) the agent scored 38.8 percent, and the cost shown beside the score was 15.15 US dollars. With GPT-5 at its medium setting (August 2025) it scored 35.2 percent at 41.69 US dollars [1]. The second row cost 2.75 times as much as the first (41.69 divided by 15.15) and scored lower.

Two rows show one comparison and no trend. The site says its leaderboard is paused, so the figures are a record of those 2025 runs and say nothing about current models. A report that listed accuracy alone would hide which run was the costly one. This page is about putting cost, time and step counts into the eval itself, each with a limit that fails a run when it is exceeded.

Why accuracy alone is the wrong score for an agent

The main evidence is a July 2024 paper titled "AI Agents That Matter", by Sayash Kapoor, Arvind Narayanan and three co-authors at Princeton University. They ran published coding agents five times each on HumanEval, a benchmark of 164 programming problems, and recorded the cost of each agent beside its accuracy [2]. Three of their results:

  • For accuracy the authors describe as similar, cost differed by "almost two orders of magnitude" [2]. Two orders of magnitude is a factor of one hundred.
  • Two complex agents, Reflexion and LDB, cost over 50 percent more than a simple comparison method the authors call the warming strategy, and a third, LATS, cost over 50 times more [2].
  • The authors found "no significant accuracy difference between our warming strategy and the best-performing agent architecture" [2].

Their conclusion is that "useful agent evaluations must control for cost", because accuracy "can be improved by scientifically meaningless methods such as retrying" [2]. The models were from the GPT-3.5 and GPT-4 period and the result comes from one benchmark, so the exact ratios belong to that test. The reasoning applies to any agent: a score can be raised by spending more, so a score reported without its cost cannot be compared with another.

A second study, the HAL paper of October 2025, covered 21,730 agent runs across 9 models and 9 benchmarks, at a total cost the authors give as 40,000 US dollars and mark as approximate. In 21 of 36 combinations of model, agent and benchmark, a higher reasoning effort setting, which lets the model do more computation before it answers, produced "equal or lower accuracy" [3].

What to measure on every run

Record five things for each run of each task.

  1. Model usage cost. Models are billed by the token, a word piece the model reads or writes. Cost is the tokens used multiplied by the provider's price, added up over every model call in the run.
  2. Wall-clock time. The elapsed time from the first request to the final result.
  3. Number of steps. One step is one turn of the agent's loop: the model chooses an action, the action runs, the model reads the result.
  4. Number of tool calls. A step can contain more than one call, and a tool can have its own charge.
  5. Whether the task passed.

Anthropic's guide to agent evals lists similar quantities as things a fixed set of tasks lets a team track: "latency, token usage, cost per task, and error rates" [4]. Latency is the waiting time for a response. The Berkeley Function Calling Leaderboard publishes an estimated cost in US dollars and latency in seconds beside each model's accuracy [5].

Report the spread as well as the average. A February 2026 paper on agent reliability ran each task 5 times and found "high variance in token and compute usage across runs" [6]. Run each task several times, as agent reliability across repeated runs describes, and record the middle value and the largest value across the runs.

A limit turns a measurement into a pass or a fail

A measurement becomes an eval when it has a threshold. Set one limit for each measure, and count a run that exceeds it as failed even when the result is correct.

Limit What it protects against How to set it
Step limit A loop, and runs that take many needless steps From the step counts of passing runs, plus a margin
Tool-call limit Repeated calls to a service that charges per call or restricts how many it accepts For each tool, from passing runs
Cost limit per task A run that costs more than the task is worth From the value of the task to the business
Time limit A user or another system left waiting From how long the user will wait

Published tests set limits of this type. tau-bench stopped each task at 30 agent actions [7]. The OSWorld benchmark, which tests agents on real desktop applications, set "the maximum steps of interaction to 15 and the maximum time limits to 30 minutes for all tasks" [8]. The authors of SWE-Agent, as cited in "AI Agents That Matter", set a cost limit of 4 US dollars per task [2]. METR runs its agents "up to a token and time limit" [9].

Researchers chose those values for a benchmark. No source used for this page recommends a step or cost limit for a product as a number, so any figure for your product comes from your own runs. The method:

  1. Run the task set with repeats, with only a high safety limit in place.
  2. Look at the passing runs alone. Note the largest step count, cost and time among them.
  3. Set each limit above that largest value by a margin you choose and write down.
  4. Run again, and read every run that now fails on a limit alone.

Here is an invented illustration. Suppose the passing runs of a refund task took between 6 and 11 steps across 40 runs. A step limit of 15 allows an unusual valid route and still ends a run that has stopped making progress.

How a step limit stops a loop

A loop is an agent repeating the same step: it calls a tool, gets a result it cannot use, and makes the identical call again. Each repeat costs tokens and time, and with no stopping rule the run continues until something outside the agent ends it.

Anthropic's guide to building agents names the control: "it's also common to include stopping conditions (such as a maximum number of iterations) to maintain control" [10]. An iteration is one turn of the loop. The same guide says the autonomous nature of agents "means higher costs, and the potential for compounding errors" [10], and a loop produces both at once.

In an eval, use the limit in three ways:

  • End the run. When the step count reaches the limit, the run stops and is recorded as failed, with the reason "step limit reached".
  • Detect the repeat earlier. A code check on the recorded steps can flag the same tool called with the same arguments several times in a row, before the limit is reached. A tool error is one situation where an agent can start repeating itself, which is why error handling is one of the five checks in tool-call evals.
  • Check what happens at the limit. An agent that reaches its limit should stop, say what it completed, and hand the task to a person. Write a case for that behaviour.

Limits work beside other controls, which are covered in permissions and confirmation steps for AI agents.

How much an agent run costs

Cost depends on the task, the model and the date, so the only figure you can rely on is one you measure. Published figures show the range.

  • tau-bench, 2024. With a gpt-4o agent and a gpt-4 simulated user on the retail tasks, the authors reported 0.38 US dollars per task for the agent and 0.23 for the simulated user, and put the cost of one trial of every task at 200 dollars, a figure the paper marks as approximate [7]. Both figures are repeated as the paper states them: 0.61 dollars multiplied by the 115 retail tasks is 70.15 dollars, and the passage does not explain the difference.
  • SWE-Agent, 2024. With a limit of 4 US dollars per task, running the agent on the entire benchmark "could cost over USD 8,000 for a single evaluation run", as cited in "AI Agents That Matter" [2].
  • HAL, 2025. 21,730 runs for a total the authors give as 40,000 US dollars and mark as approximate [3]. Dividing that total by the count gives 1.84 dollars per run. That is an estimate, because the total is approximate, and it is an average across unlike benchmarks.

One pass of an eval set costs the number of cases, times the repeats for each case, times the cost of one run. As an illustration with an invented case count: 50 cases, 5 repeats each, at the 0.61 dollars per task reported for tau-bench, comes to 50 x 5 x 0.61 = 152.50 dollars for one pass. The cost of the product itself is covered in what it costs and how long it takes to release an AI agent, and the budget for evals in what AI evals cost and who should own them.

How long a task an agent can finish: METR's time horizon

A time limit depends on how long a task an agent can complete at all. METR, a research organisation, measures this with a figure it calls the time horizon. Its paper defines the 50 percent version as "the time humans typically take to complete tasks that AI models can complete with 50% success rate" [11]. METR's page adds that the figure is "a measure of the difficulty of a task, rather than the time an AI spends to complete the task" [9]. A two-hour horizon means the agent succeeds half the time on tasks that take a skilled person two hours.

METR's figures differ between its publications, so each is given with its source and date [11] [12] [13] [9].

Source Date What it states
Paper, abstract First version 18 March 2025, read in its fourth version of 10 July 2026 Claude 3.7 Sonnet has a 50 percent horizon of 50 minutes, and the horizon has doubled every seven months from 2019 on. The abstract marks both figures as approximate.
Paper, body Same paper Doubling time of 207 days for the 50 percent horizon over 2019 to 2025, and 204 days for the 80 percent horizon
Blog post 19 March 2025 The same model's horizon is one hour, which the post marks as approximate
Time Horizon 1.1 29 January 2026 Task set grew from 170 to 228 tasks. Doubling time 196.5 days over the whole period, and 130.8 days from 2023 on (interval 107 to 161)
Data file Read 30 September 2026 Doubling time 187.778 days over the whole period, and 128.744 days from 2023 on (interval 104.428 to 158.012)

Three facts from these sources matter more than any single value.

  • Higher reliability means a shorter horizon. The paper reports that models' 80 percent horizons are 4 to 6 times shorter than their 50 percent horizons [11].
  • Success falls as tasks get longer. In March 2025 METR reported that the models of that date succeeded on almost all tasks that take a person less than 4 minutes, and on fewer than 10 percent of tasks longer than a threshold that METR gives as 4 hours and marks as approximate [12].
  • The estimates are wide and they change. METR's data file, as read on 30 September 2026, gives the model listed as gpt_5_4, released on 5 March 2026, a 50 percent estimate of 341.74 with an interval of 186.58 to 768.78, and an 80 percent estimate of 53.88. The file prints no unit, and the scale matches minutes. METR's January 2026 post says its confidence intervals are still wide [13], and its page says measurements above 16 hours "are unreliable with our current task suite" [9].

METR's March 2025 post describes its tasks as multi-step software and reasoning tasks [12], so a horizon for another type of work has to be measured on that work. For your limits, the practical reading is this: if a task takes a skilled person several hours, plan for a lower pass rate than on short tasks, test it with repeats, and set the time and step limits from passing runs. Agent benchmarks explained covers the benchmark itself.

Put the same limits into the released agent

The released agent should stop at the same step, cost and time limits that the eval used, so that the eval describes how the product behaves. A production run that reaches a limit is also a candidate for a new eval case, as turning production traces into eval cases describes. The rest of the method is in the AI agent evals guide. A related post covers what founders get wrong about AI agents.

How Reveneau applies this

At Reveneau, all code is written by AI, and every change must pass a large eval suite written from the specification before the code. For an agent, that suite records cost, time, steps and tool calls for every run, and each task has a written limit for each of them. A run that returns the right result over its limit fails. We report cost and time beside the pass rate, so that a change which raises accuracy by spending more is visible as such. The run time of the eval suite matters as well: we grade the judged checks with Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader.

We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To set limits for an agent you are planning, see AI development at Reveneau or contact us.

Best for

  • Agents billed by usage, where a correct but costly run loses money
  • Agents that can repeat a failed step
  • Comparing two agent versions whose accuracy is close

Avoid if

  • The plan is to copy a benchmark's limit, such as 30 actions, without measuring your own passing runs
  • The limits would exist in the test set and be absent from the released agent

Check before you decide

  • Every run records cost, time, steps and tool calls
  • Each limit was set from passing runs and the margin is written down
  • A case checks that the agent stops and reports at the limit

Common questions

How much does one AI agent run cost?

The cost of one agent run depends on the task, the model and the date, so measure it on your own tasks. Published examples: the tau-bench authors reported 0.38 US dollars per retail task for the agent and 0.23 for the simulated user in 2024, and the paper AI Agents That Matter cites a limit of 4 US dollars per task set by the authors of SWE-Agent. Both are benchmark figures for 2024 models.

How do I stop an AI agent from looping?

Stop an agent from looping with a step limit: a maximum number of turns after which the run ends. Anthropic's guide to building agents names a maximum number of iterations as a common stopping condition. Add a code check that flags the same tool called with the same arguments several times in a row, and test what the agent does at the limit: stop, report what was completed, and hand the task to a person.

What is a step limit for an AI agent?

A step limit is the maximum number of turns an agent may take on one task before the run is ended. One step is one turn of the agent's loop: choose an action, run it, read the result. Benchmarks set them: tau-bench stopped each task at 30 agent actions, and OSWorld at 15 steps and 30 minutes. Researchers chose those values for a test, and no source recommends a number for a product.

Should cost be part of an agent eval?

Yes. The 2024 paper AI Agents That Matter concludes that useful agent evaluations must control for cost, because accuracy can be raised by spending more, for example by retrying. On the HumanEval coding benchmark of 164 problems its authors found that cost differed by almost two orders of magnitude (two orders is a factor of one hundred) between agents of similar accuracy. Record cost for every run and report it beside the pass rate.

What is METR's time horizon?

METR's time horizon is the length of task, measured by how long it takes a human expert, that an AI agent completes at a given rate of success. The 50 percent horizon is the task length at which the agent succeeds half the time. METR's page states that the figure measures the difficulty of a task, which is a separate thing from how long the agent spends working on it.

How long a task can an AI agent finish?

METR's published answers differ by source and date. The abstract of its paper, first published in March 2025, gives Claude 3.7 Sonnet a 50 percent horizon of 50 minutes and marks the figure as approximate. METR's data file, read on 30 September 2026, lists a model released in March 2026 at 341.74, with an interval from 186.58 to 768.78; the file prints no unit and the scale matches minutes. Horizons at 80 percent reliability are 4 to 6 times shorter.

How do I choose the number for a step or cost limit?

Choose a limit from your own passing runs. Run the task set with repeats under a high safety limit, find the largest step count, cost and time among the runs that passed, and set each limit above that value by a margin you write down. No source used in this guide recommends a number for a product. Benchmark limits such as tau-bench's 30 actions were chosen by researchers for one test.

Does a higher reasoning setting give an agent better results?

A higher reasoning effort setting gave equal or lower accuracy in 21 of 36 combinations in one large study. The paper on HAL, a public ranking of AI agents run by a Princeton University team, covered 21,730 agent runs across 9 models and 9 benchmarks in October 2025. In those 21 combinations of model, agent and benchmark, the higher setting produced equal or lower accuracy. Measure accuracy and cost together on your own tasks before choosing the higher setting.

What is the difference between a step limit and a time limit?

A step limit counts the agent's turns and a time limit counts minutes on the clock. A step limit stops a loop, where the agent repeats the same action. A time limit protects a user or another system that is waiting, and it catches a slow tool that a step count would miss. The OSWorld benchmark set both: 15 steps and 30 minutes for every task.

What goes wrong if an agent eval has no limits?

Without limits, an eval counts a run as passed whatever it spent, and a loop can continue until something outside the agent ends it. The eval then rewards changes that raise accuracy by spending more. The authors of AI Agents That Matter found two complex coding agents costing over 50 percent more than the authors' simple comparison method, and a third over 50 times more. They report no significant accuracy difference between that method and the best-performing agent design.

How much does it cost to run an agent eval suite?

The cost of one pass of an agent eval suite is the number of cases, multiplied by the repeats per case, multiplied by the cost per run. As an illustration, 50 cases run 5 times each at the 0.61 US dollars per task reported for tau-bench in 2024 comes to 152.50 dollars. The paper AI Agents That Matter notes that one full run of SWE-Agent on its benchmark could cost over 8,000 US dollars.

Do these limits apply to a simple agent with one or two tools?

Yes, the limits apply to any agent that works in a loop, including one with two tools. A loop needs only one tool and one result the agent cannot use. For a small agent the work is small as well: record steps, tool calls, cost and time for each run, set each limit from the passing runs, and add one case that checks the agent stops and reports when a limit is reached.

References