AI agent evals: how to test an agent that takes actions / What to measure
Cost, time and step limits as agent evals
A run that returns the right result can still be a failed run if it took too many steps or cost too much. An agent eval should record model usage cost, elapsed time, steps and tool calls for every run, and fail any run that exceeds a written limit on one of them. The 2024 paper AI Agents That Matter found coding agents of similar accuracy whose cost differed by almost two orders of magnitude (two orders is a factor of one hundred), so accuracy reported without cost cannot be compared. No source gives a recommended limit as a number: set each one from your own passing runs. METR's time horizon figures, quoted here with their dates, show how long a task models complete.
Published September 30, 2026. Editorial.
Key takeaways
- Record five things for every agent run: model usage cost, elapsed time, number of steps, number of tool calls, and whether the task passed. Put a limit on each of the first four.
- The 2024 paper AI Agents That Matter found that, on a coding benchmark of 164 problems, cost differed by almost two orders of magnitude (two orders is a factor of one hundred) between agents of similar accuracy.
- A step limit ends a loop, which is an agent repeating the same step. Anthropic's guide to building agents names a maximum number of iterations as a common stopping condition.
- No source recommends a step or cost limit for a product as a number. Set each limit from the largest value among your own passing runs, plus a margin you write down.
- METR's paper reports that models' 80 percent time horizons are 4 to 6 times shorter than their 50 percent horizons, so a higher reliability target means giving an agent shorter tasks.
HAL is a leaderboard (a public ranking table) of AI agents run by a team at Princeton University. As read on 30 September 2026, it listed two 2025 runs on a benchmark (a public set of tasks with a fixed scoring method) named AssistantBench. Both rows used the same agent software, Browser-Use. With the o3 model at its medium setting (April 2025) the agent scored 38.8 percent, and the cost shown beside the score was 15.15 US dollars. With GPT-5 at its medium setting (August 2025) it scored 35.2 percent at 41.69 US dollars [1]. The second row cost 2.75 times as much as the first (41.69 divided by 15.15) and scored lower.
Two rows show one comparison and no trend. The site says its leaderboard is paused, so the figures are a record of those 2025 runs and say nothing about current models. A report that listed accuracy alone would hide which run was the costly one. This page is about putting cost, time and step counts into the eval itself, each with a limit that fails a run when it is exceeded.
Why accuracy alone is the wrong score for an agent
The main evidence is a July 2024 paper titled "AI Agents That Matter", by Sayash Kapoor, Arvind Narayanan and three co-authors at Princeton University. They ran published coding agents five times each on HumanEval, a benchmark of 164 programming problems, and recorded the cost of each agent beside its accuracy [2]. Three of their results:
- For accuracy the authors describe as similar, cost differed by "almost two orders of magnitude" [2]. Two orders of magnitude is a factor of one hundred.
- Two complex agents, Reflexion and LDB, cost over 50 percent more than a simple comparison method the authors call the warming strategy, and a third, LATS, cost over 50 times more [2].
- The authors found "no significant accuracy difference between our warming strategy and the best-performing agent architecture" [2].
Their conclusion is that "useful agent evaluations must control for cost", because accuracy "can be improved by scientifically meaningless methods such as retrying" [2]. The models were from the GPT-3.5 and GPT-4 period and the result comes from one benchmark, so the exact ratios belong to that test. The reasoning applies to any agent: a score can be raised by spending more, so a score reported without its cost cannot be compared with another.
A second study, the HAL paper of October 2025, covered 21,730 agent runs across 9 models and 9 benchmarks, at a total cost the authors give as 40,000 US dollars and mark as approximate. In 21 of 36 combinations of model, agent and benchmark, a higher reasoning effort setting, which lets the model do more computation before it answers, produced "equal or lower accuracy" [3].
What to measure on every run
Record five things for each run of each task.
- Model usage cost. Models are billed by the token, a word piece the model reads or writes. Cost is the tokens used multiplied by the provider's price, added up over every model call in the run.
- Wall-clock time. The elapsed time from the first request to the final result.
- Number of steps. One step is one turn of the agent's loop: the model chooses an action, the action runs, the model reads the result.
- Number of tool calls. A step can contain more than one call, and a tool can have its own charge.
- Whether the task passed.
Anthropic's guide to agent evals lists similar quantities as things a fixed set of tasks lets a team track: "latency, token usage, cost per task, and error rates" [4]. Latency is the waiting time for a response. The Berkeley Function Calling Leaderboard publishes an estimated cost in US dollars and latency in seconds beside each model's accuracy [5].
Report the spread as well as the average. A February 2026 paper on agent reliability ran each task 5 times and found "high variance in token and compute usage across runs" [6]. Run each task several times, as agent reliability across repeated runs describes, and record the middle value and the largest value across the runs.
A limit turns a measurement into a pass or a fail
A measurement becomes an eval when it has a threshold. Set one limit for each measure, and count a run that exceeds it as failed even when the result is correct.
| Limit | What it protects against | How to set it |
|---|---|---|
| Step limit | A loop, and runs that take many needless steps | From the step counts of passing runs, plus a margin |
| Tool-call limit | Repeated calls to a service that charges per call or restricts how many it accepts | For each tool, from passing runs |
| Cost limit per task | A run that costs more than the task is worth | From the value of the task to the business |
| Time limit | A user or another system left waiting | From how long the user will wait |
Published tests set limits of this type. tau-bench stopped each task at 30 agent actions [7]. The OSWorld benchmark, which tests agents on real desktop applications, set "the maximum steps of interaction to 15 and the maximum time limits to 30 minutes for all tasks" [8]. The authors of SWE-Agent, as cited in "AI Agents That Matter", set a cost limit of 4 US dollars per task [2]. METR runs its agents "up to a token and time limit" [9].
Researchers chose those values for a benchmark. No source used for this page recommends a step or cost limit for a product as a number, so any figure for your product comes from your own runs. The method:
- Run the task set with repeats, with only a high safety limit in place.
- Look at the passing runs alone. Note the largest step count, cost and time among them.
- Set each limit above that largest value by a margin you choose and write down.
- Run again, and read every run that now fails on a limit alone.
Here is an invented illustration. Suppose the passing runs of a refund task took between 6 and 11 steps across 40 runs. A step limit of 15 allows an unusual valid route and still ends a run that has stopped making progress.
How a step limit stops a loop
A loop is an agent repeating the same step: it calls a tool, gets a result it cannot use, and makes the identical call again. Each repeat costs tokens and time, and with no stopping rule the run continues until something outside the agent ends it.
Anthropic's guide to building agents names the control: "it's also common to include stopping conditions (such as a maximum number of iterations) to maintain control" [10]. An iteration is one turn of the loop. The same guide says the autonomous nature of agents "means higher costs, and the potential for compounding errors" [10], and a loop produces both at once.
In an eval, use the limit in three ways:
- End the run. When the step count reaches the limit, the run stops and is recorded as failed, with the reason "step limit reached".
- Detect the repeat earlier. A code check on the recorded steps can flag the same tool called with the same arguments several times in a row, before the limit is reached. A tool error is one situation where an agent can start repeating itself, which is why error handling is one of the five checks in tool-call evals.
- Check what happens at the limit. An agent that reaches its limit should stop, say what it completed, and hand the task to a person. Write a case for that behaviour.
Limits work beside other controls, which are covered in permissions and confirmation steps for AI agents.
How much an agent run costs
Cost depends on the task, the model and the date, so the only figure you can rely on is one you measure. Published figures show the range.
- tau-bench, 2024. With a gpt-4o agent and a gpt-4 simulated user on the retail tasks, the authors reported 0.38 US dollars per task for the agent and 0.23 for the simulated user, and put the cost of one trial of every task at 200 dollars, a figure the paper marks as approximate [7]. Both figures are repeated as the paper states them: 0.61 dollars multiplied by the 115 retail tasks is 70.15 dollars, and the passage does not explain the difference.
- SWE-Agent, 2024. With a limit of 4 US dollars per task, running the agent on the entire benchmark "could cost over USD 8,000 for a single evaluation run", as cited in "AI Agents That Matter" [2].
- HAL, 2025. 21,730 runs for a total the authors give as 40,000 US dollars and mark as approximate [3]. Dividing that total by the count gives 1.84 dollars per run. That is an estimate, because the total is approximate, and it is an average across unlike benchmarks.
One pass of an eval set costs the number of cases, times the repeats for each case, times the cost of one run. As an illustration with an invented case count: 50 cases, 5 repeats each, at the 0.61 dollars per task reported for tau-bench, comes to 50 x 5 x 0.61 = 152.50 dollars for one pass. The cost of the product itself is covered in what it costs and how long it takes to release an AI agent, and the budget for evals in what AI evals cost and who should own them.
How long a task an agent can finish: METR's time horizon
A time limit depends on how long a task an agent can complete at all. METR, a research organisation, measures this with a figure it calls the time horizon. Its paper defines the 50 percent version as "the time humans typically take to complete tasks that AI models can complete with 50% success rate" [11]. METR's page adds that the figure is "a measure of the difficulty of a task, rather than the time an AI spends to complete the task" [9]. A two-hour horizon means the agent succeeds half the time on tasks that take a skilled person two hours.
METR's figures differ between its publications, so each is given with its source and date [11] [12] [13] [9].
| Source | Date | What it states |
|---|---|---|
| Paper, abstract | First version 18 March 2025, read in its fourth version of 10 July 2026 | Claude 3.7 Sonnet has a 50 percent horizon of 50 minutes, and the horizon has doubled every seven months from 2019 on. The abstract marks both figures as approximate. |
| Paper, body | Same paper | Doubling time of 207 days for the 50 percent horizon over 2019 to 2025, and 204 days for the 80 percent horizon |
| Blog post | 19 March 2025 | The same model's horizon is one hour, which the post marks as approximate |
| Time Horizon 1.1 | 29 January 2026 | Task set grew from 170 to 228 tasks. Doubling time 196.5 days over the whole period, and 130.8 days from 2023 on (interval 107 to 161) |
| Data file | Read 30 September 2026 | Doubling time 187.778 days over the whole period, and 128.744 days from 2023 on (interval 104.428 to 158.012) |
Three facts from these sources matter more than any single value.
- Higher reliability means a shorter horizon. The paper reports that models' 80 percent horizons are 4 to 6 times shorter than their 50 percent horizons [11].
- Success falls as tasks get longer. In March 2025 METR reported that the models of that date succeeded on almost all tasks that take a person less than 4 minutes, and on fewer than 10 percent of tasks longer than a threshold that METR gives as 4 hours and marks as approximate [12].
- The estimates are wide and they change. METR's data file, as read on 30 September 2026, gives the model listed as gpt_5_4, released on 5 March 2026, a 50 percent estimate of 341.74 with an interval of 186.58 to 768.78, and an 80 percent estimate of 53.88. The file prints no unit, and the scale matches minutes. METR's January 2026 post says its confidence intervals are still wide [13], and its page says measurements above 16 hours "are unreliable with our current task suite" [9].
METR's March 2025 post describes its tasks as multi-step software and reasoning tasks [12], so a horizon for another type of work has to be measured on that work. For your limits, the practical reading is this: if a task takes a skilled person several hours, plan for a lower pass rate than on short tasks, test it with repeats, and set the time and step limits from passing runs. Agent benchmarks explained covers the benchmark itself.
Put the same limits into the released agent
The released agent should stop at the same step, cost and time limits that the eval used, so that the eval describes how the product behaves. A production run that reaches a limit is also a candidate for a new eval case, as turning production traces into eval cases describes. The rest of the method is in the AI agent evals guide. A related post covers what founders get wrong about AI agents.
How Reveneau applies this
At Reveneau, all code is written by AI, and every change must pass a large eval suite written from the specification before the code. For an agent, that suite records cost, time, steps and tool calls for every run, and each task has a written limit for each of them. A run that returns the right result over its limit fails. We report cost and time beside the pass rate, so that a change which raises accuracy by spending more is visible as such. The run time of the eval suite matters as well: we grade the judged checks with Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To set limits for an agent you are planning, see AI development at Reveneau or contact us.
Best for
- Agents billed by usage, where a correct but costly run loses money
- Agents that can repeat a failed step
- Comparing two agent versions whose accuracy is close
Avoid if
- The plan is to copy a benchmark's limit, such as 30 actions, without measuring your own passing runs
- The limits would exist in the test set and be absent from the released agent
Check before you decide
- Every run records cost, time, steps and tool calls
- Each limit was set from passing runs and the margin is written down
- A case checks that the agent stops and reports at the limit
Common questions
How much does one AI agent run cost?
The cost of one agent run depends on the task, the model and the date, so measure it on your own tasks. Published examples: the tau-bench authors reported 0.38 US dollars per retail task for the agent and 0.23 for the simulated user in 2024, and the paper AI Agents That Matter cites a limit of 4 US dollars per task set by the authors of SWE-Agent. Both are benchmark figures for 2024 models.
How do I stop an AI agent from looping?
Stop an agent from looping with a step limit: a maximum number of turns after which the run ends. Anthropic's guide to building agents names a maximum number of iterations as a common stopping condition. Add a code check that flags the same tool called with the same arguments several times in a row, and test what the agent does at the limit: stop, report what was completed, and hand the task to a person.
What is a step limit for an AI agent?
A step limit is the maximum number of turns an agent may take on one task before the run is ended. One step is one turn of the agent's loop: choose an action, run it, read the result. Benchmarks set them: tau-bench stopped each task at 30 agent actions, and OSWorld at 15 steps and 30 minutes. Researchers chose those values for a test, and no source recommends a number for a product.
Should cost be part of an agent eval?
Yes. The 2024 paper AI Agents That Matter concludes that useful agent evaluations must control for cost, because accuracy can be raised by spending more, for example by retrying. On the HumanEval coding benchmark of 164 problems its authors found that cost differed by almost two orders of magnitude (two orders is a factor of one hundred) between agents of similar accuracy. Record cost for every run and report it beside the pass rate.
What is METR's time horizon?
METR's time horizon is the length of task, measured by how long it takes a human expert, that an AI agent completes at a given rate of success. The 50 percent horizon is the task length at which the agent succeeds half the time. METR's page states that the figure measures the difficulty of a task, which is a separate thing from how long the agent spends working on it.
How long a task can an AI agent finish?
METR's published answers differ by source and date. The abstract of its paper, first published in March 2025, gives Claude 3.7 Sonnet a 50 percent horizon of 50 minutes and marks the figure as approximate. METR's data file, read on 30 September 2026, lists a model released in March 2026 at 341.74, with an interval from 186.58 to 768.78; the file prints no unit and the scale matches minutes. Horizons at 80 percent reliability are 4 to 6 times shorter.
How do I choose the number for a step or cost limit?
Choose a limit from your own passing runs. Run the task set with repeats under a high safety limit, find the largest step count, cost and time among the runs that passed, and set each limit above that value by a margin you write down. No source used in this guide recommends a number for a product. Benchmark limits such as tau-bench's 30 actions were chosen by researchers for one test.
Does a higher reasoning setting give an agent better results?
A higher reasoning effort setting gave equal or lower accuracy in 21 of 36 combinations in one large study. The paper on HAL, a public ranking of AI agents run by a Princeton University team, covered 21,730 agent runs across 9 models and 9 benchmarks in October 2025. In those 21 combinations of model, agent and benchmark, the higher setting produced equal or lower accuracy. Measure accuracy and cost together on your own tasks before choosing the higher setting.
What is the difference between a step limit and a time limit?
A step limit counts the agent's turns and a time limit counts minutes on the clock. A step limit stops a loop, where the agent repeats the same action. A time limit protects a user or another system that is waiting, and it catches a slow tool that a step count would miss. The OSWorld benchmark set both: 15 steps and 30 minutes for every task.
What goes wrong if an agent eval has no limits?
Without limits, an eval counts a run as passed whatever it spent, and a loop can continue until something outside the agent ends it. The eval then rewards changes that raise accuracy by spending more. The authors of AI Agents That Matter found two complex coding agents costing over 50 percent more than the authors' simple comparison method, and a third over 50 times more. They report no significant accuracy difference between that method and the best-performing agent design.
How much does it cost to run an agent eval suite?
The cost of one pass of an agent eval suite is the number of cases, multiplied by the repeats per case, multiplied by the cost per run. As an illustration, 50 cases run 5 times each at the 0.61 US dollars per task reported for tau-bench in 2024 comes to 152.50 dollars. The paper AI Agents That Matter notes that one full run of SWE-Agent on its benchmark could cost over 8,000 US dollars.
Do these limits apply to a simple agent with one or two tools?
Yes, the limits apply to any agent that works in a loop, including one with two tools. A loop needs only one tool and one result the agent cannot use. For a small agent the work is small as well: record steps, tool calls, cost and time for each run, set each limit from the passing runs, and add one case that checks the agent stops and reports when a limit is reached.
References
- [1] SAgE team, Princeton University, HAL leaderboard site (undated, read 30 September 2026): two AssistantBench rows read on 30 September 2026 (Browser-Use with o3 Medium, 38.8 percent and 15.15 US dollars; with GPT-5 Medium, 35.2 percent and 41.69 US dollars), and the notice that the leaderboard is paused.
- [2] Kapoor, Stroebl, Siegel, Nadgir and Narayanan, AI Agents That Matter (arXiv, submitted 1 July 2024): agents run five times on 164 HumanEval problems; cost differing by almost two orders of magnitude at similar accuracy; Reflexion, LDB and LATS costs against the warming strategy; evaluations must control for cost; the SWE-Agent limit of 4 US dollars per task and over 8,000 US dollars for one full run.
- [3] Kapoor, Stroebl, Kirgis and 28 co-authors, Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation (arXiv, submitted 13 October 2025): 21,730 agent runs across 9 models and 9 benchmarks at a total cost the authors give as 40,000 US dollars, marked as approximate; equal or lower accuracy from higher reasoning effort in 21 of 36 combinations.
- [4] Anthropic, Demystifying evals for AI agents (9 January 2026): latency, token usage, cost per task and error rates can be tracked on a fixed set of tasks.
- [5] Gorilla LLM team, UC Berkeley, Berkeley Function-Calling Leaderboard live page (last updated 12 April 2026, read 30 September 2026): the leaderboard reports an estimated cost in US dollars and latency in seconds beside accuracy.
- [6] Rabanser, Kapoor, Kirgis, Liu, Utpala and Narayanan, Towards a Science of AI Agent Reliability (arXiv, submitted 18 February 2026): each task run 5 times; high variance in token and compute usage across runs.
- [7] Yao, Shinn, Razavi and Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, submitted 17 June 2024): a limit of 30 agent actions per task; 0.38 and 0.23 US dollars per retail task for the agent and the simulated user, and 200 dollars, marked as approximate in the paper, for one trial per task.
- [8] Xie and 16 co-authors, OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments (arXiv, first version 11 April 2024): a maximum of 15 steps and 30 minutes for every task.
- [9] METR, Task-Completion Time Horizons of Frontier AI Models, live page and data file (last updated 8 May 2026, read 30 September 2026): the definition of the time horizon as task difficulty; runs up to a token and time limit; measurements above 16 hours are unreliable; data file doubling times and the gpt_5_4 entry as read on 30 September 2026.
- [10] Anthropic, Building effective agents (19 December 2024): stopping conditions such as a maximum number of iterations; agents bring higher costs and compounding errors.
- [11] Kwa and 25 co-authors (METR), Measuring AI Ability to Complete Long Software Tasks (arXiv, first version 18 March 2025, fourth version 10 July 2026): the definition of the 50 percent time horizon; 50 minutes for Claude 3.7 Sonnet; doubling times of 207 and 204 days; 80 percent horizons 4 to 6 times shorter.
- [12] METR, Measuring AI Ability to Complete Long Tasks (blog post, 19 March 2025): one hour for Claude 3.7 Sonnet; success by task length for the models of March 2025.
- [13] METR, Time Horizon 1.1 (29 January 2026): task set grown from 170 to 228; doubling times of 196.5 and 130.8 days; confidence intervals still wide.
Related reading
How autonomous are AI coding agents, really?
Engineers at the leading AI labs now say a model writes one hundred percent of their code. Read the quotes closely and a person is still involved in every one of them. Here is what the 2026 evidence supports, and what it does not.
What founders get wrong about AI agents
An impressive agent demo and a reliable agent are two different things. Most of the work, and most of the risk, is in the final step before production, which nobody shows in the demo.
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
More in What to measure
Tool-call evals: did the agent pick the right tool and the right arguments
A tool call is one request an AI agent sends to another program: the name of a tool and the values to pass to it. A tool-call eval grades that single step with five checks. Was a tool needed, was the right tool named, do the arguments fit the tool's written definition, are the argument values correct, and did the agent respond sensibly when the tool returned an error. Four of the five can be decided by code, which makes them cheap to run on every change. The Berkeley Function Calling Leaderboard checks the name and the arguments this way, and the tau-bench paper found in 2024 that its best agent usually chose the right tool and filled in one or more arguments incorrectly.
Agent reliability: why one passing run proves little
An agent that passes a task once can fail the same task on the next run. The tau-bench paper of June 2024 measured this with pass^k, the chance that all k runs of a task pass: its best agent passed 61.2 percent of single retail runs and had a pass^8 below 25 percent. One passing run is therefore weak evidence. Run every case several times, at least 5 before a release decision, and report two numbers: the pass rate across all runs and the share of tasks that passed every run. A user of the released product gets one run, so the second number describes what users experience.