AI agent evals: how to test an agent that takes actions / What to measure
Tool-call evals: did the agent pick the right tool and the right arguments
A tool call is one request an AI agent sends to another program: the name of a tool and the values to pass to it. A tool-call eval grades that single step with five checks. Was a tool needed, was the right tool named, do the arguments fit the tool's written definition, are the argument values correct, and did the agent respond sensibly when the tool returned an error. Four of the five can be decided by code, which makes them cheap to run on every change. The Berkeley Function Calling Leaderboard checks the name and the arguments this way, and the tau-bench paper found in 2024 that its best agent usually chose the right tool and filled in one or more arguments incorrectly.
Published September 30, 2026. Editorial.
Key takeaways
- A tool-call eval grades one step of an agent with five checks: was a tool needed, was the right tool chosen, do the arguments fit the tool's written definition, are the values correct, and was a tool error handled.
- Four of the five checks are comparisons that code can make, so they can run on every change without a model grader or a person.
- Strict mode is a vendor setting that makes a tool call fit the tool's written definition. By OpenAI's and Anthropic's own documentation it covers the shape of the arguments. Tool choice and the correctness of each value still need their own checks.
- In the 2024 tau-bench paper the gpt-4o agent usually chose the right tool and filled in one or more arguments incorrectly. It made 0.46 calls per retail task with identifiers, such as order numbers, that did not exist in the test data.
- Include cases where the correct action is to call no tool. The Berkeley Function Calling Leaderboard tests this and expects no function call as the output.
In June 2024 the authors of the tau-bench benchmark read through the failed runs of the best agent in their paper. The agent was built on gpt-4o and worked as a customer service assistant for a simulated shop. Of 115 runs on the retail tasks, 40 failed, and the authors traced 36 of the 40 to the agent and 4 to faults in the instructions given to the simulated customer [1]. Their summary of the agent's failures is one sentence: the agent "usually makes the right type of tool call(s) but fills in one or more arguments incorrectly" [1].
That sentence describes a failure in one step of a long task: the right tool, sent a wrong value. A test that only reads the agent's final message can miss it, because the message may still say the job is done. A tool-call eval looks at the single step and asks whether the request the agent sent was the right request. This page sets out its five checks, which of them code can decide, and how to build a first set of cases.
What a tool call is and why it is the smallest thing to test
A large language model (LLM) produces text. To act on anything, it has to ask another program to do the work. A tool call is that request: the model writes the name of a tool, such as "look up order", and the values to send with it, such as the order number. Your software runs the tool and returns the result to the model. The Berkeley researchers who maintain a public test of this skill define it this way: "Function calling, also called tool use, refers to an LLM's ability to invoke external functions, APIs, or user-defined tools in response to user queries" [2]. An API (application programming interface) is the set of requests one program accepts from another.
Each tool comes with a written definition called a schema. The schema lists the tool's name, what it does, which values it accepts, the type of each value (a number, a date, a piece of text), and which values are required. The values sent with a call are called arguments.
An agent's run is a sequence of these requests, and the full record of one run is called a trace. Why testing an agent is different from testing one model answer explains how an error in one step affects the steps after it. The tool call is the smallest part of a run that can be right or wrong by itself, so it is the first place to test. The rest of the method is in the AI agent evals guide.
The five checks for one step
For a single step, a tool-call eval asks five questions in order.
| Check | How it is decided | Example failure |
|---|---|---|
| Was a tool needed? | Code: compare "called" or "did not call" with the case | A customer asks for opening hours and the agent calls the refund tool |
| Was the right tool chosen? | Code: compare the tool name with the accepted names | The agent calls "cancel order" when the case expects "look up order" |
| Do the arguments fit the schema? | Code: validate the request against the schema | A quantity is sent as the text "2" where a number is required |
| Are the argument values correct? | Code for exact values, judgment for free text | Order number 48213 is sent and the customer's order is 48231 |
| Did the agent handle a tool error? | Code for the next action, judgment for the wording | The tool answers "order not found" and the agent says the refund is done |
The shop, the tools and the order numbers in the table are an invented illustration.
Check 1: was a tool needed at all
Some requests need no tool. An agent that calls one anyway spends money and may change a record that should have stayed as it was.
The Berkeley Function Calling Leaderboard, known as BFCL, is a leaderboard (a public ranking table) of models on tool use, and it tests this case directly. Its first version included what the authors call relevance detection: "we design scenarios where none of the provided functions are relevant and supposed to be invoked. We expect the model's output to be no function call" [3]. The later paper says the benchmark measures "the ability of models to abstain" [2], meaning to decline to call a tool. This page uses the benchmark's method and quotes no model scores.
Copy that design. For every tool your agent has, write at least one case where the request looks similar to that tool's purpose and the correct action is to call nothing. The check is one line of code: the case passes when the step contains no tool call.
Check 2: was the right tool chosen
When a tool is needed, the first thing to compare is its name. BFCL's checking procedure starts there: it "first extracts the function name and verifies that it is consistent with the one in possible answer" [3]. OpenAI's documentation lists "Did the agent pick the right tool?" among the questions its trace grading is meant to answer [4].
Write each case so that the accepted tools are a stated list, because some requests can be met correctly by more than one tool. Put tools with similar purposes into the same cases: an order lookup and a customer lookup are a pair an agent can confuse.
The sources used for this page give no verified limit on how many tools one agent can handle. The tau-bench retail setting gave its agent 15 tools, 7 that change data and 8 that only read it [1], which describes one test and sets no maximum. Measure it on your own agent: add the tools and re-run the selection cases.
Check 3: do the arguments fit the schema
After the name, BFCL checks the arguments against the tool's definition in two steps. It confirms that "all required parameters, as identified by the "required" attribute in the function documentation, are present in the model output", and it "ensures that only parameters exist in the function doc are used, flagging the model hallucination outputs" [3]. A hallucination here means a field or a value the model made up. The paper calls this an Abstract Syntax Tree method: the call is broken into its parts (the name and each argument) and the parts are compared one by one [2].
Model vendors now offer a setting that enforces the schema while the model writes. OpenAI's documentation says that setting strict to true "will ensure function calls reliably adhere to the function schema, instead of being best effort", and recommends always turning it on [5]. Anthropic's documentation says its strict setting "guarantees Claude's tool inputs match your JSON Schema", and that without it the model "might return incompatible types ("2" instead of 2) or omit required fields" [6]. JSON (JavaScript Object Notation) is a text format for structured data. Each statement is a company's description of its own product.
Keep the schema check in your eval with the setting on. OpenAI's documentation lists cases where output can still fail to match the schema: when the model refuses a request for safety reasons, and when a response stops at the maximum length and is incomplete [7]. It also says strict mode is disabled when a fine-tuned model (one given extra training on a customer's own examples) calls several functions in one turn [5]. The check costs one validation per step.
Check 4: are the argument values correct
A request can fit the schema and still be wrong. The strict settings constrain the shape of the arguments: the right fields with the right types. Whether the order number is this customer's order number is a separate question. That reading of the vendor documentation is ours.
The tau-bench paper measured this failure. Its authors counted tool calls that used identifiers (for a user, a product, an order or an item) that were absent from the test database. The gpt-4o agent made 0.46 such calls per retail task, and the two gpt-3.5-turbo agents made 2.08 and 6.34 [1]. Those are 2024 models on one benchmark, and they show that this failure can be counted.
Test values in three ways:
- Exact values. Where the case has one correct value (an identifier, an amount, a date), compare in code. Decide in advance whether formats such as "2026-09-30" and "30 September 2026" both count.
- Existence. For every identifier, check in code that it exists in the test data. This catches invented identifiers in cases that have no expected answer.
- Free text. For a search query or a message to a customer, many wordings are right. Use a person or a model grader working from written scoring rules, as described in the three kinds of LLM eval.
Check 5: what the agent does when a tool returns an error
A tool can fail: a record is missing, a service is busy, a value is rejected. Anthropic's guide to building agents says an agent needs to learn the real state of the system from its environment at each step, and names tool call results as one place that information comes from [8]. An error message is one of those results.
The sources for this page report no measured rate for this failure, so the example is an invented illustration. Take a support agent for an invented furniture shop whose order tool answers "order not found". Three next actions are acceptable: look up the order a different way, ask the customer to confirm the number, or stop and pass the conversation to a person. Three are failures: sending the identical call again, making up the contents of the order, or telling the customer the task succeeded.
To test this, replace the real tool with a version that returns a scripted error, which test environments for agents describes. The next action can be checked in code: the case fails if the following call is identical to the failed one, or if the run ends with a success message. An agent that repeats a failed call again and again is in a loop, and cost, time and step limits are the control that ends it.
Why these checks are cheap, and what they leave out
Four of the five checks are comparisons that code can make, with no model grader and no person, so they can run on every change. A model grader costs money on every decision, which the real price of an LLM judge sets out with public figures.
Passing every single-step check is weaker evidence than a passing task. The BFCL paper's own finding is that leading models "excel at singleturn calls" while memory and long multi-step reasoning "remain open challenges" [2]. Whole-task grading is covered in outcome evals vs trajectory evals, and the business overview is how to evaluate an AI agent before you trust it.
A tool-call eval judges a step where the case defines the right request. Anthropic's guide to agent evals warns that checking a fixed sequence of calls across a whole task gives tests that fail when an agent finds another valid route [9]. Use step checks where one answer is right, and grade the route only for rules that must always hold.
How to build a first tool-call set
- List every tool and its schema, and mark the tools that change data.
- For each tool, write cases of four types: a clear request for that tool, a request that belongs to a similar tool, a request that needs no tool, and a request followed by a scripted error.
- Take the inputs from real requests where you have them. LangSmith's documentation describes an agent test example as holding "correct tool selection and proper argument formatting" [10], and Anthropic suggests starting an agent suite with 20 to 50 simple tasks drawn from real failures [9].
- Write the accepted tools, the accepted values and the forbidden next actions into each case before running anything.
- Run each case several times, for the reason agent reliability across repeated runs explains.
- Report each check as its own pass rate, because wrong tool choice and wrong values need different fixes.
How Reveneau applies this
At Reveneau, all code is written by AI, and every change must pass a large eval suite that is written from the specification before the code exists. When the product is an agent, tool-call cases are part of that suite. We write the accepted tools, the accepted argument values and the forbidden next actions for each case from the specification, before the prompt and the tool definitions are written. The four checks that code can decide run on every change. The judged checks, such as whether a free-text argument says the right thing, are graded by Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To have tool-call evals written for an agent you are planning, see AI development at Reveneau or contact us.
Best for
- Agents with tools that change data, where a wrong argument has a cost
- Teams that want checks cheap enough to run on every change
- Finding out whether failures come from tool choice or from argument values
Avoid if
- The assistant calls no tools and only writes answers
- Step checks would be the only test, with no grading of the whole task
Check before you decide
- Every case states the accepted tools and values before the run
- The set includes requests where the right action is to call nothing
- Each of the five checks is reported as its own pass rate
Common questions
What is a tool call in an AI agent?
A tool call is a request an AI agent sends to another program, made of a tool's name and the values to pass to it. The model writes the request, your software runs the tool and returns the result, and the model reads that result before choosing its next step. A tool-call eval grades this single request: whether a tool was needed, which tool was named, and whether the values were valid and correct.
What is function calling, and is it different from tool use?
Function calling and tool use are two names for the same ability. The Berkeley Function Calling Leaderboard paper defines function calling, which it says is also called tool use, as a language model's ability to invoke external functions, application programming interfaces or user-defined tools in response to a user's request. Vendors use one term or the other in their documentation, and a test written for one applies to the other without change.
How do I test whether an agent chooses the right tool?
Test tool selection by writing cases where the accepted tools are a stated list, then comparing the tool name the agent sent with that list in code. The Berkeley benchmark's checking procedure starts the same way, by extracting the function name and comparing it with the expected one. Include tools with similar purposes in the same cases, because an order lookup and a customer lookup are a pair an agent can confuse.
How do I test the arguments an agent sends to a tool?
Test tool arguments in two stages. First validate the request in code against the tool's written definition, called its schema: every required value present, no invented fields, each value of the right type. Then compare the values with the case's expected values, and check that every identifier exists in the test data. The tau-bench paper counted calls with non-existent identifiers and found 0.46 per retail task for its gpt-4o agent in 2024.
Does strict mode or structured outputs make tool-call evals unnecessary?
No. Strict mode, by OpenAI's and Anthropic's own documentation, makes the arguments match the tool's written definition, called its schema, which covers one of the five checks. A request can fit the schema and still name the wrong tool or contain a wrong order number. OpenAI's documentation also lists cases where output can fail to match, including a safety refusal and a response that stops at the maximum length, so the schema check stays in the suite as well.
What should an agent do when a tool returns an error?
When a tool returns an error, an agent should read the error and change its next action: try a different lookup, ask the user to confirm a value, or stop and hand the task to a person. Sending the identical call again, making up a result and reporting success are failures. Test the behaviour by replacing the tool with a version that returns a scripted error and checking the following step in code.
How many tools can one agent handle?
The sources used for this guide give no verified maximum number of tools for one agent. The tau-bench retail tests gave the agent 15 tools, 7 that change data and 8 that only read it, which describes one benchmark setup and sets no limit. Measure the answer on your own agent: add the tools, re-run the tool selection cases, and watch whether the selection pass rate falls.
How much do tool-call evals cost to run?
Tool-call evals cost less than most agent tests because four of the five checks are decided by code: a name against a list, a request against a schema, a value against an expected value, and a next action against a list of forbidden ones. Only free-text arguments need a person or a model grader. The remaining cost is the model usage of running each step, multiplied by the number of repeats.
How is a tool-call eval different from a trajectory eval?
A tool-call eval grades one step against a case that defines the right request for that step. A trajectory eval compares the whole sequence of steps in a run with an expected sequence. Anthropic's guide to agent evals warns that checking a fixed order of calls across a task produces tests that fail when the agent finds another valid route, so step checks suit moments with one right answer.
What does the Berkeley Function Calling Leaderboard measure?
The Berkeley Function Calling Leaderboard, known as BFCL, measures how accurately language models call tools. Its first version checked the function name, the required arguments and any invented arguments by comparing the parts of each call, and included cases where no tool should be called. The 2025 paper about the benchmark reports that leading models do well on single calls while long multi-step work is still unsolved. This guide uses the benchmark's method and quotes no model scores.
Do I need tool-call evals if my assistant only answers questions?
Tool-call evals apply from the moment an assistant can search, read a record or change anything through another program. An assistant that only writes answers from its instructions makes no tool calls, and the methods for grading text answers cover it. When tools are added, start with the ones that change data, because a wrong call to those costs more than a wrong read, and include cases where the right action is to call nothing.
What should I test after the tool-call checks pass?
After the tool-call checks pass, test whole tasks. The Berkeley benchmark paper found that leading models do well on single calls while long multi-step work is still unsolved, so correct steps alone leave the task result unproven. Grade the final state the agent leaves behind, repeat every case several times, and set limits on steps and cost. Each of those has its own page in the AI agent evals guide.
References
- [1] Yao, Shinn, Razavi and Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, submitted 17 June 2024): of 115 gpt-4o retail runs 40 failed, 36 traced to the agent and 4 to the user instruction; the agent usually made the right type of tool call and filled in arguments incorrectly; 0.46, 2.08 and 6.34 calls with non-existent identifiers per retail task; retail has 7 write tools and 8 non-write tools.
- [2] Patil and six co-authors, The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models (Proceedings of Machine Learning Research volume 267, ICML 2025): the definition of function calling, the Abstract Syntax Tree checking method, the test of abstaining, and the finding that single-turn calls are strong while multi-step work remains open.
- [3] Gorilla LLM team, UC Berkeley, Berkeley Function Calling Leaderboard blog on version 1 (page last updated 19 August 2024): the three checking steps (function name, required parameters, invented parameters) and relevance detection, where no function should be called.
- [4] OpenAI, Evaluate agent workflows (API documentation, undated, read 30 September 2026): the questions trace grading is meant to answer, including whether the agent picked the right tool.
- [5] OpenAI, Function calling (API documentation, undated, read 30 September 2026): OpenAI's own description of strict mode, its recommendation to enable it, and the exception for fine-tuned models calling several functions in one turn.
- [6] Anthropic, Strict tool use (Claude Platform documentation, undated, read 30 September 2026): Anthropic's own description of strict tool use and of the type and missing-field errors that occur without it.
- [7] OpenAI, Structured model outputs (API documentation, undated, read 30 September 2026): the cases where output can still fail to match the schema: a refusal, or an incomplete response at the maximum length.
- [8] Anthropic, Building effective agents (19 December 2024): agents need to learn the real state of the system from the environment at each step, such as tool call results.
- [9] Anthropic, Demystifying evals for AI agents (9 January 2026): the warning against checking a fixed sequence of tool calls, and the suggestion to start with 20 to 50 simple tasks drawn from real failures.
- [10] LangChain, Evaluation concepts (LangSmith documentation, undated, read 30 September 2026): what an agent dataset example holds: correct tool selection and proper argument formatting.
Related reading
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
How to write an acceptance test a machine can run
Most acceptance criteria are written for a human reader who will add the missing details. A machine adds nothing. Here is how to write the sentence so an automated check can enforce it.
More in What to measure
Agent reliability: why one passing run proves little
An agent that passes a task once can fail the same task on the next run. The tau-bench paper of June 2024 measured this with pass^k, the chance that all k runs of a task pass: its best agent passed 61.2 percent of single retail runs and had a pass^8 below 25 percent. One passing run is therefore weak evidence. Run every case several times, at least 5 before a release decision, and report two numbers: the pass rate across all runs and the share of tasks that passed every run. A user of the released product gets one run, so the second number describes what users experience.
Cost, time and step limits as agent evals
A run that returns the right result can still be a failed run if it took too many steps or cost too much. An agent eval should record model usage cost, elapsed time, steps and tool calls for every run, and fail any run that exceeds a written limit on one of them. The 2024 paper AI Agents That Matter found coding agents of similar accuracy whose cost differed by almost two orders of magnitude (two orders is a factor of one hundred), so accuracy reported without cost cannot be compared. No source gives a recommended limit as a number: set each one from your own passing runs. METR's time horizon figures, quoted here with their dates, show how long a task models complete.