Taking AI agents from prototype to production / Make it reliable
How to evaluate an AI agent before you trust it
Evaluation is the part of agent work that decides whether everything else is possible. Without a written definition of correct and a set of real cases to score against, you cannot tell whether a change helped, you cannot defend a launch decision, and you cannot notice when the agent quietly gets worse. It is also the part teams most often postpone, because it produces no visible progress in a demo.
Published August 8, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- Define correctness in writing before building: what must always be true, what must never happen, what is good enough.
- Build the evaluation set from real user requests, including the messy ones, not from invented examples.
- Score the outcome and the path separately. An agent can reach a right answer through a route you would not accept.
- Every production failure should become a permanent case in the set, so the same regression cannot happen twice.
Ask a team how they know their agent is improving and you learn quickly whether the project is going well. The strong answer sounds like "we score against four hundred real cases and last week's change moved us from this to that". The weak answer sounds like "it feels better".
Evaluation is what turns agent work from opinion into engineering. Here is how to build it.
Start with a written definition of correct
Before any scoring, write down what a correct run looks like. Not a vague goal. A specific, testable description with three parts.
What must always be true. The output is in the required format. Any figure it reports appears in the source it cites. It never invents a customer record.
What must never happen. It does not reveal another customer's data. It does not take an irreversible action without confirmation. It does not claim certainty it does not have.
What counts as good enough. This is the judgment part and it cannot be skipped. For open-ended tasks there are many acceptable answers, and you have to describe the qualities that separate acceptable from unacceptable: does it address the actual question, is it grounded in real data, would a competent person be satisfied.
Writing this down is often the moment a team discovers that different stakeholders disagree about what the agent is for. That disagreement is much cheaper to resolve now than after launch.
Build the set from real cases
The most common evaluation mistake is testing against examples the team invented. Invented examples reflect what the builders imagined, which is a much simpler mix of requests than real use produces.
Use real requests. If the agent is not live yet, use the real requests people currently send to whatever the agent will replace: support tickets, emails, search queries, internal forms. If you genuinely have none, run a limited internal release first with a human reviewing everything, purely to collect them.
Include the difficult ones deliberately. Ambiguous requests. Requests with contradictory information. Requests where the correct response is to refuse or ask a question. Requests that are hostile or attempt to manipulate the agent. Requests in the wrong language or format. Real traffic contains all of this, and an evaluation set without it will make an agent look far more ready than it is.
A few hundred well-chosen cases is a serious asset. Coverage of the variety of things users actually do matters more than raw count.
Score the outcome and the path separately
For agents specifically, the final answer is only part of what matters. An agent can produce a correct result while doing something you would never approve of on the way: calling a tool it should not have used, retrieving far more data than needed, or taking eleven steps for a two step problem.
Score both. Outcome quality asks whether the result was right. Path quality asks whether it got there acceptably: were the tool calls appropriate, was the number of steps reasonable, did it stay inside its permissions.
Path problems are early warnings. An agent taking a strange route to a right answer today will take that route to a wrong answer eventually.
How to score when there is no exact expected answer
Deterministic checks come first, because they are cheap and reliable. Is the output valid JSON. Does the cited figure match the source. Did it stay under the step limit. Did it avoid forbidden tools. A surprising amount of quality can be captured this way.
For the judgment part, two approaches work together. Human review is the reference standard: have a qualified person grade a sample against your written criteria. It is slow and expensive, so use it deliberately rather than continuously.
Model-assisted grading, where a separate model scores outputs against your criteria, scales far better and is genuinely useful. The important discipline is to check the grader against human judgment on a sample. If the model grader and your reviewers disagree often, the grader is not trustworthy yet and its scores should not drive decisions.
Run it continuously, not once
Evaluation has to keep running after launch. Agents change, because the model, the APIs, and the input distribution all change without you touching your code.
Run the evaluation set on a schedule and after every meaningful change. Track the score over time so a slow decline is visible before it becomes a complaint. Watch cost and latency alongside quality, because fixing an accuracy problem by adding steps has a price that should be a conscious tradeoff.
Every production failure becomes a permanent case
This habit adds more value over time than any other. When the agent fails in production, do not just fix it. Add that exact case to the evaluation set first.
The set then grows to match the failures that real use produces, and any fix is verified against the real failure rather than a reconstruction of it. Most importantly, that failure can never silently return, because the set now tests for it permanently. Over a year, this turns into a body of knowledge about how your agent fails that no amount of upfront design would have produced.
An AI agent in production engagement builds this scoring in from the start, because teams need exactly this visibility into whether their AI systems are improving or regressing. The tooling varies; the discipline does not.
What good looks like
A team with evaluation under control can answer these without hesitation: how often is the agent right, on what set of cases, how has that changed over the last month, what are the top three remaining failure modes, and what happens when it is wrong.
Those answers are what make a launch decision defensible. Without them, launching an agent is a guess rather than a decision based on evidence, which is exactly why so many agents never launch at all.
Common mistakes worth avoiding
A few patterns show up often enough to name.
Evaluating only normal use. A set made of reasonable requests will tell you the agent is ready when it is not. The value is concentrated in the awkward cases, so a set without them is measuring the wrong thing.
Treating the score as the goal. Once a number exists, there is pressure to move it, and it is possible to improve a score in ways that do not improve the product, particularly if cases get copied into the instructions the agent is given. Keep some cases held back and review real outputs directly rather than trusting the aggregate alone.
Averaging away the failures that matter. An overall accuracy figure hides the distribution. An agent that is right ninety five percent of the time may be failing almost entirely on one specific category of request, which matters much more than the overall number. Break scores down by request type.
Evaluating once. A single pre-launch measurement tells you about a system that no longer exists a month later, since the model, the dependencies, and the input mix all change. Scheduled runs are what turn evaluation from a one-time check into an ongoing control.
Waiting for perfect before starting. A rough evaluation set built this week is worth more than a comprehensive one planned for next quarter. Start with fifty real cases and grow it from production failures.
A worked example: scoring a support agent
Consider an agent that answers account questions by looking up a customer's record and replying in plain language. A written definition of correct for this task might read: the reply must state only facts present in the record it retrieved, it must not disclose another customer's information under any circumstance, and it must ask a clarifying question rather than guess when the request is ambiguous about which account is meant.
Given that definition, a deterministic check can catch a large share of failures automatically. Does the reply mention an account number that differs from the one the customer is authenticated as. Does every dollar figure in the reply match a figure present in the retrieved record. Did the agent call the lookup tool with a customer identifier rather than a name, since names are ambiguous and identifiers are not. Each of these is a programmatic check with no need for human judgment, and running them against every case in the set catches a meaningful share of problems before anyone has to read an output.
The judgment layer sits on top of that. Some fraction of replies will pass every deterministic check and still read poorly: technically accurate but confusing, or accurate but missing the part of the question the customer actually cared about. That is where a human reviewer, or a model grader validated against human reviewers, earns its cost, because no fixed rule can capture "did this actually answer the question the customer asked."
Evaluating the tools separately from the agent
A failure that looks like an agent reasoning problem is often a tool problem, and evaluation should be structured to tell the two apart, because the fix for each is completely different.
Build a small, separate test for each tool the agent calls: given this input, does the tool return the expected result, and does it fail cleanly and informatively on a bad input. If a tool itself is unreliable, slow, or inconsistently formatted, no amount of agent-level evaluation will isolate that, because every agent-level failure caused by the tool will look like a reasoning mistake in the aggregate score. Testing tools independently means a bad score on the full agent evaluation can be traced back to a specific broken component rather than treated as a vague sign the model needs to be smarter. Permissions and guardrails for AI agents covers the related discipline of validating what a tool accepts and rejects, which is worth building alongside the tool-level tests described here.
Regression testing across model and prompt changes
Every change to the model version, the prompt, or a tool is a chance to make things better or worse, and the only way to know which happened is to run the same evaluation set before and after the change and compare.
This sounds obvious and is nonetheless the step teams skip under time pressure, particularly when a model provider ships an update that appears to be a strict improvement. Model providers describe their own updates in aggregate terms, across benchmarks that may have nothing to do with your specific task, so a general improvement can still be a regression on your narrow use case. Running your evaluation set against the new model before switching, rather than after, is the only way to catch that before your users do. Keep the previous model version available to switch back to quickly if the new one scores worse, because a rollback plan decided during the emergency is a worse rollback plan than one decided in advance.
What a good evaluation report looks like
A team with evaluation under control should be able to produce something concrete when asked, not just describe the practice in the abstract. A useful report names the size and composition of the test set, states the overall score and the score broken down by request category, lists the specific cases that failed and a one-line reason for each, and compares the current score against the previous version so a reader can see the direction of movement rather than only a snapshot.
That report is the artifact that turns an internal impression into something a stakeholder outside the engineering team can evaluate on its own terms. It is also what makes a launch decision defensible after the fact, if a regulator, a customer, or an internal audit ever asks how the team knew the agent was ready.
Evaluation for agents that act, not just agents that answer
Everything above focuses on the content of a reply, but an agent that takes actions, sending a message, updating a record, issuing a refund, needs evaluation of the action itself, not only the text that explains it.
For an acting agent, build cases where the correct behavior is to take no action at all, and score whether the agent recognizes that. This category is easy to skip, because most evaluation examples are naturally built around cases where some action is expected, and an agent that has never been tested on "the right answer here is to do nothing" will often invent a reason to act rather than sit still. That specific failure, acting when the correct behavior was restraint, is one of the more expensive categories of agent mistake precisely because it is rare enough in a typical test set to go unnoticed until it happens with a real customer.
This page is the overview. The method for each part, from grading tool calls and repeating every run to setting cost limits and testing for attacks, is in our guide to AI agent evals.
Best for
- Any agent that real users will rely on, where a wrong answer has a cost
- Teams that need to defend a launch decision with evidence rather than impressions
- Systems expected to run for months, where slow, unnoticed decline is the main risk
Avoid if
- The task genuinely has no definition of correct, in which case solve that before building
- You are still deciding whether an agent is the right approach at all
Check before you decide
- Confirm the evaluation set is built from real requests, not invented examples
- Confirm any model-assisted grader agrees with human reviewers on a sample
- Confirm the set runs on a schedule, not only before releases
Common questions
How many test cases does an AI agent evaluation set need?
Coverage matters more than count. A few hundred cases that represent the real variety of requests, including ambiguous, contradictory, and hostile ones, is far more useful than thousands of similar easy examples. Grow the set from real production failures rather than trying to guess the right size upfront.
Can you use a model to grade another model's output?
Yes, and it scales much better than human review. The discipline is to validate the grader: have qualified people grade a sample and compare. If the grader disagrees with human judgment often, its scores should not be driving decisions until that gap closes.
What do you measure besides whether the answer was right?
Path quality, cost, and latency. An agent can reach a correct answer through an unacceptable route, such as calling tools it should not use or retrieving more data than needed, and that is an early warning. Cost and latency matter because accuracy fixes often add steps, and that tradeoff should be deliberate.
When should evaluation start?
Before the build. Defining what correct means is what tells you whether the task is even suitable for an agent, and it frequently brings out disagreement between stakeholders about the agent's purpose. Resolving that early is much cheaper than discovering it after launch.
What is the difference between evaluating a workflow and evaluating an agent?
A fixed workflow can be tested by asserting that a given input produces a given output, the way ordinary software is tested. An agent can take a different path on a similar request from one day to the next, so evaluation means scoring the quality of outcomes and the path taken across many real cases rather than asserting one exact result.
Is it risky to launch an AI agent without an evaluation set?
Yes. Without a written definition of correctness and real cases to score against, there is no way to tell whether the agent is improving, no way to defend a launch decision with evidence, and no way to notice when the agent quietly gets worse after models or dependencies change. That risk shows up as a slow, unmeasured decline rather than a single visible failure.
How much does it cost to build an evaluation set for an AI agent?
It is genuine, ongoing work rather than a one-time setup cost: defining correctness, collecting real cases, building the scoring, and validating that the scoring agrees with human judgment. It produces no visible progress in a demo, which is why teams underestimate it, but a rough set of fifty real cases built this week is worth more than a comprehensive one planned for later.
Can an AI agent be trusted if it passes its evaluation set once?
Not indefinitely. A single pre-launch measurement describes a system that no longer exists a month later, because the model, the dependencies, and the mix of real requests all shift over time. Running the evaluation set on a schedule, not only before releases, is what turns a one-time score into an ongoing basis for trust.
References
Related reading
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.