Guide

Taking AI agents from prototype to production

An agent that works in a demo and an agent that works in production are different systems. The demo has to succeed once, with someone friendly at the keyboard. Production has to succeed repeatedly, for people who did not read the instructions, with real money and real data affected by each call.

Published August 8, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • The gap between an agent demo and an agent in production is not model quality. It is evaluation, permissions, error handling, and the cost of being wrong.
  • You cannot release what you cannot measure. Before real users rely on an agent, you need a written definition of what a correct run looks like and a way to score it repeatedly.
  • Give an agent the narrowest permissions that let it do its job. Most serious agent incidents are permission problems, not intelligence problems.
  • Start with the simplest architecture that works. One agent with good tools works better than a multi-agent system almost every time at the start.
  • Design for the agent being wrong, because it will be. The question is whether a wrong answer is caught, contained, and recoverable.
  • Keep a human in the loop wherever an action is expensive to undo. Reversibility, not confidence, should decide what an agent is allowed to do alone.

Most teams building with AI right now have the same experience. Someone puts together an agent in a couple of days, it does something genuinely impressive in a demo, and the people watching get excited. Then the work of making it real begins, and it turns out to be much harder and much slower than the demo suggested. The agent that answered five questions beautifully answers the sixth one confidently and wrongly. It calls a tool with the wrong argument. It loops. It works fine for a week and then quietly stops working when an API it depends on changes its response format.

This is not a sign that the team is doing something wrong. It is the normal pattern for this kind of problem. An agent is a system that makes decisions, and every system that makes decisions needs the same things around it: a way to tell whether the decisions were good, limits on what it can do when the decisions are bad, and a plan for what happens when something breaks. None of that is needed for a demo. All of it is needed in production.

This guide is about that second part. It assumes you already believe an agent is the right approach for your problem, and it focuses on what it takes to run one with real users: how to evaluate it, how to keep it reliable, what permissions to give it, how to choose an architecture, and how to think about the budget and timeline honestly.

What an agent actually is, and why the definition matters

An agent is a system where a model decides what to do next, not just what to say next. The model is given a goal, a set of tools it can call, and some context, and it works through steps until it decides it is done. That is the whole idea. The important part is the word "decides", because that is what separates an agent from a normal software feature and it is where all the difficulty comes from.

In ordinary software, you write the steps. If the program does something wrong, the wrong behavior is in the code, and you can find it and fix it. In an agent, the steps are chosen at runtime, and the same input can produce a different path on a different day. You are no longer debugging a fixed sequence. You are trying to shape the behavior of something that reasons, and that requires different methods.

This distinction matters commercially too. Plenty of products marketed as agents are really a model call inside a fixed workflow, which is a perfectly good thing to build and often the correct choice. It is worth being honest internally about which one you are building, because a fixed workflow with a model in it is much easier to test, much cheaper to run, and much less likely to surprise you. If a predictable sequence solves the problem, build the predictable sequence. Use an agent when the sequence genuinely cannot be known in advance.

Why agents stall between demo and production

The pattern is consistent enough to describe. A demo is judged on whether an impressive thing happened. Production is judged on whether an unacceptable thing never happens. Those are almost opposite standards, and moving from one to the other is where most of the work is.

A demo runs a handful of times, usually with inputs chosen by the person who built it, in front of an audience willing to interpret a rough answer generously. Nothing is at stake. If it fails, someone reruns it.

Production runs constantly, with inputs nobody anticipated, for users who will not excuse its mistakes, in situations where a wrong action can cost money or trust. The agent will meet inputs that are ambiguous, contradictory, malicious, or simply strange. It will meet an API that times out. It will meet a user who asks it to do something it should refuse.

The work that gets an agent from demo to production is plain and unglamorous: a test set, a scoring method, permission boundaries, logging, retries, fallbacks, and a decision about what a human still has to approve. None of them make a good demo. All of them decide whether the agent keeps working in production.

Start by writing down what "correct" means

This is the step teams skip, and skipping it is why so many agent projects run for months without ever being ready. Before you improve an agent, you need to be able to tell whether a change made it better or worse. That means writing down, in plain language, what a correct run looks like.

For some tasks this is easy. If the agent is extracting a date from a document, correctness is whether the date is right. For most real tasks it is harder, because there is more than one acceptable answer and the difference between good and bad is a matter of judgment. That does not mean it cannot be defined. It means you have to do the work of defining it: what must always be true, what must never happen, and what counts as good enough.

Once you have that, collect real cases. Not invented ones. Pull the actual requests users are making, including the messy and unreasonable ones, and build a set you can run the agent against repeatedly. This set is the single most valuable asset in an agent project, and its value grows over time. Every time the agent fails in a new way in production, that failure becomes a new case in the set, and the agent can never quietly regress on it again.

An AI agent in production engagement makes scoring the main focus for exactly this reason: teams need visibility into whether their AI systems are getting better or worse. The pattern holds regardless of tooling. You cannot improve what you are not scoring.

Reliability is a design problem, not a model problem

When an agent behaves badly, the first reaction is to try a better model or a better prompt. Sometimes that helps. More often the lasting fix is structural, because the failure was not really about intelligence.

Narrow the job. An agent asked to do one well-defined thing is far more reliable than one asked to handle anything. If your agent has grown into a general assistant, splitting it into a few specific agents with clear boundaries usually improves accuracy more than any prompt change.

Make the tools hard to misuse. Most agent errors we see in practice are tool-call errors: the right intention with the wrong argument. Tools with strict, well-described inputs that validate and reject bad calls turn a wrong action that nobody would notice into a clear error the agent can recover from.

Constrain the output. If the next step in your system needs structured data, require structured data and validate it before acting on it. Free text passed between steps is where small misunderstandings become large ones.

Plan the failure path. Every external call can fail. Every model call can return something unusable. Decide in advance what happens then: retry, fall back to a simpler path, or stop and ask a person. An agent with no defined failure path does the worst possible thing, which is to continue confidently.

Permissions matter more than capability

This is the part that gets least attention and causes the most damage. The interesting question about an agent is not what it can do, it is what it is allowed to do, and those should be very different lists.

The rule we use is simple: give an agent the narrowest permissions that let it do its actual job, and treat every expansion as a decision that needs a reason. If an agent only needs to read, it should not have write access. If it only needs one customer's records, it should not be able to query the whole table. If it can send an email to a user, someone should have decided explicitly that this is acceptable and bounded.

The reason this matters more for agents than for ordinary software is that an agent's behavior is not fully predictable in advance, and it can be influenced by the content it processes. An agent that reads untrusted input and also holds broad permissions is a combination worth taking seriously, because the input can shape what the agent decides to do next. Separating those two things, so that anything reading untrusted content has limited power and anything with real power reads only trusted input, removes a whole category of problem.

Then log everything. Every tool call, every input, every decision, kept in a form you can actually search after an incident. When an agent does something unexpected in production, the log is the only way to find out why.

Choose the simplest architecture that works

There is a lot of enthusiasm for multi-agent systems, and a lot of it is premature. Multiple agents mean multiple things to evaluate, multiple failure modes, and a coordination problem in addition. The complexity is real and immediate, while the benefit is often theoretical.

Start with one agent and good tools. Most problems that look like they need a team of agents are actually one agent that needs better tools or a narrower job. Add a second agent when you have a specific reason: genuinely separate domains of knowledge, a real need to isolate permissions, or a step that must be independently verified.

When you do split, keep the boundaries clear and the handoffs explicit and structured. The systems that work tend to look like ordinary software architecture, with clear responsibilities and defined interfaces, rather than a conversation between several models hoping to reach agreement.

The same restraint applies to frameworks. Frameworks are useful, and they also change quickly and make design choices for you that are hard to reverse later. Understand what your agent does without one before you adopt one, so the framework is a convenience rather than the only thing that makes the system work.

Decide what a human still approves

The most useful question in agent design is not "can the agent do this" but "what does it cost if the agent gets this wrong, and can we undo it".

Actions that are cheap and reversible are good candidates for full automation. Drafting something a person reviews, retrieving information, categorizing a request, preparing work for approval: if a mistake costs a few seconds, let the agent run.

Actions that are expensive or irreversible deserve a person in the loop, at least until you have evidence from your evaluation set that the agent handles them reliably. Moving money, sending something to a customer, deleting data, changing a production system. The point is not that agents cannot do these things. It is that reversibility, not confidence, should decide.

This is a design decision with real product consequences, and it is worth making explicitly rather than by default. In document extraction, the AI can pull structured data out of a page while users have no way to tell a confident extraction from an uncertain one, so everything needs re-checking manually and the automation delivers little benefit. Showing the model's confidence directly in the workflow changed that: high-confidence results moved through quickly, and uncertain ones got a quick check by a person. The lesson generalizes to agents. The value often comes not from removing the human but from making it obvious where the human is genuinely needed.

What it costs and how long it takes

We will not quote a figure here, because an honest one depends on the task, and anyone quoting a single number for "an AI agent" should not be trusted. What we can describe is where the cost actually goes, which is usually not where teams expect.

The prototype is the cheap part. Getting an agent working well enough to demo is often a matter of days, and that speed is exactly what creates the false expectation for everything after.

The expensive parts are evaluation, integration, and strengthening the system against failure. Building the test set and the scoring method takes real work. Connecting the agent properly to the systems it needs, with correct permissions and sensible error handling, takes real work. Handling the many rare, strange inputs takes real work, and it never fully ends, because new rare inputs keep appearing.

Then there is running cost, which behaves differently from ordinary software. An agent that takes several steps per request, each one a model call, has a per-request cost that grows with use rather than staying the same. That is worth modeling early with realistic volumes, because an economic model that works at demo scale can fail at production scale.

The practical advice: budget the prototype in weeks and the production system in months, and treat evaluation as a permanent line item rather than a phase that finishes.

Launch is the middle of the project, not the end

Ordinary software is fairly stable once it works. You release it, and unless someone changes it, it keeps doing the same thing. Agents do not behave that way, and this often surprises teams.

An agent depends on things that change. The model provider updates the model. An API it calls changes a response format. The kind of requests users send changes as they learn what the agent is good at, and as new users arrive with different habits. None of that changes your code, and all of it can change your results. An agent that was accurate in March can be noticeably worse in June with nothing in your repository having changed.

So the monitoring has to be different too. Uptime and error rates still matter, but they will not catch the failure that matters most, which is an agent that is running perfectly and answering badly. Watch quality directly: run your evaluation set on a schedule rather than only before a release, and track the mix of real requests so you can see when the inputs start to differ from what you tested.

Give users an easy way to say an answer was wrong, and actually read what comes back. User reports are the cheapest source of the failure cases you did not imagine, and each one belongs in the evaluation set. Watch cost per request alongside quality, because improving one often makes the other worse, and it is common to fix an accuracy problem by adding steps without noticing what that did to the economics.

The teams that do well here treat an agent like a system that needs ongoing maintenance rather than a feature that gets finished. That is a real ongoing commitment, and it is better to plan for it than to discover it.

How to start

Pick one task that is narrow, valuable, and forgiving. Narrow so you can define correctness. Valuable so the work is worth doing. Forgiving so early mistakes are cheap.

Write down what correct looks like before building. Collect real examples, including the awkward ones. Build the simplest version that could work, then run it against your examples and look honestly at the failures. Fix the structural causes rather than patching individual cases. Give it the minimum permissions it needs. Release it to a small group with a human check on anything expensive, and remove that check only when your evaluation gives you a reason to.

That is a slower path than the demo suggested. It is also the one that ends with something people can actually depend on.

Explore the guide

Common questions

What is the difference between an AI agent and an AI feature?

An AI feature runs a model inside a sequence you wrote. An agent decides the sequence itself: it is given a goal and a set of tools, and it chooses what to do next. The practical consequence is that a feature can be tested like normal software, while an agent needs evaluation, permission limits, and a defined failure path, because the same input can produce a different path on a different day.

Why do so many agent projects stop after the prototype?

Because a prototype is judged on whether an impressive thing happened, and production is judged on whether an unacceptable thing never happens. Those are close to opposite standards. Getting from one to the other means building an evaluation set, defining permissions, handling errors, and deciding what a human still approves. That work is unglamorous, it does not improve the demo, and teams routinely underestimate it.

How do you test an AI agent when the output is not deterministic?

You define what a correct run looks like in writing, then score against that definition rather than against an exact expected string. Collect real cases, including the messy ones, and run the agent against them repeatedly so you can tell whether a change helped or hurt. Every production failure should become a new case in that set, which is what stops the agent quietly regressing later.

Should we build a multi-agent system?

Usually not at the start. Multiple agents multiply what you have to evaluate and add a coordination problem on top. Most problems that look like they need several agents are one agent that needs better tools or a narrower job. Add a second agent when you have a specific reason, such as genuinely separate domains, isolated permissions, or a step that must be independently verified.

What permissions should an AI agent have?

The narrowest set that lets it do its actual job, with every expansion treated as a deliberate decision. If it only needs to read, it should not have write access. This matters more than for ordinary software because an agent's behavior is shaped partly by the content it processes, so an agent that reads untrusted input while holding broad permissions is a combination worth avoiding.

How much does it cost to put an AI agent into production?

The cost scales with the task, so a single number quoted for 'an AI agent' should be treated with suspicion. What is predictable is where the cost goes: the prototype is cheap and fast, while evaluation, integration, permissions, and handling the many rare, unusual inputs are where the real time goes. Running cost also scales with use, because a multi-step agent makes several model calls per request.

How do you start building an AI agent for production?

Pick one task that is narrow, valuable, and forgiving of early mistakes, then write down what a correct run looks like before building anything. Collect real examples, including the awkward ones, build the simplest version that could work, and run it against those examples to look honestly at the failures. Give it the minimum permissions it needs and release it to a small group with a human check on anything expensive.

Is it safe to let an AI agent act without a human reviewing it first?

The answer depends on what the action costs if the agent is wrong and whether that action can be undone. Cheap, reversible actions such as drafting a response or retrieving information are reasonable candidates for full automation early. Actions that move money, reach a customer, or delete data deserve a person in the loop until evaluation data shows the agent handles that specific action reliably, because reversibility, not confidence, should decide.

What is the biggest mistake teams make when moving an AI agent to production?

Treating the production system as a continuation of the demo rather than a different kind of work. A demo only has to succeed a few times with a friendly person watching, while production has to behave acceptably on inputs nobody chose, with a defined response for failure. Skipping the evaluation set is the most common specific error, because without one no change can be shown to have helped or hurt.