Taking AI agents from prototype to production / Start here
What an AI agent actually is, and when you need one
An AI agent is a system where the model decides what to do next, not just what to say next. It is given a goal and a set of tools, and it works through steps until it decides it is finished. That single property, deciding rather than following, is what makes agents powerful and what makes them hard to run in production. Plenty of products called agents are really a model call inside a fixed sequence, which is often the better choice.
Published August 8, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- An agent chooses its own steps at runtime. An AI feature runs a model inside a sequence you wrote in advance.
- The tradeoff is control: agents handle situations you cannot script, but you give up the ability to test a fixed path.
- If you can write down the steps reliably, write down the steps. A predictable workflow is cheaper, faster, and easier to trust.
- Reach for an agent when the right sequence genuinely cannot be known before the request arrives.
The word "agent" is used for many different things in the market right now, and that loose use causes real problems. Teams commit to building an agent before deciding whether the problem needs one, and they take on all of the operational difficulty of agents for a task a simple workflow would have handled better.
Here is the definition worth using. An agent is a system where a model decides what to do next. You give it a goal, a set of tools it can call, and some context. It reasons about the goal, picks a tool, looks at the result, and decides what to do after that. It keeps going until it concludes it is done or hits a limit you set.
Compare that to what most AI features actually are. A summarizer takes a document and returns a summary. A classifier takes a message and returns a category. There is a model involved, and it may be doing something genuinely difficult, but the sequence is fixed. You wrote it. The model fills in one step of a path that was decided in advance.
Why the distinction has practical consequences
This is not a vocabulary argument. The two designs behave differently in every way that matters once real users are involved.
Testing. A fixed workflow can be tested the way software is normally tested. Given this input, the system does these steps and produces this output. An agent may take a different path on Tuesday than it took on Monday for a similar request, so testing means evaluating the quality of outcomes across many cases rather than asserting one exact result.
Debugging. When a workflow misbehaves, the mistake is somewhere in code you wrote, and you can find it. When an agent misbehaves, you have to reconstruct why it chose the path it chose, which requires having logged every step and every tool call. Without that record, you are guessing.
Cost. A workflow with one model call costs one model call. An agent that reasons through six steps makes at least six calls, plus whatever the tools cost. The difference is small in a demo and significant at production volume.
Failure modes. A broken workflow usually fails visibly. A confused agent often fails invisibly: it produces something plausible, in the right format, that happens to be wrong. That is a harder failure to catch and a more dangerous one to release.
When a fixed workflow is the right answer
Choose the predictable path when you can write down the steps. If the process is "take the document, pull out these six fields, validate them, write them to the database, notify the owner", you do not need an agent. You need those five steps and a model call for the extraction. That system will be faster, cheaper, easier to test, and far easier to explain to whoever has to trust it.
This is the case in more situations than the current enthusiasm suggests. A great deal of useful AI work is extraction, classification, summarization, drafting, and routing, all of which fit comfortably into sequences you control. Building those as agents adds unpredictability with no matching benefit.
A good test: try to write the steps out as a flowchart. If you can, and the flowchart does not have dozens of branches, build the flowchart. The moment you find yourself writing "and then it depends on what came back, in ways I cannot fully enumerate", that is the signal that an agent may genuinely be warranted.
When an agent is worth its complexity
Agents make sense when the sequence cannot be known in advance. A support request might need a database lookup, or an API call, or a document search, or all three in an order that depends on what the earlier steps returned. A research task might require following wherever the evidence leads. A debugging assistant cannot know which file matters until it has read some of them.
What these cases share is genuine branching that depends on intermediate results, where enumerating every path in advance would be impractical. In those situations, letting the model decide the sequence is not a shortcut. It is the only reasonable approach, and the operational work described across this guide is the cost of it.
Agents also make sense when the tool set is broad and the right combination varies per request. If there are twenty tools and any given task might need three of them in an order you cannot predict, a model choosing among them is a reasonable design.
The mixed design most production systems end up with
In practice, the systems that work well are usually neither pure workflow nor pure agent. They are a workflow with agentic steps inside it.
The overall shape is fixed, because you know the general process. Within that shape, one or two steps are genuinely open-ended, and those steps get an agent with a narrow job and a small set of tools. The result keeps most of the predictability of a workflow while allowing flexibility exactly where the problem requires it.
This is worth aiming for deliberately, because it gives you the benefits of both. The fixed parts can be tested normally. The agentic parts are small enough to evaluate carefully and constrain tightly. When something goes wrong, you only need to search one limited step rather than an entire open-ended process.
Deciding for your own case
Ask three questions before committing.
First, can you write down the steps? If yes, write them down and build that. Do not add an agent because agents are interesting.
Second, what does a wrong answer cost? If the answer is "little, and a person sees it before anything happens", you can safely experiment. If a wrong answer moves money or reaches a customer directly, the standard for letting it act alone is much higher and you should expect to keep a person in the loop for a while.
Third, can you tell whether an answer was right? If you cannot define correctness for this task, you cannot evaluate an agent doing it, which means you cannot safely improve it or know when it degrades. Solve that before building. How to evaluate an AI agent before you trust it covers what that written definition needs to contain and how to score against it once you have one.
Teams that work through those three questions honestly usually end up building something smaller than they first imagined, and releasing it much sooner.
A worked example: a returns request
Consider a retailer handling a customer message that says a product arrived damaged and asks for a refund. Walk through what a fixed workflow would do and where an agent would differ.
A workflow reads the message, extracts the order number and the stated reason, checks the order against a return policy table, and either approves a refund automatically or routes the case to a person, depending on the amount and the policy rule that matched. Every one of those steps is known in advance. The model is doing one job inside it, reading unstructured text and turning it into a structured reason code, and everything after that is ordinary logic. This is fast to build, cheap to run, and easy to test: given a message describing a damaged product under fifty dollars, the system should approve automatically, and you can assert that directly.
Now suppose the message is more tangled. The customer says the product arrived damaged, mentions in passing that a second item from the same order never arrived, and asks whether they can exchange one and keep credit on the other. There is no fixed order to these sub-problems. The system has to notice that two separate issues were raised, decide whether to look up shipping status for the missing item, decide whether the damaged item needs a photo before a refund can proceed under policy, and decide how to phrase a single reply that resolves both without contradicting itself. That is a case where the right sequence of lookups depends on what the message actually contains, which is the exact signal described above for reaching for an agent rather than a workflow.
The lesson from this pair of examples generalizes. Most support and operations messages look like the first case and should be built as workflows. A minority look like the second case, where the branching is genuine rather than something you failed to enumerate. Sorting incoming requests into these two buckets before building anything, rather than assuming every message needs the same kind of system, is itself one of the more valuable early design decisions a team can make.
An edge case worth planning for: the malformed request
Whichever design you choose, plan for input that fits neither your workflow's assumptions nor an agent's available tools. A message in an unexpected language, a request that names a product the company does not sell, or a message that is actually spam routed into the same queue by mistake.
A fixed workflow handles this by falling through to a default branch, usually "send to a person", and that default is easy to write because the workflow's steps are already enumerated. An agent has to be given an explicit instruction that some inputs are out of scope and that recognizing this is itself a correct outcome, because otherwise it will attempt to force an unrelated request into the tools it has, which produces a confident, wrong response rather than a clean refusal. Why AI agent pilots stall before production covers why teams that skip planning for this case are often surprised by how many real requests fall outside what they designed for.
Best for
- Tasks where the right sequence of steps depends on intermediate results and cannot be enumerated in advance
- Problems with a broad tool set where the useful combination changes per request
- Work where a human reviews the output before anything irreversible happens
Avoid if
- You can draw the process as a flowchart without dozens of unpredictable branches
- The task is a single well-defined operation such as extraction, classification, or summarization
- You cannot yet define what a correct result looks like for the task
Check before you decide
- Write the steps out first and confirm they genuinely cannot be enumerated
- Confirm you can define and score correctness before building anything
- Check what a wrong answer costs and whether it can be undone
Common questions
Is a chatbot an AI agent?
Only if it decides what to do next rather than only what to say next. A chatbot that answers from a model or retrieves documents through a fixed retrieval step is an AI feature. It becomes an agent when it can choose among tools and take actions whose order was not decided in advance.
Is it wrong to build a fixed workflow instead of an agent?
No, and it is often the better engineering decision. A workflow you can draw as a flowchart is cheaper to run, easier to test, and far easier to debug. Use an agent when the sequence genuinely cannot be known before the request arrives, not because the label is more appealing.
Can we start with a workflow and move to an agent later?
Yes, and that is usually the sensible order. Building the workflow first forces you to understand the task, define correctness, and connect the tools properly. All of that work carries over directly if you later decide one step needs to be open-ended.
What tools does an AI agent need to be given?
The specific functions it needs to complete its goal and nothing more, such as a database lookup, an API call, or a document search. Each tool should have a precise description of what it does and when to use it, because the model reasons from that description when deciding what to call next. A broad or poorly described tool set is a common source of agent errors that looks like a reasoning failure but is really a design failure.
How do you know if a task genuinely needs an agent instead of a workflow?
Try to write the task out as a flowchart first. If you can draw it without dozens of unpredictable branches, build that flowchart with a model call inside it rather than an agent. The signal that an agent is genuinely warranted is finding yourself writing that the next step depends on what came back in ways you cannot fully enumerate in advance.
Does an AI agent cost more to run than a fixed AI workflow?
Usually, because an agent that reasons through several steps makes several model calls per request, plus whatever the tools cost, while a fixed workflow with one model call costs one model call. The difference is small in a demo and becomes significant at production volume, which is one reason to default to a workflow when the steps can be written down in advance.
What happens if you build an agent for a task that only needed a workflow?
The system becomes harder to test, harder to debug, and more expensive to run for no matching benefit. A workflow can be tested by asserting that a given input produces a given output, while an agent has to be evaluated across many cases because its path can vary. Unpredictability without a reason to accept it only adds cost and risk.
How do you debug an AI agent compared to a normal application?
In a normal application, a mistake is somewhere in code you wrote, and you can find it directly. In an AI agent, you have to reconstruct why it chose the path it chose, which requires having logged every step and every tool call it made along the way. Without that record, debugging an agent becomes guesswork rather than a search through known code.
Related reading
What founders get wrong about AI agents
An impressive agent demo and a reliable agent are two different things. Most of the work, and most of the risk, is in the final step before production, which nobody shows in the demo.
Can you trust an AI agent with real work yet?
An agent that answers a question and an agent that takes an action are not the same risk. Here is how we decide where an agent is ready to act, and where it is not.