Agentic AI framework design: how to architect a system that acts, not just answers

A few months ago a team walked us through an agent they had built to sort incoming sales leads: read the form submission, look up the company, decide a tier, and send it to the right sales rep. It worked well in the demo. Then we asked what happens when the lookup tool times out. Nobody had an answer, because the agent's plan was "call the tool, use the result," and nothing in the code described what to do when there was no result. That missing step is the whole subject of this post. Agentic AI framework design is not about picking a smarter model. It is about deciding, before you write the first line of code, how much autonomy the system gets, what it is allowed to touch, and what it does the moment something does not go as planned.
What "agentic" actually means
The word gets used loosely, so it is worth being precise. A system is agentic when it plans a sequence of steps and calls tools to carry them out, deciding what to do next based on what the previous step returned, without a human approving each individual step along the way. A chatbot takes one input and produces one output. An agent takes an input, decides on an action, observes what happened, and decides on the next action, some number of times, before it hands control back to a person.
That loop, plan, act, observe, replan, is the entire difference. It is also the entire source of new risk, because every one of those turns is a place the system can go wrong in a way a chatbot that answers once cannot. A chatbot that gives a bad answer wastes someone's time reading it. An agent that takes a wrong action three steps into a plan has already made a real change, and the fourth step will build on that mistake unless something catches it. We wrote about the difference between an answer and an action in more detail in can you trust an AI agent. This post is about the architecture decisions that determine whether the agent deserves that trust in the first place.
The first decision: single agent, or multiple
Most teams choose a multi-agent architecture before they need one, usually because the word "multi-agent" sounds more sophisticated than "one agent with a few tools." In our experience the opposite is true. A single agent with a small, well-scoped tool set handles the large majority of real workflows, and it is dramatically easier to reason about, test, and monitor than a system with several agents talking to each other.
The case for a single agent is simple: one plan, one context, one place to look when something breaks. You can trace exactly why it chose an action, because there is one decision-maker to trace. The case for multiple agents shows up when a single agent's job would require holding too much context or too many tools at once, to the point where it starts choosing the wrong tool because the list got too long, or losing track of an earlier step because the plan grew past what it can reliably track. At that point, splitting the work into a coordinator agent and a small number of narrow specialist agents, each with its own small tool set, is a real architectural improvement, not a fashion choice.
The mistake we see most is going straight to a group of equal agents with no coordinator. That multiplies the ways the system can fail without a matching gain in capability. If you are choosing multi-agent, choose a hierarchy: one agent decides what needs to happen and delegates, and specialists execute a narrow task and report back. That is a system you can debug. A flat network of agents negotiating with each other is not.
The decision that actually matters: autonomy per step, not per agent
Here is what's changed in how we think about this over the last couple of years of building these systems: autonomy is not a single setting you choose for the whole agent. It is a decision you make action by action. The question is never "how autonomous is this agent." It is "how autonomous is this agent allowed to be for this specific action."
Reading data can usually be broad. An agent that looks things up, summarizes, and drafts a response for a person to review carries low risk, because nothing has happened in the world yet. Writing data needs to be narrower, scoped to exactly the records and fields the task requires. Anything expensive or impossible to undo, sending an email, issuing a refund, deleting a record, deserves the least autonomy of all: a confirmation step, a narrower tool, or a person approving it before it runs.
This is why we design agents around a graduated scope rather than a single permission level. Start an agent on the lowest-risk version of its job, watch it, and measure how often it is right. Only widen what it can do once you have evidence it deserves the wider scope, one action at a time, not all at once. An agent that has proven itself at drafting a reply has told you nothing about whether it is safe to let it send that reply unsupervised. Those are different actions with different failure costs, and they should be evaluated separately.
Structuring tool access and permissions
Treat every tool you give an agent the way you would treat an API key: scope it to the smallest set of actions and data it needs, not the largest set it might eventually use. This gets skipped constantly, usually because it is faster to give the agent one broad "database" tool than to build three narrow ones.
The narrow set is worth the extra work. An agent with ten broad tools has many more ways to fail than one with three narrow tools, even when the narrow tools took longer to build, because each broad tool lets a bad decision reach more data. Two questions matter here, and they are not the same question: can the agent call this tool at all, and separately, can the agent call this tool on this particular piece of data, this account, this record. A support agent that can look up any customer's order history is a different risk than one that can only look up the order history of the customer whose ticket it is currently handling, even though both use the same underlying tool.
Logging belongs at every tool call, as well as at the agent's final output. Every tool call, its arguments, and its result should be recorded, because when something goes wrong three steps into a plan, the tool log tells you which step actually broke, rather than guessing from the final state of the system.
State and memory across steps
An agent's context window is not its memory. It holds a limited and incomplete record of what has happened, and long agent loops often summarize or drop earlier parts of it to make room for new tool results. If the agent's only record of what it already tried lives in that shrinking context window, it will eventually forget that it already attempted a step and failed, and try it again.
The fix is a persistent state record that sits outside the model's context: what the current plan is, which steps have completed, what each tool call returned, and what has already failed. This is plain engineering, not a model capability. It is also what makes an agent's behavior explainable after the fact. When someone asks why the agent did what it did, "here is the state at each step" is a real answer. "The model decided to" is not.
The tradeoff: more autonomy means more ways to fail, so control has to grow with it
The core tradeoff in agentic system design is this: more autonomy means more ways to fail. Every additional action you let an agent take without a check is an additional way the system can go wrong in production, in a way nobody reviewed in advance. That is not an argument against autonomy. It is an argument that observability, logging, evaluation, and human review need to grow at the same rate autonomy does, and keep up with it.
We have seen the failure mode that comes from ignoring this: a team grants a wide scope of action because the agent tested well on a narrow set of cases, then discovers in production that the real world produces inputs the test set never covered. The model did not get worse. The scope grew faster than the checks built to catch its mistakes.
When a simpler pipeline is the right choice
Not every workflow that touches an LLM should be an agent, and this is worth saying because "agentic" has become the default answer even when the actual job does not need it. If the sequence of steps and their order is known ahead of time and does not depend on a decision the model has to make while it runs, write it as a fixed pipeline: a sequence of function calls with an LLM doing one job inside it, like classifying an input or drafting a paragraph, and plain code handling the rest.
A pipeline is cheaper to run, faster to respond, and easier to test, because its behavior does not vary from run to run the way an agent's planning does. We build plenty of AI features this way: a document goes through extraction, then a model classifies it, then a fixed set of rules routes it, with no agent deciding what happens next because nothing here actually requires that decision to be made while it runs. Use an agent when the number and order of steps genuinely cannot be known in advance, not because the word sounds more advanced in a presentation.
Designing for the failure modes
Assume the agent will fail at some step, because it will, and design the failure path before you release rather than discovering it in production. Three failure modes show up constantly, and each needs a specific answer built in from the start.
A step fails. The tool times out, returns an error, or gives back something the agent did not expect. The agent needs an explicit path for this: retry with backoff for something transient, a different tool or approach for something structurally wrong, and a stop-and-ask-a-human path when neither resolves it. An agent with no defined failure path defaults to trying again, which is exactly how a single bad response turns into a stuck loop.
The agent loops. Put a fixed limit on the whole run: a maximum number of steps, a maximum time, or a maximum cost, after which it stops and reports instead of continuing indefinitely. Pair that limit with a check for repeated identical actions, since a looping agent is usually calling the same tool with the same arguments more than once.
A tool call has an unintended side effect. This is the one that actually costs money or trust, and the fix has to exist before the tool is ever called, not after. Scope what data the tool can touch, make destructive actions reversible where the underlying system allows it, and route anything irreversible through a confirmation step. No amount of model intelligence prevents this kind of failure. Only the limits you built around the tool do.
How we build these
When we design an agentic system for a client, we start with the job, not the architecture. What does this agent actually need to do, and what is the smallest system that can do it reliably. Most of the time that answer is a single agent with a narrow, well-permissioned tool set, staged autonomy that widens only after we can measure that it deserves that trust, and a state record and failure path built as part of the same build rather than added after something breaks. That is the same discipline behind our broader AI development work: prove the narrow case, measure it honestly, and only then widen the scope. The architecture diagram matters less than whether the system knows what to do the day a tool call fails, because it will, and that is the day the design either works or does not.
Thanks to the teams who have shown us their agents mid-build and asked the harder question before release instead of after. An agent earns autonomy the same way a person does: one action at a time, with someone watching, until the evidence says it does not need watching quite as closely anymore.
Related guide: Taking AI agents from prototype to production.
We also build them with generated code, which is worth stating because it is the same trust question applied to our own code. The agent's implementation is written by a model and read by a named engineer before it is used. We would not ask you to give an agent write access to your systems because of a good demo, and we do not merge our own code just because it compiles.


