AI

Agentic AI framework design: how to architect a system that acts, not just answers

Editorial · Reveneau · July 28, 2026

Agentic AI framework design: how to architect a system that acts, not just answers

A few months ago a team walked us through an agent they had built to sort incoming sales leads: read the form submission, look up the company, decide a tier, and send it to the right sales rep. It worked well in the demo. Then we asked what happens when the lookup tool times out. Nobody had an answer, because the agent's plan was "call the tool, use the result," and nothing in the code described what to do when there was no result. That missing step is the whole subject of this post. Agentic AI framework design is not about picking a smarter model. It is about deciding, before you write the first line of code, how much autonomy the system gets, what it is allowed to touch, and what it does the moment something does not go as planned.

What "agentic" actually means

The word gets used loosely, so it is worth being precise. A system is agentic when it plans a sequence of steps and calls tools to carry them out, deciding what to do next based on what the previous step returned, without a human approving each individual step along the way. A chatbot takes one input and produces one output. An agent takes an input, decides on an action, observes what happened, and decides on the next action, some number of times, before it hands control back to a person.

That loop, plan, act, observe, replan, is the entire difference. It is also the entire source of new risk, because every one of those turns is a place the system can go wrong in a way a chatbot that answers once cannot. A chatbot that gives a bad answer wastes someone's time reading it. An agent that takes a wrong action three steps into a plan has already made a real change, and the fourth step will build on that mistake unless something catches it. We wrote about the difference between an answer and an action in more detail in can you trust an AI agent. This post is about the architecture decisions that determine whether the agent deserves that trust in the first place.

The first decision: single agent, or multiple

Most teams choose a multi-agent architecture before they need one, usually because the word "multi-agent" sounds more sophisticated than "one agent with a few tools." In our experience the opposite is true. A single agent with a small, well-scoped tool set handles the large majority of real workflows, and it is dramatically easier to reason about, test, and monitor than a system with several agents talking to each other.

The case for a single agent is simple: one plan, one context, one place to look when something breaks. You can trace exactly why it chose an action, because there is one decision-maker to trace. The case for multiple agents shows up when a single agent's job would require holding too much context or too many tools at once, to the point where it starts choosing the wrong tool because the list got too long, or losing track of an earlier step because the plan grew past what it can reliably track. At that point, splitting the work into a coordinator agent and a small number of narrow specialist agents, each with its own small tool set, is a real architectural improvement, not a fashion choice.

The mistake we see most is going straight to a group of equal agents with no coordinator. That multiplies the ways the system can fail without a matching gain in capability. If you are choosing multi-agent, choose a hierarchy: one agent decides what needs to happen and delegates, and specialists execute a narrow task and report back. That is a system you can debug. A flat network of agents negotiating with each other is not.

The decision that actually matters: autonomy per step, not per agent

Here is what's changed in how we think about this over the last couple of years of building these systems: autonomy is not a single setting you choose for the whole agent. It is a decision you make action by action. The question is never "how autonomous is this agent." It is "how autonomous is this agent allowed to be for this specific action."

Reading data can usually be broad. An agent that looks things up, summarizes, and drafts a response for a person to review carries low risk, because nothing has happened in the world yet. Writing data needs to be narrower, scoped to exactly the records and fields the task requires. Anything expensive or impossible to undo, sending an email, issuing a refund, deleting a record, deserves the least autonomy of all: a confirmation step, a narrower tool, or a person approving it before it runs.

This is why we design agents around a graduated scope rather than a single permission level. Start an agent on the lowest-risk version of its job, watch it, and measure how often it is right. Only widen what it can do once you have evidence it deserves the wider scope, one action at a time, not all at once. An agent that has proven itself at drafting a reply has told you nothing about whether it is safe to let it send that reply unsupervised. Those are different actions with different failure costs, and they should be evaluated separately.

Structuring tool access and permissions

Treat every tool you give an agent the way you would treat an API key: scope it to the smallest set of actions and data it needs, not the largest set it might eventually use. This gets skipped constantly, usually because it is faster to give the agent one broad "database" tool than to build three narrow ones.

The narrow set is worth the extra work. An agent with ten broad tools has many more ways to fail than one with three narrow tools, even when the narrow tools took longer to build, because each broad tool lets a bad decision reach more data. Two questions matter here, and they are not the same question: can the agent call this tool at all, and separately, can the agent call this tool on this particular piece of data, this account, this record. A support agent that can look up any customer's order history is a different risk than one that can only look up the order history of the customer whose ticket it is currently handling, even though both use the same underlying tool.

Logging belongs at every tool call, as well as at the agent's final output. Every tool call, its arguments, and its result should be recorded, because when something goes wrong three steps into a plan, the tool log tells you which step actually broke, rather than guessing from the final state of the system.

State and memory across steps

An agent's context window is not its memory. It holds a limited and incomplete record of what has happened, and long agent loops often summarize or drop earlier parts of it to make room for new tool results. If the agent's only record of what it already tried lives in that shrinking context window, it will eventually forget that it already attempted a step and failed, and try it again.

The fix is a persistent state record that sits outside the model's context: what the current plan is, which steps have completed, what each tool call returned, and what has already failed. This is plain engineering, not a model capability. It is also what makes an agent's behavior explainable after the fact. When someone asks why the agent did what it did, "here is the state at each step" is a real answer. "The model decided to" is not.

The tradeoff: more autonomy means more ways to fail, so control has to grow with it

The core tradeoff in agentic system design is this: more autonomy means more ways to fail. Every additional action you let an agent take without a check is an additional way the system can go wrong in production, in a way nobody reviewed in advance. That is not an argument against autonomy. It is an argument that observability, logging, evaluation, and human review need to grow at the same rate autonomy does, and keep up with it.

We have seen the failure mode that comes from ignoring this: a team grants a wide scope of action because the agent tested well on a narrow set of cases, then discovers in production that the real world produces inputs the test set never covered. The model did not get worse. The scope grew faster than the checks built to catch its mistakes.

When a simpler pipeline is the right choice

Not every workflow that touches an LLM should be an agent, and this is worth saying because "agentic" has become the default answer even when the actual job does not need it. If the sequence of steps and their order is known ahead of time and does not depend on a decision the model has to make while it runs, write it as a fixed pipeline: a sequence of function calls with an LLM doing one job inside it, like classifying an input or drafting a paragraph, and plain code handling the rest.

A pipeline is cheaper to run, faster to respond, and easier to test, because its behavior does not vary from run to run the way an agent's planning does. We build plenty of AI features this way: a document goes through extraction, then a model classifies it, then a fixed set of rules routes it, with no agent deciding what happens next because nothing here actually requires that decision to be made while it runs. Use an agent when the number and order of steps genuinely cannot be known in advance, not because the word sounds more advanced in a presentation.

Designing for the failure modes

Assume the agent will fail at some step, because it will, and design the failure path before you release rather than discovering it in production. Three failure modes show up constantly, and each needs a specific answer built in from the start.

A step fails. The tool times out, returns an error, or gives back something the agent did not expect. The agent needs an explicit path for this: retry with backoff for something transient, a different tool or approach for something structurally wrong, and a stop-and-ask-a-human path when neither resolves it. An agent with no defined failure path defaults to trying again, which is exactly how a single bad response turns into a stuck loop.

The agent loops. Put a fixed limit on the whole run: a maximum number of steps, a maximum time, or a maximum cost, after which it stops and reports instead of continuing indefinitely. Pair that limit with a check for repeated identical actions, since a looping agent is usually calling the same tool with the same arguments more than once.

A tool call has an unintended side effect. This is the one that actually costs money or trust, and the fix has to exist before the tool is ever called, not after. Scope what data the tool can touch, make destructive actions reversible where the underlying system allows it, and route anything irreversible through a confirmation step. No amount of model intelligence prevents this kind of failure. Only the limits you built around the tool do.

How we build these

When we design an agentic system for a client, we start with the job, not the architecture. What does this agent actually need to do, and what is the smallest system that can do it reliably. Most of the time that answer is a single agent with a narrow, well-permissioned tool set, staged autonomy that widens only after we can measure that it deserves that trust, and a state record and failure path built as part of the same build rather than added after something breaks. That is the same discipline behind our broader AI development work: prove the narrow case, measure it honestly, and only then widen the scope. The architecture diagram matters less than whether the system knows what to do the day a tool call fails, because it will, and that is the day the design either works or does not.

Thanks to the teams who have shown us their agents mid-build and asked the harder question before release instead of after. An agent earns autonomy the same way a person does: one action at a time, with someone watching, until the evidence says it does not need watching quite as closely anymore.

Related guide: Taking AI agents from prototype to production.

We also build them with generated code, which is worth stating because it is the same trust question applied to our own code. The agent's implementation is written by a model and read by a named engineer before it is used. We would not ask you to give an agent write access to your systems because of a good demo, and we do not merge our own code just because it compiles.

Common questions

What does "agentic" actually mean in AI development?

A system is agentic when it plans a multi-step sequence of actions and calls tools to carry them out, deciding what to do next based on what just happened, rather than answering a single prompt and stopping. The test is not whether it uses an LLM. The test is whether it can take more than one turn of action without a human approving each step.

When should we build a single agent with tools instead of a multi-agent system?

Start with a single agent with a small, well-defined tool set for almost every real task. A single agent is easier to reason about, debug, and monitor, and it covers the large majority of workflows we see. Move to multiple agents only when one agent's context or tool list would grow so large that it starts choosing the wrong tool or losing track of the plan, and even then, one coordinator plus a few narrow specialists works better than a group of equal agents with nobody in charge.

What is the biggest architectural decision in agentic AI framework design?

How much autonomy to grant at each step, not which model to use. Autonomy is not one setting for the whole agent. It is a decision made action by action: read access can be broad, write access should be narrow, and anything expensive to undo should require a human or a second check before it happens.

How should we structure tool access and permissions for an AI agent?

Give the agent the smallest tool set that can complete the task, scope each tool's permissions the way you would scope an API key, and treat "can call this tool" and "can call this tool on this data" as two separate questions. An agent with ten broad tools has many more ways to fail than one with three narrow ones, even if the narrow set takes more engineering to define.

How does state and memory management work across agent steps?

The agent needs a record of what it has already tried, what the tool calls returned, and what the current plan is, distinct from the model's own context window, which is finite and gets summarized or dropped. Without persistent state, an agent that retries a failed step cannot tell it already failed the same way twice, which is exactly how loops start.

What is the tradeoff between agent autonomy and control?

More autonomy means more ways to fail, so the controls and observability around the agent have to grow at the same rate that autonomy grows. An agent with broad, unsupervised write access needs far more logging, evaluation, and human review paths than one that only reads data and drafts a message for a person to send.

When is a non-agentic pipeline actually the better choice?

When the steps and their order are known in advance and do not depend on decisions the model has to make while it runs. If you can write the workflow as a fixed sequence of function calls with an LLM step doing the classification or drafting inside it, that is a pipeline, not an agent, and it is cheaper, faster, and far easier to test.

What happens when a step in an agent's plan fails?

The agent needs an explicit failure path for each tool call: retry with backoff for a transient error, a different approach for a tool that returned an unexpected result, or a stop-and-ask-a-human path when neither works. Without a defined failure path, the default behavior of most agent loops is to keep trying, which is how a single bad API response turns into a loop that never stops.

How do you stop an agent from looping indefinitely?

Put a fixed limit on the loop itself, as well as on individual tool calls: a maximum number of steps, a maximum time budget, or a maximum cost, after which the agent stops and reports rather than continuing. Pair that with a check for repeated identical actions, since an agent stuck in a loop is usually calling the same tool with the same arguments more than once.

How do you handle a tool call with an unintended side effect?

Design the tool's permission limits before the agent ever calls it, not after something goes wrong: scope what data it can touch, make destructive actions reversible where you can, and route anything irreversible through a confirmation step. The agent's intelligence does not prevent this kind of failure. Only the limits around the tool do.

Does a better model make an agentic system more reliable?

A better model can make the agent's reasoning better, but it does not build the safety checks, the state tracking, or the human review paths that make an agent safe to depend on. Most of the reliability work in an agentic system sits in the infrastructure around the model, not in the model itself.

How does Reveneau approach agentic AI framework design for clients?

We start by asking what the agent needs to do, not what an agent can do in general, then choose the smallest architecture (often a single agent with a narrow tool set) that fits, and build the state tracking, permission boundaries, and failure handling as part of the same build rather than as an afterthought once something breaks in production.