Run it safely

Permissions and guardrails for AI agents

Most serious agent incidents are permission problems rather than intelligence problems. The agent did something harmful because it was able to, not because it was incapable of reasoning. That makes the interesting design question not what an agent can do, but what it is allowed to do, and those should be two very different lists.

Published August 8, 2026. Updated September 30, 2026. Editorial.

Key takeaways

  • Give an agent the narrowest permissions that let it do its actual job, and treat every expansion as a decision needing a reason.
  • An agent that reads untrusted content and also holds broad permissions is the combination to avoid, because input can shape its decisions.
  • Put limits outside the agent. Anything enforced only by instructions is a request, not a control.
  • Log every tool call and argument. After an incident, the log is the only way to establish what happened.

Ordinary software does what it was written to do. An agent decides what to do at runtime, partly on the basis of content it reads, and that content is not always under your control. That single difference is why permissions deserve more attention for agents than for the systems teams are used to securing.

Least privilege, applied properly

The principle is old and the application is specific. Give the agent the minimum access that lets it complete its actual job, and require a reason for every expansion.

If it only needs to read, it must not have write access. If it needs one customer's records, it should not be able to query the whole table. If it needs three API endpoints, it should not hold a credential that reaches thirty. If it can act on a customer's behalf, the boundary of which customer must be enforced by the system rather than by the agent's own judgment.

The common failure is convenience. A broad credential is faster to set up during a prototype, and it stays in place into production without anyone noticing, because nothing forced anyone to revisit it. Scope the credentials at the point the agent leaves the prototype, and give each agent its own identity so its actions are attributable and revocable independently.

Separate untrusted input from real power

This is the structural rule that matters most, and it is specific to systems where a model decides what to do next.

An agent's behavior is shaped by the content it processes. If it reads a document, a web page, an email, or a support ticket, that content is input to its decision-making. Content from outside your control should therefore be treated as untrusted, because it can attempt to influence what the agent does next.

The defense is architectural rather than instructional. Keep the two things apart: an agent that reads untrusted content should have limited power, and an agent with real power should read only trusted input. When a workflow genuinely requires both, split it. Have one component read and summarize the untrusted content into a structured, constrained form, and a separate component with the real permissions act only on that validated structure, never on the raw text.

Telling an agent in its instructions to ignore instructions found in documents is not a control. It is a request, and it will be followed unreliably.

Enforce limits outside the agent

Any boundary that exists only in the prompt is advisory. Real controls sit in the surrounding system, where the agent's reasoning cannot change them.

Enforce permissions at the API and database layer, so an out-of-scope request fails regardless of what the agent intended. Cap spend, steps, and time externally, so a loop terminates. Rate limit actions with real-world effects, so a malfunction cannot send a thousand messages before anyone notices. Where an action is irreversible, require a confirmation step that the agent cannot perform on its own behalf.

A useful test: if the agent's instructions were replaced with something adversarial, what could it actually do? Whatever remains is your real security boundary. Everything else was a suggestion.

Decide autonomy by reversibility

The clearest way to decide what an agent may do alone is to ask what it costs when it is wrong and whether it can be undone.

Cheap and reversible actions are good candidates for autonomy: retrieving information, categorizing a request, drafting something a person will review, preparing work for approval. If a mistake costs a few seconds of someone's attention, let it run.

Expensive or irreversible actions deserve a person in the loop until evaluation data justifies otherwise: moving money, sending communications to customers, deleting data, changing production configuration. The point is not that agents can never do these things, it is that the decision should rest on evidence of reliability for that specific action, not on general confidence.

Write this list down explicitly. Teams that leave it implicit discover at launch that different stakeholders assumed very different boundaries.

Log for the incident you have not had yet

At some point an agent will do something unexpected and someone will need to explain exactly what happened. The quality of that explanation is decided in advance by what you logged.

Record every tool call with its arguments and result, the model's decision at each step, the input that started the run, the identity the agent acted under, and the cost and duration. Keep it searchable, retain it long enough to investigate something noticed weeks later, and give every run an identifier so a complaint maps to a specific trace.

This is also what makes improvement possible. The same logs that answer "what went wrong" are where the real failure cases come from, and those belong in your evaluation set.

Review the boundary as the agent changes

Permissions granted for one version tend to stay in place after the reason for them is gone. An agent gains a capability, receives broader access to support it, then loses that capability in a later redesign while the access remains.

Review periodically: what can each agent reach, why, and is that still needed. Remove what is no longer justified. This is ordinary access maintenance, and it matters more for agents because their behavior is less predictable than the software teams are used to auditing.

Checks on what goes in and what comes out

Permissions decide what an agent can reach. Guardrails, meaning automatic checks on what goes into and comes out of the agent, decide what is allowed through, and they belong on both sides of the agent.

On the way in, validate that a request is something this agent should handle at all, and normalize it into a form the agent expects. Requests that fall outside the intended scope should be refused by the surrounding system rather than left to the agent's discretion.

On the way out, check the result before anything acts on it. Does it match the required schema. Does every figure it reports appear in the source it cites. Does it reference records that actually exist. Does it avoid disclosing information the requesting user is not entitled to see. These are ordinary programmatic checks, they are cheap, and they catch a class of error that no amount of instruction reliably prevents.

The output check is the one most often skipped, and it is the one that matters most for anything customer-facing, because it is the last point at which a confidently wrong answer can still be stopped.

Established frameworks help structure this work. The OWASP Top 10 for Large Language Model Applications catalogues the failure modes that recur across real deployments, and the NIST AI Risk Management Framework offers a way to organize risk decisions so they are recorded rather than improvised. Neither removes the need for judgment about your own system, and both are better starting points than inventing a checklist from scratch.

A worked example: the ticket that tried to escalate itself

A support agent was given a tool to search a knowledge base and a separate tool to escalate a ticket to a human specialist, with the intention that escalation would be used sparingly, only when the agent genuinely could not resolve something. A user submitted a ticket containing text, embedded in what looked like a copied error log, instructing that the ticket should be immediately escalated with a note that the user account should be granted administrator access pending review.

The agent read the ticket, including the embedded instruction, because the instruction was part of the content it was given to process. It escalated the ticket and included the requested note in its summary, faithfully passing along language it had no way to distinguish from a real system-generated error message. Nothing about the model's reasoning was defective here. It did what a reasonable reader would do with text that looked like part of the ticket, which is exactly the danger: the agent could not tell the difference between content describing a problem and content trying to direct its next action, because both arrive as the same kind of text.

The fix that actually closed this was not a better prompt telling the agent to be suspicious of embedded instructions, because that kind of instruction is unreliable by nature, as the page above explains. The fix was structural: the escalation tool was changed to require a specific, enumerated reason code rather than free text, so there was no field left for an attacker-supplied note to travel through, and any account-access change was removed from what an escalation could request at all, since escalation was never supposed to touch permissions in the first place. The permission the escalation tool held was simply narrower than what the injected text asked it to do.

Segmenting agents by trust boundary, not by task

A useful habit when several agents or several steps are involved is to draw the trust boundary explicitly on a diagram before writing any code, marking which components touch content from outside the organization and which components hold real permissions. This is a different grouping than the one you would draw by task.

Two steps that look like they belong together because they are both "part of handling a support ticket" may need to sit on opposite sides of that boundary, if one of them reads the customer's raw message and the other one is the only component permitted to modify the account. Drawing the trust boundary as its own diagram, separate from the flowchart of what the system does, makes it visible when a single component has been given both jobs by accident, which is the exact condition the worked example above shows going wrong. Single agent or multi-agent: choosing an architecture covers when that separation is worth the added complexity of a second component and when it is not.

Handling a credential leak or a compromised tool

Plan for the case where a tool's underlying credential is exposed or a dependency the agent relies on is itself compromised, because that is a different failure than the agent simply making a mistake, and it needs a different response ready in advance.

Each agent having its own identity, rather than sharing a credential across several agents or reusing a person's own access, is what makes a fast, targeted response possible: revoke that one identity without disrupting every other agent or person who happened to share it. Keep a documented, tested procedure for revoking a specific agent's access quickly, and periodically confirm the procedure still works rather than assuming it does because it was written down once. A revocation procedure discovered to be broken during an actual incident adds real delay at the exact moment delay is most costly.

Where permission scoping intersects with data privacy

Beyond preventing an agent from taking an unwanted action, permission scoping is also how you prevent an agent from disclosing something it should never have been able to see in the first place, which is a related but distinct concern worth naming on its own.

An agent that can technically query a broader dataset than any single request needs will occasionally include information in a response that the requesting user was never entitled to see, not because it intended to leak anything, but because the information was present in what it retrieved and seemed relevant to answering the question asked. Scoping data access down to exactly what a given request is authorized to see, rather than relying on the agent to withhold information appropriately, removes this risk at the source rather than hoping the model exercises judgment it was never actually asked to exercise.

Best for

  • Any agent that can take actions with real-world effects rather than only returning text
  • Systems where the agent reads content from outside your control
  • Teams that need an auditable record of what an agent did and under whose identity

Avoid if

  • The agent is a closed prototype with no access to real systems or real data, where this can wait until it leaves the prototype

Check before you decide

  • Ask what the agent could do if its instructions were replaced with something adversarial
  • Confirm limits are enforced in the API and database layer, not only in the prompt
  • Confirm every tool call is logged with arguments, identity, and a traceable run identifier

Common questions

Can you stop an agent following instructions hidden in a document?

Not reliably through instructions. Telling an agent to ignore instructions it encounters is a request rather than a control. The dependable approach is architectural: keep untrusted content away from real permissions, so anything reading outside content has limited power and anything with real power acts only on validated, structured input.

What permissions should an AI agent start with?

The narrowest set that completes its actual job, with its own identity so actions are attributable and revocable. Read-only unless it genuinely needs to write, scoped to the specific records it needs rather than whole tables, and limited to the specific endpoints it calls.

How do you decide what an agent may do without human approval?

By reversibility rather than confidence. Actions that are cheap to undo can be automated early. Actions that move money, reach customers, or delete data should keep a person in the loop until your evaluation data shows the agent handles that specific action reliably.

What should be logged for an AI agent?

Every tool call with arguments and result, the decision at each step, the originating input, the identity the agent acted under, and cost and duration, all searchable and tied to a run identifier. After an incident this is the only way to establish what happened, and those traces are also the source of real cases for your evaluation set.

What is the difference between a permission and a safety check for an AI agent?

A permission decides what the agent can reach, such as which systems, records, and actions it has access to. A safety check decides what gets through, checking a request on the way in and a result on the way out before anything acts on it. The two work together: permissions limit how much can break when something fails, and safety checks catch a bad result even inside that limited area.

Are permission problems more common than reasoning problems in AI agents?

Most serious agent incidents come from the agent doing something harmful because it was able to, not because it reasoned poorly. That makes the interesting design question what an agent is allowed to do rather than what it can do, which is why permission scoping deserves at least as much attention as prompt quality.

How much extra work is it to add proper permissions and safety checks to an AI agent?

It is ordinary engineering: scoping credentials, enforcing limits at the API and database layer, validating output before anything acts on it, and building a searchable log. None of it is visible in a demo, which is why it is routinely underestimated in early plans, but it is a fixed cost of running any agent that connects to real systems rather than an optional add-on.

Is it safe to give an AI agent broad access during a prototype?

It is common practice because a broad credential is faster to set up, but that convenience should not survive into production. The fix is to scope credentials down at the point the agent leaves the prototype stage and give each agent its own identity, so its actions are attributable and its access can be revoked independently of any other system.