Can you trust an AI agent with real work yet?

Last year a team showed us an agent they were proud of. You typed a request in plain language, and it went off and did the work: it read the customer's account, decided what needed to change, and made the change. In the demo it was a pleasure to watch. Then someone asked the question nobody had asked. What happens the day it changes the wrong account? Nobody spoke, because nobody had an answer. That moment shows the whole problem with agents. An agent that answers a question and an agent that takes an action are not the same risk, and most of the excitement is about the first while most of the danger comes from the second.
We work on a lot of AI, and we now use a simple way to decide whether an agent can be trusted with a real task. It is not one question about the agent. It is four questions about the task: can you measure whether it did the task right, what does it do when it is unsure, how hard is the action to undo, and who is accountable when it goes wrong. An agent can pass all four for one task and fail all four for a similar task. So the useful question is never "are agents ready." It is "ready for what."
An action is a different risk than an answer
A chatbot that gives a wrong answer costs the user a minute. They read it, they see it is wrong, they ignore it. An agent that takes a wrong action costs something real. It sends the email to the wrong person, it approves the refund that should have been denied, it moves the record that other systems depend on. The model underneath can be the same. The consequence is not.
This is why "the agent is smart enough now" is the wrong place to start. Smart is not the risk. Acting is the risk. When we look at an agent, the first thing we do is list every real action it can actually take, not every question it can answer. Reading data is one thing. Writing data is another. Sending something to a customer is another again. The trust you need grows with the harm the action can cause, and that has almost nothing to do with how clever the model sounds in a demo.
Here is the pattern, with the identifying details removed. A team built an agent to handle a common support request end to end, from reading the ticket to issuing the fix. In testing it was right most of the time, and everyone felt good. Then it received a ticket that looked routine but belonged to an unusual account, and it confidently made a change that took a person two days to undo. The idea was fine. The scope was wrong. They had given the agent the power to act before they had the evidence to trust it with that action.
The four questions, one at a time
Take the four in order, because they build on each other.
Can you measure it. With a normal program, you look at the output and you know if it is correct. With an agent, "correct" is often a judgment, and the action may leave no clear signal of whether it was right. If you cannot measure whether the agent acted correctly, on examples you trust, you are working without any evidence. You will not notice it getting worse until a user is hurt. So measurement comes first, because without it the other three questions cannot even be checked.
What does it do when unsure. A good agent knows the limits of its own knowledge and stops there. A dangerous one acts on every input, including the ones it does not understand. The single most useful behavior we build into an agent is the ability to say "I am not sure about this one" and hand it to a person. An agent that asks for help on the hard five percent is far safer than one that acts confidently on all of it, and it is usually not much slower.
How reversible is the action. This is the one teams give the least weight to. Drafting a reply that a human approves is safe, because nothing has happened yet and a person is the last step. Sending a payment is not safe, because there is often no clean way to take it back. The more reversible the action, the more freedom the agent can have. The harder it is to undo, the more a person should approve it first. This matches how the strongest engineering teams already work: the industry's long-running research on high-performing teams, DORA's State of DevOps program, keeps finding that the best teams release small changes and recover from failure fast. An agent should follow the same rule. Small, reversible actions first, and easy recovery when something goes wrong.
Who is accountable. When the agent makes a mistake, "the AI did it" is not an answer a customer or a regulator will accept. Someone has to own each action the agent can take. Before it is put into use, we decide who is responsible, and we make sure that person can see what the agent is doing and stop it. An agent with no clear owner is not an efficiency. It is a risk with a nice interface.
Most of the trust comes from the system around the model
There is a comfortable belief that a better model will make the agent trustworthy. It will not, and we have the evidence in our own history as an industry. When Google engineers studied real machine learning systems, they found that the model itself is only a small fraction of the code, and the vast majority is the infrastructure around it: the data pipelines, the serving, the monitoring, the configuration. Agents make this even more true, because now you are adding code around the model that decides what it is allowed to do, checks what it did, and catches its mistakes.
So when we make an agent trustworthy, most of the work is not the model. It is the limits on what actions it can take. It is the measurement that tells us if it acted correctly. It is the monitoring that shows what it is doing right now. It is the human review path for the cases it should not handle alone. A stronger model can make the agent smarter. It cannot add the safety checks that nobody built.
This is also why agent projects fail for the same reason other AI projects fail. A RAND study of AI project failures found that more than 80 percent of them do not succeed, about twice the rate of other technology projects, and the leading cause was not weak technology. It was misunderstanding the real problem. With agents the misunderstanding is usually about scope. The team gives the agent too much power too soon, skips the measurement and the safety checks, and is surprised when it acts wrongly on a case nobody planned for. The model was never the thing that failed.
How to start without risking the business
You do not build trust in an agent by arguing about it. You build it by watching it act, in a place where being wrong is cheap. So start small and start reversible. Pick a low-risk action that is easy to undo, have a person approve each action at first, and measure how often the agent is right before you widen what it can do. Each time it proves itself on a narrow task, you can extend its scope a little, with the four questions checked again for the new action.
We think of this as the same disciplined approach as any AI product build. It is the core of our AI development work, and it is part of the broader way we approach any custom software build: prove the idea, then do the unglamorous work that makes it safe to depend on. If you want the longer version of why the exciting demo is only the beginning, we wrote about the gap between an AI demo and a real product. Both of those lessons matter most for agents, because with an agent the failure is not a bad answer on a screen. It is a real action.
The honest summary is that agents are ready for some work and not for other work, and the task decides which is which. An agent that reads and organizes, that drafts what a person approves, that takes small steps you can undo, with a clear owner watching, can be trusted today. An agent that moves money on its own, with no measure of whether it was right and no one accountable, cannot, no matter how good the demo looked.
Thanks to the engineers who have worked with us and asked the hard question early, before an agent was put into use instead of after. Trust an agent with the actions you can measure, undo, and take responsibility for, and not one action more than that.
Related guide: Taking AI agents from prototype to production.
The same reasoning governs how we work internally, which is a useful test of whether we believe our own advice. Every line of code we deliver is generated, and none of it reaches a client branch without a named engineer reading it first. We do not extend more trust to a model writing our code than we would tell you to extend to an agent handling your customers. The rule is the same in both cases: the model does the work, a person is accountable for the result, and the trust is built step by step with evidence rather than granted because the output looked convincing.
Sources
- RAND (2024), "The Root Causes of Failure for AI Projects". Supports the figure that more than 80 percent of AI projects fail, about twice the rate of non-AI technology projects, with misunderstanding the real problem named as the leading cause rather than the technology.
- Sculley et al. (2015), "Hidden Technical Debt in Machine Learning Systems," NeurIPS. Supports the point that the model is only a small fraction of a real machine learning system, and the surrounding infrastructure (safety checks, measurement, monitoring, review) is most of the work.
- DORA, State of DevOps research program. Supports the point that the strongest engineering teams release small changes often and recover from failure fast, which matches giving an agent small, reversible scope first.


