What founders get wrong about AI agents

We have built a lot of AI features over the past few years, and AI agents are the thing founders ask about most right now. The conversation almost always starts the same way. Someone saw a demo, an agent that booked a meeting or closed a support ticket entirely on its own, and it looked almost impossible. They want that. The problem is not that the excitement is wrong. The demo really was impressive. The problem is that the demo and a product you can actually release are two different things, and the work between them is where most of the work, most of the cost, and almost all of the risk is.
Here is how we think about agents, and where we see founders get it wrong.
The demo is the easy part
A good agent demo is designed to go well. Someone picked the task, picked the input, and ran it on a clean example where the data was there and the request was clear. In that example, a modern model looks excellent. It reads the request, calls the right tool, takes the action, and reports back. It feels finished.
It is not finished. It is the start. A demo shows you the 80 percent of cases that go right. A product has to handle the other cases: the request phrased in a way nobody expected, the field that is missing, the tool that returns something strange, the user who asks for two things at once. On those inputs, the model does not stop and say "I am not sure." It acts on its best guess, and its best guess is sometimes wrong. In a chatbot, a wrong guess is an awkward answer. In an agent, a wrong guess is a wrong action: the wrong meeting booked, the wrong refund issued, the wrong record deleted. That is a real cost, and it is the reason agents are harder to release than chat features.
We wrote about this gap in more detail in why the final step from demo to production is the real job. The short version: the impressive part is cheap now, and the reliable part is the whole engineering problem.
A narrow agent works better than a general one, by a large margin
The most common wrong instinct we see is wanting the agent to do everything. A founder imagines one smart assistant that handles support, sales, scheduling, and internal questions, all in one. It sounds efficient. It is a mistake.
The wider you make an agent's job, the harder it is to know whether it is working. If an agent can take a hundred different actions, you cannot test all the ways those actions combine, and you cannot define what "correct" even means across all of them. A narrow agent has a small, clear job. Answer this one type of question. Sort these requests into three groups. Draft the first reply to this kind of ticket. When the job is narrow, you can write down what correct looks like, build tests for it, measure how often it succeeds, and catch it when it fails. That is what matters. You cannot trust what you cannot measure, and you cannot measure a job that has no clear limits.
So when someone asks us to build a broad, do-everything agent, our first step is to make it smaller. Find the one task that is repetitive, well-defined, and valuable, and build an agent that does only that, reliably. A narrow agent that people trust is worth more than a general one they have to double-check, because an agent people double-check is not saving anyone time. This is the same instinct we bring to product scope in general, which we wrote about in knowing what to build: success comes from doing one thing well before doing many things badly.
The final step to production is the real work
Say your agent works on 8 out of 10 real inputs. That sounds close to done. It is not. Getting from 8 out of 10 to 99 out of 100 is usually harder than building the first version, and it takes longer.
The reason is that each remaining failure is a rare, specific case. One agent guesses wrong when a date is written in an unusual format. Another calls a tool that takes too long to respond and does not handle the retry. Another gets a request that is technically two requests and only answers half. None of these show up in a demo, and none of them are covered by the core feature. Handling them is slow, unexciting work: reading real transcripts, finding the cases that went wrong, and adding the specific handling that catches each one. This is where the engineering time actually goes, and it is the part founders most often forget to budget for.
There is a useful lesson from machine learning here. Researchers at Google pointed out that in a real ML system, the clever model everyone talks about is only a tiny fraction of the code. The rest is the surrounding code that connects the parts: the data handling, the monitoring, the code that keeps it reliable. Agents follow the same pattern. The model call is small. The work that makes it trustworthy is the system around it. If you plan only for the model call, you have planned for the demo, not the product.
Build the agent to notice when it is unsure
Here is a design idea that separates agents you can trust from ones you cannot. A good agent is built to know its own limits. When it gets an input it is not confident about, it should stop and pass it to a person, not continue and guess.
This sounds obvious and it is rare in practice. The default behavior of a language model is to always produce an answer, even a wrong one, and to sound confident either way. Left alone, an agent inherits that habit: it will take an action on an uncertain guess with the same tone it uses when it is certain. So part of building a reliable agent is teaching it to detect the cases it should not handle, and to send those to a human instead. An agent that handles 70 percent of cases well and correctly passes the other 30 percent to a person is more useful than one that handles 90 percent well and silently gets the last 10 percent wrong, because you can trust the first one and you cannot trust the second.
This is also why we usually start with a human checking the agent's work. Early on, let the agent propose the action and have a person approve it. You catch mistakes before they cost anything, and every approval or correction teaches you where the agent is weak. As the evidence grows that it handles a certain task well, you remove the human from that task and keep them on the risky ones. Trust is earned with data, not assumed from a good demo. We covered the trust question on its own in can you trust an AI agent, because it is the question that decides whether an agent is a real product or only an interesting experiment.
How to actually release one that works
Put it together and a working approach looks like this. Pick one narrow, repetitive task where the value is clear and a mistake is cheap to catch. Build the first version fast, because that part really is quick now. Then spend the larger share of your time on the final step before production: run the agent on a big sample of real inputs, measure how often it does the right thing, find the cases where it fails, and add handling for them one by one. Keep a person checking the agent's actions while you gather that evidence, and only remove that person from the tasks the data says are safe.
The founders who succeed with agents are the ones who treat the demo as the start of the work, not the end. They pick a small job, set a high standard for reliability, and put the work into the unexciting middle stage where trust is built. The ones who struggle are the ones who saw the impressive demo, assumed the demo was the whole thing, and released an agent that looks impressive until a real user gives it something the demo never saw. If you want help deciding those limits, deciding what an agent should and should not do, and building it so people actually trust it, that is a large part of what our AI development work is.
Build the agent narrow. Measure it against real work. Let it earn trust one task at a time. That is the whole difference between a demo and a product.
Related guide: Taking AI agents from prototype to production.
One last thing founders get wrong, and it happens before the agent is built. Because a working agent can be produced so quickly now, teams skip the step where they write down exactly what it should do in each case, and go straight to building. The agent then makes those decisions for them, without anyone seeing, at every point the specification did not cover. When the behaviour turns out to be wrong, it looks like a model problem and gets fixed with a better prompt or a bigger model. It is almost never a model problem. It is an unwritten decision.
Sources
- Sculley et al. (2015), "Hidden Technical Debt in Machine Learning Systems," NeurIPS: https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-systems


