AI

What founders get wrong about AI agents

Editorial · Reveneau · August 10, 2026

What founders get wrong about AI agents

We have built a lot of AI features over the past few years, and AI agents are the thing founders ask about most right now. The conversation almost always starts the same way. Someone saw a demo, an agent that booked a meeting or closed a support ticket entirely on its own, and it looked almost impossible. They want that. The problem is not that the excitement is wrong. The demo really was impressive. The problem is that the demo and a product you can actually release are two different things, and the work between them is where most of the work, most of the cost, and almost all of the risk is.

Here is how we think about agents, and where we see founders get it wrong.

The demo is the easy part

A good agent demo is designed to go well. Someone picked the task, picked the input, and ran it on a clean example where the data was there and the request was clear. In that example, a modern model looks excellent. It reads the request, calls the right tool, takes the action, and reports back. It feels finished.

It is not finished. It is the start. A demo shows you the 80 percent of cases that go right. A product has to handle the other cases: the request phrased in a way nobody expected, the field that is missing, the tool that returns something strange, the user who asks for two things at once. On those inputs, the model does not stop and say "I am not sure." It acts on its best guess, and its best guess is sometimes wrong. In a chatbot, a wrong guess is an awkward answer. In an agent, a wrong guess is a wrong action: the wrong meeting booked, the wrong refund issued, the wrong record deleted. That is a real cost, and it is the reason agents are harder to release than chat features.

We wrote about this gap in more detail in why the final step from demo to production is the real job. The short version: the impressive part is cheap now, and the reliable part is the whole engineering problem.

A narrow agent works better than a general one, by a large margin

The most common wrong instinct we see is wanting the agent to do everything. A founder imagines one smart assistant that handles support, sales, scheduling, and internal questions, all in one. It sounds efficient. It is a mistake.

The wider you make an agent's job, the harder it is to know whether it is working. If an agent can take a hundred different actions, you cannot test all the ways those actions combine, and you cannot define what "correct" even means across all of them. A narrow agent has a small, clear job. Answer this one type of question. Sort these requests into three groups. Draft the first reply to this kind of ticket. When the job is narrow, you can write down what correct looks like, build tests for it, measure how often it succeeds, and catch it when it fails. That is what matters. You cannot trust what you cannot measure, and you cannot measure a job that has no clear limits.

So when someone asks us to build a broad, do-everything agent, our first step is to make it smaller. Find the one task that is repetitive, well-defined, and valuable, and build an agent that does only that, reliably. A narrow agent that people trust is worth more than a general one they have to double-check, because an agent people double-check is not saving anyone time. This is the same instinct we bring to product scope in general, which we wrote about in knowing what to build: success comes from doing one thing well before doing many things badly.

The final step to production is the real work

Say your agent works on 8 out of 10 real inputs. That sounds close to done. It is not. Getting from 8 out of 10 to 99 out of 100 is usually harder than building the first version, and it takes longer.

The reason is that each remaining failure is a rare, specific case. One agent guesses wrong when a date is written in an unusual format. Another calls a tool that takes too long to respond and does not handle the retry. Another gets a request that is technically two requests and only answers half. None of these show up in a demo, and none of them are covered by the core feature. Handling them is slow, unexciting work: reading real transcripts, finding the cases that went wrong, and adding the specific handling that catches each one. This is where the engineering time actually goes, and it is the part founders most often forget to budget for.

There is a useful lesson from machine learning here. Researchers at Google pointed out that in a real ML system, the clever model everyone talks about is only a tiny fraction of the code. The rest is the surrounding code that connects the parts: the data handling, the monitoring, the code that keeps it reliable. Agents follow the same pattern. The model call is small. The work that makes it trustworthy is the system around it. If you plan only for the model call, you have planned for the demo, not the product.

Build the agent to notice when it is unsure

Here is a design idea that separates agents you can trust from ones you cannot. A good agent is built to know its own limits. When it gets an input it is not confident about, it should stop and pass it to a person, not continue and guess.

This sounds obvious and it is rare in practice. The default behavior of a language model is to always produce an answer, even a wrong one, and to sound confident either way. Left alone, an agent inherits that habit: it will take an action on an uncertain guess with the same tone it uses when it is certain. So part of building a reliable agent is teaching it to detect the cases it should not handle, and to send those to a human instead. An agent that handles 70 percent of cases well and correctly passes the other 30 percent to a person is more useful than one that handles 90 percent well and silently gets the last 10 percent wrong, because you can trust the first one and you cannot trust the second.

This is also why we usually start with a human checking the agent's work. Early on, let the agent propose the action and have a person approve it. You catch mistakes before they cost anything, and every approval or correction teaches you where the agent is weak. As the evidence grows that it handles a certain task well, you remove the human from that task and keep them on the risky ones. Trust is earned with data, not assumed from a good demo. We covered the trust question on its own in can you trust an AI agent, because it is the question that decides whether an agent is a real product or only an interesting experiment.

How to actually release one that works

Put it together and a working approach looks like this. Pick one narrow, repetitive task where the value is clear and a mistake is cheap to catch. Build the first version fast, because that part really is quick now. Then spend the larger share of your time on the final step before production: run the agent on a big sample of real inputs, measure how often it does the right thing, find the cases where it fails, and add handling for them one by one. Keep a person checking the agent's actions while you gather that evidence, and only remove that person from the tasks the data says are safe.

The founders who succeed with agents are the ones who treat the demo as the start of the work, not the end. They pick a small job, set a high standard for reliability, and put the work into the unexciting middle stage where trust is built. The ones who struggle are the ones who saw the impressive demo, assumed the demo was the whole thing, and released an agent that looks impressive until a real user gives it something the demo never saw. If you want help deciding those limits, deciding what an agent should and should not do, and building it so people actually trust it, that is a large part of what our AI development work is.

Build the agent narrow. Measure it against real work. Let it earn trust one task at a time. That is the whole difference between a demo and a product.

Related guide: Taking AI agents from prototype to production.

One last thing founders get wrong, and it happens before the agent is built. Because a working agent can be produced so quickly now, teams skip the step where they write down exactly what it should do in each case, and go straight to building. The agent then makes those decisions for them, without anyone seeing, at every point the specification did not cover. When the behaviour turns out to be wrong, it looks like a model problem and gets fixed with a better prompt or a bigger model. It is almost never a model problem. It is an unwritten decision.

Sources

Common questions

What is an AI agent?

An AI agent is software that uses a language model to take actions on its own, not just answer questions. It can call tools, read data, and complete a task across several steps, such as booking a meeting or resolving a support ticket. The hard part is not making it act, it is making it act correctly every time, because a wrong action has a real cost.

Why does a great agent demo not mean a working product?

A demo shows the agent on a clean, chosen example where everything goes right. A product has to handle the messy inputs, the missing data, and the strange requests that real users bring. The work between the two is the final step before production, and it is where most of the engineering time actually goes.

What is the final-step problem with AI agents?

The final step is the work between an agent that works most of the time and one you can trust to run without a human watching. Getting from 80 percent reliable to 99 percent reliable is often harder than building the first version, because each remaining failure is a rare unusual case that needs its own handling.

Why is a narrow agent better than a general one?

A narrow agent has a small, clear job, so you can define what correct looks like, test it, and catch mistakes. A general agent has a huge range of possible actions, which makes it almost impossible to test or trust. Narrow scope is not a limitation, it is what makes reliability possible.

How do I know if an agent is reliable enough to release?

Measure it against real tasks, not demo tasks. Run it on a large sample of actual inputs, count how often it does the right thing, and decide what failure rate you can accept given the cost of a mistake. If a single wrong action is expensive, your standard has to be high.

Should a human check an AI agent's actions?

Often yes, at least at first. Letting the agent propose an action that a person approves catches mistakes while you learn where it fails. As you gather evidence that it handles a task well, you can remove the human from the low-risk tasks and keep them on the risky ones.

What tasks are a good fit for an AI agent today?

Tasks that are repetitive, well-defined, and where a mistake is cheap to catch and fix. Drafting a first response, sorting requests, pulling data together, and similar work are good fits. Tasks where one wrong action causes real harm need much more care and usually a human check.

How long does it take to build a reliable AI agent?

The first working version can come together in weeks. Getting it reliable enough to trust in production takes longer, because that time goes into unusual cases, testing, and safety checks, not the core feature. Founders who budget only for the demo are surprised by the second half.

Why do AI agents fail in production?

They usually fail on inputs the builder did not anticipate: unusual phrasing, missing fields, or a tool that returns something unexpected. The model does not stop and ask, it acts on its best guess, and a wrong guess becomes a wrong action. Good agents are built to notice uncertainty and stop instead of guessing.

Do I need a huge model to build a good agent?

A smaller, well-scoped setup with good safety checks often works better than a bigger model given a vague task. The quality of an agent comes more from clear scope, good testing, and sensible limits than from raw model size, since a narrow job is one you can measure and trust, while a broad job given to a large model is still impossible to fully test.

What is the biggest mistake founders make with AI agents?

Treating the demo as the end of the work. The demo is the easy 80 percent, run on a clean example where the input was chosen to go well. The reliability, the unusual cases, and the safety checks are the harder work that follows, and skipping them releases an agent that looks impressive until a real user gives it something the demo never saw.