Strategy

An AI demo is not a product

Editorial · Reveneau · July 21, 2026

An AI demo is not a product

The most dangerous moment in an AI project is the first demo that works. Someone connects a model to a real problem, it produces a genuinely impressive answer, and everyone in the meeting is excited. We have been in that meeting many times. The mistake that follows is almost always the same: everyone assumes the hard part is done. It is not. The demo is the first five percent. The other ninety-five is the reason so many AI projects stop making progress between "look what it did" and "people rely on it every day." Here is the work between those two points.

Our own work shows this, and so does research. When Google engineers studied their own machine learning systems, they found that the model itself is a small part of a much larger system: in a mature system, only a fraction of the code is the machine learning part, and the rest is the infrastructure around it, the data pipelines, the serving, the monitoring, the configuration. The demo shows you the small part. People call the rest "the last mile": all the work needed around the model before people can safely depend on it.

1. A demo works once. A product works every time.

A demo has an easy job. It has to work one time, on an input you picked, in front of people who want it to succeed. A product has a much harder job. It has to work on the input a confused user types at midnight, the input that is slightly malformed, the input nobody imagined, and it has to do that thousands of times without a human standing by to pick the good result.

That difference is not a detail. It is most of the engineering. The same feature that took an afternoon to demo can take months to make trustworthy, because "trustworthy" means handling all the cases the demo skipped. We have seen teams show a perfect demo on Monday and spend the next quarter on the inputs that demo never tested.

Here is the pattern in one anonymized example. A team builds an assistant that reads a customer support message and drafts a reply. In the demo, the messages are clean: one clear question, good grammar, one topic. The drafts look excellent. Then real messages arrive. One email is three questions stacked together, one is a forwarded thread with five people in it, one is angry and asks for a refund the policy does not allow. The model that looked ready now needs rules for what to do with each of these, and none of that logic existed in the demo. The idea was sound. The product was still mostly unbuilt.

If you take one thing from this: a demo proves the idea can work. It does not prove the idea works. Those are different claims, and the work needed to go from one to the other is the project.

2. The work after the demo is evaluation, unusual inputs, cost, and failure

When we help teams finish that work, it almost always belongs to one of four groups, and none of them are visible in the demo.

Evaluation comes first. In normal software, you can look at the output and know if it is right. With AI, "right" is often a judgment call, so you need a real way to measure whether answers are good, or you cannot tell if the feature works. Then unusual inputs: the many strange inputs that the normal-use demo never met, which is usually where users decide whether to trust the feature. Then cost, because a model call that is trivial once becomes a large cost when many people use it, and the simple version that worked in the demo can be too expensive to run for real. And finally failure handling: deciding what the system does when it is unsure, because a confident wrong answer is worse than an honest "I do not know."

Evaluation needs more explanation, because it is the one teams skip most often. In ordinary software, a test is simple: the function returns 4, you expected 4, the test passes. With AI, the answer is a paragraph, and "correct" can mean accurate, complete, on-tone, and safe all at once. So the real work is building a set of examples you trust, with the answer you would accept for each, and a way to score new answers against them. That set costs time to make. It takes people who know the subject area sitting down and writing what "good" looks like, case by case. Teams that skip this step cannot tell whether a change to the prompt or the model made the product better or worse, so they end up guessing on the one thing users judge them by.

This work is dull, and it is most of the project. A team that plans for these four from the start finishes the product. A team that treats them as small tasks for the end stops making progress.

3. Budget for the eighty percent that starts after the demo

The planning mistake is treating the demo as the end and the remaining work as a short list of small fixes. It is the reverse. The demo is the first step, and the real build starts once the idea is proven.

So when we scope an AI project now, we are explicit about it. The demo shows the project is worth continuing. The plan, the time, and the people are aimed at everything that comes after: measuring quality, making the feature reliable on unusual inputs, controlling cost, and handling the moments the model gets it wrong. Framed that way, nobody is surprised in month three, because month three was always the point.

The cost of getting this wrong appears in research numbers. A RAND study of why AI projects fail found that more than 80 percent of them do not succeed, about twice the failure rate of other technology projects, and one of the named root causes is underinvesting in the infrastructure needed to deploy and keep a model running. That is the work after the demo, described by outside researchers. The projects that fail are not usually the ones with a bad idea. They are the ones that paid for the demo and did not pay enough for the work that comes after it.

This is not a reason to slow down at the start. Get to a demo fast, because the demo is how you learn the idea is worth pursuing. Just do not confuse reaching it with being nearly done.

4. The work after the demo is where your real advantage comes from

There is a better reason to care about the work after the demo than avoiding failure. It is where you build the part that belongs only to you. The demo is easy to copy, because the model is the same one your competitor can call. Anyone can reach the impressive first answer in an afternoon. What is hard to copy is the set of trusted examples you built to measure quality, the handling you wrote for the strange inputs your users actually send, the cost work that lets you run the feature at a price that makes sense, and the honest behavior you designed for the moments the model is unsure.

Think about two teams releasing the same idea. Both have the same model and the same demo. One stops there and releases it. The other spends the quarter on the four kinds of work above. Six months in, the first team's feature still breaks easily and users have learned not to trust it. The second team's feature keeps working on real, badly formed inputs, so people use it and rely on it. The difference between them was never the model. It was the extra work the second team chose to do. When we tell teams that good execution is what competitors cannot easily copy, this is what we mean in an AI project: the model is shared, so the advantage has to come from the work around it, and that work comes after the demo.

There is a second, less obvious benefit. The measurement system you build to evaluate the feature tells you if today's version is good, and it also lets you improve safely. When a new model comes out, or you change a prompt, you can run it against your trusted examples and see whether the product got better or worse before any user does. The teams that built this can adopt new models with confidence. The teams that skipped it are back to guessing every time something changes, which in AI is often.

What this means for you

When your team shows a demo that works, celebrate for a minute, then reset expectations. Name the four kinds of work still ahead: how you will measure correctness, how you will handle the strange inputs, what it will cost at real volume, and what happens when the system is unsure. Put those on the plan as the main work. And judge progress by how the feature behaves on the inputs you did not choose, because that is what your users will give it.

Thanks to the engineers who have worked with us through the dull middle part of these projects, making features reliable on unusual inputs long after the demo stopped being exciting. The demo is the easy part. The work after it still needs to be done, and that work is what makes the product.

We should be direct about our own position in this, because it could be read two ways. We write one hundred percent of our code with AI, so we are not sceptics about what models can produce. That is exactly why we speak directly about the work after the demo: we see what a first draft written by AI looks like every day, and the pattern is consistent. It handles normal use well and handles little else. The demo is honest about the easy eighty percent. The remaining work is what makes the product.

References

Related guide: How to build an AI product.

Common questions

Why is an AI demo so much easier than a production AI feature?

A demo only has to work once, on inputs you chose, in front of a friendly audience. A product has to work every time, on inputs you did not expect, for users who stop trusting it after a wrong answer. Most of the engineering effort goes into closing the difference between those two.

What work remains after an AI demo works?

It is everything after the idea is proven: evaluating whether answers are actually correct, handling the unusual inputs that break normal use, controlling cost when many people use it, and deciding what the system does when it is unsure. This is usually the majority of the work.

How should teams plan for the work after the demo?

Assume the demo is roughly the first fifth of the work, not the last. Budget time and people for evaluation, unusual inputs, cost, and failure handling from the start, and treat a working demo as the beginning of the real build rather than the end.

Why do so many AI projects stop making progress after a good demo?

Because the team reads the demo as proof the work is done, when it only proves the idea can work. The project then stops making progress between "look what it did" and "people rely on it every day," because that is exactly where the dull, necessary work is.

How do you evaluate whether an AI feature is actually good?

In normal software you can look at the output and know if it is right. With AI, "right" is often a matter of judgment, so you need a real way to measure answer quality against examples you trust. Without that measure you cannot tell how well the feature does the one thing users care about most.

What should an AI system do when it is not sure?

An AI system should be able to say so instead of guessing. A confident wrong answer is worse than an honest "I do not know," so failure handling means deciding in advance what happens when the model is unsure and building that behavior on purpose. This is one of the four kinds of work a demo never shows, alongside evaluation, unusual inputs, and cost, and skipping it is how a feature that breaks easily reaches real users.

Why does cost matter more in production than in a demo?

A single model call is trivial once, but the same call becomes a large cost when thousands of users make it every day. The simple first version that looked fine in the demo can be too expensive to run for real, so cost has to be part of the design from the start.

What are unusual inputs and why do they matter so much?

Unusual inputs are the many strange, badly formed, or unexpected inputs the normal-use demo never met, such as a message that stacks three questions together or a forwarded thread with five people in it. They matter because that is usually where users decide whether to trust the feature, and a demo skips almost all of them, which is why the same feature that took an afternoon to demo can take months to make trustworthy.

Should a slow, careful start be the answer to this?

No, get to a demo fast, because the demo is how you learn the idea is worth pursuing. The goal is to stop treating a working demo as proof the project is nearly done, since the demo is roughly the first fifth of the work. Budget time and people for evaluation, unusual inputs, cost, and failure handling from the start, so the real build begins the moment the idea is proven.

How does Reveneau scope an AI project differently now?

We treat the demo as the first step, which shows the project is worth continuing. The plan, the time, and the people are aimed at what comes after: measuring quality, making the feature reliable on unusual inputs, controlling cost, and handling the moments the model gets it wrong.

How should a team judge progress on an AI feature?

Judge it by how the feature behaves on the inputs you did not choose, because those are the inputs your users will give it. A perfect result on your own examples tells you less than a good result on the cases you never planned for.

Where does a real competitive advantage come from in an AI feature?

Your competitor can call the same model, so the advantage comes from the work around it: the trusted examples you built to measure quality, the handling for strange inputs, the cost work, and the honest behavior when the model is unsure. That work is hard to copy, and it is the work that comes after the demo.