AI

How to tell if an AI feature idea is worth building

Editorial · Reveneau · August 8, 2026

How to tell if an AI feature idea is worth building

Every week someone shows us an AI feature idea that works beautifully in a demo, and asks how fast we can release it. The demo really is impressive. The model does exactly the clever thing, the audience nods, and it feels like the hard part is done and only the building remains. That feeling is almost always wrong. A demo has to work once, on a friendly example, in front of people who want it to succeed. A product has to work every day, on messy real input, for users who will give it things nobody imagined. The work between those two is where most AI features quietly fail. So before we build one, we test the idea with four questions, and they remove many ideas that looked great in a demo.

Here is how we decide whether an AI feature is worth building.

Does it solve a real problem

Start with the least exciting question, because it rejects the most ideas. Does this feature solve a real problem that someone actually has. Not "would this be cool," not "does the competitor have one," not "can we put AI on the slide." Does it remove a genuine problem from a real user's day.

A lot of AI ideas fail this quietly, because they start from the technology and go looking for a use. Someone sees what a model can do and works backwards to a feature, which produces things that are clever and unwanted. The better direction is the reverse: start from a problem your users have, then ask whether AI is the best way to remove it. Sometimes the honest answer is that the problem does not need AI at all. A simple rule, a better search, or a clearer interface can solve it with none of the unpredictability that AI adds. Use AI when the problem genuinely needs judgment over messy input, and use something simpler and more reliable when it does not. This is the same discipline we wrote about in knowing what to build, applied to AI.

Is roughly right good enough

Now the question that decides more AI features than any other, and that people skip because it feels technical. AI systems are usually right most of the time, not all of the time. So the real question is whether your specific use can accept being occasionally wrong.

The clearest way to see it is with two examples. An AI that drafts a first version of an email for a person to edit can be roughly right and still useful, because a human reads it and fixes anything wrong before it goes out. An AI that calculates how much tax someone owes cannot be roughly right, because in that job roughly right is just wrong, and a confident wrong number is worse than no number at all. Same underlying technology, completely different fit, and the only thing that changed is how much error the use can accept.

So for any AI feature, ask honestly how much wrongness the use tolerates. Features where a person reviews the output, or where a wrong answer just means a worse suggestion, have a lot of room. Features where the output goes straight into something that must be exact, or straight into an action nobody checks, have almost none. Match the feature to the error the situation can accept, and you avoid the most common way AI features hurt people: being trusted for a job they were never accurate enough to do.

What does it cost to be wrong

Closely related, but worth its own question, is the cost of a mistake. Not how often the AI is wrong, but how bad it is when it is.

Do this concretely. Imagine the worst plausible mistake the feature could make on a normal day, and ask who gets hurt and how much. If the worst case is a worse suggestion that the user ignores, the cost is low, and you have room to release and improve. If the worst case is lost money, a safety problem, a broken legal promise, or a user misled about something that matters, the cost is high, and the feature needs safety checks, human review, or a different design before it reaches real users.

This is really a question about how much you let the AI act on its own. A feature where an AI agent takes low-risk, reversible actions with a person able to review is reasonable. A feature where an agent takes high-risk, irreversible actions by itself is not, at least not without a lot of care. We covered exactly this trust question in detail in can you trust an AI agent, and the short version is: decide how much autonomy the feature can safely have before you build it, based on the cost of being wrong, not after you have already released it and found out.

The demo-to-product gap

Now the one that surprises even good teams. The gap between a demo that works and a product that works is almost always much bigger than it looks, and underestimating it is the most common AI planning mistake we see.

A demo hides the hard parts by design. It runs on a curated example, picked because it works, in a controlled setting. The real world sends rare inputs, users who do unexpected things, adversarial users who try to break it, and many rare situations nobody thought of. Handling all of that reliably is usually most of the actual engineering, and none of it is visible in the five minutes where the demo works well. We wrote about this plainly in the AI demo versus the last mile: the demo is the first step, and the final step, making it reliable on every input real users give it, is where the work and the risk are.

The practical fix is cheap and honest. Before committing the full budget, build the smallest version that runs on real, messy input rather than a hand-picked demo, and see how often it is wrong and how bad the wrong cases are. That one test tells you how much work it will take to make the feature reliable, before you spend the money to do that work. A quick, rough test against real data is more useful than a polished demo on easy examples every time, because it measures the thing that actually determines whether the feature can be released.

Putting the four questions together

None of these questions needs deep machine learning knowledge, which is the point. Does it solve a real problem. Is roughly right good enough. What does it cost to be wrong. How big is the gap between the demo and something reliable. Test any AI idea with those four, and the answer becomes clear.

The strong ideas have things in common. They remove a real problem, they tolerate being occasionally wrong, the cost of a mistake is low or limited by a person who reviews the output, and the work of making them reliable, while real, is possible. Build those. The weak ideas also have things in common. They exist mostly to look modern, or they demand accuracy the technology cannot give, or a wrong answer is expensive, or the demo hid reliability work as large as the whole project. Be careful with those, no matter how good the demo looked, because the demo was never the thing you were releasing. Keep the final decision with the people who understand both the product and the cost of being wrong, not just whoever was most impressed on the day. The question was never whether the AI can do something clever once. It is whether it can do the right thing, reliably enough, for a price you can accept when it fails. Answer that, and you will build the AI features that last instead of the ones that only ever worked in a demo.

There is a mistake that has appeared only recently, and good teams make it. Building the feature is now so cheap that "let us just try it" feels like the sensible answer to any of these four questions. Sometimes it is. But a feature you can release in three days is still a feature you have to maintain, monitor, evaluate, and eventually explain to a customer when it behaves oddly. The build was never the expensive part of a bad idea. Cheap building does not make a weak feature worth having, it just removes the delay that used to give you time to reconsider.

Common questions

How do I know if an AI feature is worth building?

Ask four things: does it solve a real problem people have, is roughly right good enough for this use, what does it cost when the AI is wrong, and how big is the gap between a demo and something reliable. If the idea solves a real problem, tolerates being occasionally wrong, and the cost of a mistake is low, it is a strong candidate. If any of those fail, be careful.

Why do so many AI features fail after a great demo?

Because a demo only has to work once on an easy example, and a product has to work every day on messy real input. The work between those two is making the feature reliable, and it is where most of the real engineering happens. A feature that impresses people in a five-minute demo can still be too unreliable to release.

What does roughly right good enough mean for AI features?

AI systems are usually right most of the time, not all of the time, so the question is whether your use can tolerate being occasionally wrong. Suggesting a draft email is fine when roughly right, because a person edits it. Calculating someone's tax owed is not, because there roughly right is just wrong. Match the feature to how much error it can accept.

How do I judge the cost of an AI being wrong?

Imagine the worst plausible mistake the feature could make and ask who gets hurt and how badly. If a wrong answer means a worse suggestion, the cost is low and you have room to release it. If a wrong answer means lost money, a safety issue, or a broken legal promise, the cost is high and the feature needs safety checks, human review, or a new design.

What is the demo-to-product gap in AI?

It is the difference between an AI feature that works in a controlled demo and one that works reliably for real users at scale. Demos hide the hard parts: rare inputs, adversarial users, unusual cases, and the many rare situations you did not think of. Getting from the demo to a reliable product is usually most of the work, and underestimating it is the most common AI planning mistake.

Should I build an AI feature just because competitors have one?

No. Copying a competitor's AI feature without checking whether it solves a real problem for your users just adds cost and risk. The right reason to build is that it removes a genuine problem for your customers, not that it lets you put AI on a slide. Many AI features exist to look modern rather than to help anyone.

How much should I trust an AI agent inside a feature?

Trust it in proportion to the cost of it being wrong and your ability to catch mistakes. An agent taking low-risk, reversible actions with a person able to review is reasonable; an agent taking high-risk, irreversible actions on its own is not. Decide how much autonomy the feature can safely have before you build it, not after.

What is the first thing to check on any AI feature idea?

Whether it solves a real problem someone actually has. This sounds obvious, but many AI ideas start from the technology and look for a use, which produces features that are clever and unwanted. Start from a problem your users have, then ask whether AI is the best way to remove it, not the other way around.

Do I need to build AI to solve the problem?

Often not, and that is worth checking early. Some problems that look like AI problems are solved better by a simple rule, a search, or a better interface, with none of the unpredictability AI adds. Use AI when the problem genuinely needs judgment over messy input, and use something simpler and more reliable when it does not.

How do I test an AI feature idea cheaply before committing?

Build the smallest version that runs on real, messy input, not a curated demo, and see how often it is wrong and how bad the wrong cases are. This tells you how much work it will take to make it reliable before you spend the full budget. A quick, honest test against real data is more useful than a polished demo on hand-picked examples every time.

Who should decide whether an AI feature is released?

People who understand both the product and the cost of the feature being wrong, not just whoever was impressed by the demo. The decision depends on whether the error rate and the cost of errors are acceptable for real users, which is a judgment about the business, not just the technology. Keep that decision with the people who own the outcome.