How to add LLM integration to an existing product without breaking it

Most of the AI requests we get do not start with "we want to build an AI product." They start smaller than that. A team has a working product, real users, and one workflow that is clearly a fit for a model: a support queue that needs summarizing, a pile of documents that needs classifying, a form that AI could fill in. They do not want a rebuild. They want one good feature, added carefully, that does not put the rest of the product at risk. This is the version of AI work that actually reaches production, and it deserves its own step-by-step method, separate from the agent-and-RAG conversation that gets most of the attention.
Start narrow, not broad
The single biggest predictor of whether a first AI feature is released is how it was scoped. "Add AI to our product" is not a feature. It is only a wish. Nobody can build it, and nobody can tell you afterward whether it worked, because there was never a definition of working.
The fix is to pick something narrow enough that you could write down, in one sentence, what a correct answer looks like. "Summarize this support ticket into three bullet points that a human agent can scan in five seconds." "Classify this incoming lead as sales-ready, needs-nurturing, or spam." "Draft a first-pass reply to this email that a human will edit before sending." Each of those has a form you can grade. You can look at ten outputs and say which are good and which are not, and that single ability is what separates a feature you can release from an idea you can only demo.
Narrow also means you can actually measure improvement. If the feature is "summarize tickets," you can time how long agents spend per ticket before and after, and you have a real number. If the feature is "make support smarter," you have no before-and-after to compare, because you never defined the after. We have seen teams spend months on the second kind of project and come out with nothing they can point to, not because the engineering was bad, but because the goal was never specific enough to reach.
One more reason narrow is better: it is the only way to build a useful evaluation set, which is the thing that tells you if a change made the feature better or worse. A broad, vague feature cannot be evaluated because there is no fixed target to score against. A narrow one can, from the start.
Isolate the AI code so a bad model call cannot cause your product to fail
Once you know what you are building, the next decision is architectural, and it matters more than which model you pick. The AI code has to be kept behind a separate interface the rest of your product does not need to know about.
In practice that means two things. First, the model call is kept behind a single function or service with a clear interface: give it an input, get back an output or a defined failure. Nothing else in your codebase should call the model directly, the same way nothing else should query your database directly instead of going through your data layer. This is ordinary software design, and it is more important here than almost anywhere else, because the thing behind that interface is the least predictable dependency your product has ever had. A database query fails in a small number of well-understood ways. A model call can fail, time out, return malformed output, or return a confident answer that is simply wrong, and your code needs a plan for each.
Second, put the feature behind a flag. Release it switched off, turn it on for a small share of traffic, watch it, and only widen the rollout once you trust it. A flag also gives you an instant off switch. If the feature misbehaves at 2 a.m., you turn it off in seconds instead of rolling back a deploy, and the product keeps running exactly as it did before you touched it. This is not a new idea. It is how any team releases a risky change safely. AI features are simply the case where skipping it costs the most, because the failure mode is not a crash you can see. It is a wrong answer shown to a user with total confidence.
The fallback deserves its own sentence, because it is the part teams design last, if at all. When the model call fails, what does the user see? The honest answer should be: whatever they saw before this feature existed. The plain text box instead of the AI-drafted reply. The unsorted list instead of the AI-ranked one. If your fallback is a loading icon that never stops or an error page, you have built a feature that can cause the rest of the product to fail, and that is the exact risk a team with an existing product cannot afford to take on. The whole point of adding one feature to something that already works is that the something that already works keeps working.
Hosted API or run it yourself: the real tradeoff
For a first feature, this decision is usually simpler than it looks. Call a hosted model API. The reasons add up quickly: no GPU capacity to set up, no model weights to keep updated, no on-call rotation for a service you would now own, and a cost structure that scales with usage instead of sitting there as a fixed bill whether you use it or not.
The tradeoff is real, though, and worth naming honestly rather than ignoring. A hosted API charges per request, and at high enough volume that per-request cost adds up to more than the fixed cost of running your own model would have been. It also means sending your data to a third party, which is a real consideration for some industries and some data, not a formality. And you are depending on someone else's uptime and someone else's pricing changes.
Self-hosting reverses each of those. You control the data path, and at high enough volume the fixed cost of GPU time can be lower than the per-request bill. But you take on the operational work: keeping a model server running, handling updates, capacity planning for spikes, and building the monitoring you would have gotten for less effort from a hosted provider. For a single first feature, that work is rarely worth it. It becomes worth revisiting only once you have a real number: your actual request volume, your actual per-request hosted cost, and the actual fixed cost of running the alternative. Decide with that number, not with a guess that self-hosting must be cheaper because you own the hardware.
Latency is the other tradeoff, and it changes how the feature feels to users more than people expect. A database read returns in milliseconds. A model call commonly takes one to several seconds, sometimes longer for a longer output. That is not a defect, it is the nature of the work, but it means the interface has to be designed around it: a visible loading state, streaming the answer token by token so the user sees progress instead of a frozen screen, or setting the expectation that this is a draft to review rather than an instant result. A feature that ignores this and just blocks the UI until the model responds will feel broken even when it is working exactly as intended.
The separate interface we described in the last section helps again here. If the model call is kept behind a clean interface, switching from a hosted API to a self-hosted model later, or from one provider to another, is a change inside that interface. The rest of your product does not need to change at all.
How to know if it is working after launch
Releasing the feature is not the end of the work. The question that matters is whether it is actually good, and that question needs an answer that is not a feeling.
Build a small evaluation set before launch: a set of real, representative inputs paired with the answer you would accept for each. This does not need to be large to be useful. Even a few dozen well-chosen cases, covering the normal case and the unusual cases you can already think of, gives you something to score new outputs against every time you change a prompt, swap a model, or adjust how much context you send. Without this, you cannot tell whether a change made the feature better or worse. You can only guess, and guessing is exactly the habit that turns a promising feature into one that stops improving.
After launch, watch two different things, because they answer two different questions. First, the outcome metric the feature exists to improve: time per ticket, reply acceptance rate, whatever the original narrow goal was. That tells you whether the feature is delivering value. Second, the operational signals: how often the fallback path is used, how often the model returns something your safety checks reject, how latency and cost move as real traffic arrives instead of your test traffic. That tells you whether the feature is working properly, which is a different question from whether it is valuable, and you need both answers.
Real users will hand the feature inputs you did not think of. That is not a failure of your planning, it is the reason evaluation has to continue after launch instead of stopping once the feature is released. The teams that keep watching this catch a feature that is slowly getting worse in a week. The teams that release and move to other work find out from a support ticket three months later, after users have already learned not to trust the feature.
Where Reveneau fits
Adding one dependable AI feature to a product that already works is one of the most common requests inside our AI development work, and it is a smaller, more contained job than a full AI product build. The scoping discipline, the isolation behind a clear interface, and the evaluation habits above are the same ones we bring to every AI engagement, just applied to a single feature instead of a whole system. If you want the longer argument for why a working demo and a production-ready feature are not the same thing, we wrote about that gap directly in an AI demo is not a product.
Thanks to the engineers and support teams who have shown us the messy real inputs that arrived after launch, the ones no demo ever received. The model gets you an impressive first answer. The integration is what turns that answer into a feature your users can depend on without your team worrying about it.
It is also the kind of job where generated code helps most and needs watching most. The integration itself is common work and is finished quickly. The parts specific to your product, where exactly the feature sits, what it must not touch, what happens when the model is unavailable, are the parts a model will fill in with something reasonable if your specification does not say. Those are the sentences worth writing carefully, and they are the ones we spend the first days on.


