AI

How to add LLM integration to an existing product without breaking it

Editorial · Reveneau · July 28, 2026

How to add LLM integration to an existing product without breaking it

Most of the AI requests we get do not start with "we want to build an AI product." They start smaller than that. A team has a working product, real users, and one workflow that is clearly a fit for a model: a support queue that needs summarizing, a pile of documents that needs classifying, a form that AI could fill in. They do not want a rebuild. They want one good feature, added carefully, that does not put the rest of the product at risk. This is the version of AI work that actually reaches production, and it deserves its own step-by-step method, separate from the agent-and-RAG conversation that gets most of the attention.

Start narrow, not broad

The single biggest predictor of whether a first AI feature is released is how it was scoped. "Add AI to our product" is not a feature. It is only a wish. Nobody can build it, and nobody can tell you afterward whether it worked, because there was never a definition of working.

The fix is to pick something narrow enough that you could write down, in one sentence, what a correct answer looks like. "Summarize this support ticket into three bullet points that a human agent can scan in five seconds." "Classify this incoming lead as sales-ready, needs-nurturing, or spam." "Draft a first-pass reply to this email that a human will edit before sending." Each of those has a form you can grade. You can look at ten outputs and say which are good and which are not, and that single ability is what separates a feature you can release from an idea you can only demo.

Narrow also means you can actually measure improvement. If the feature is "summarize tickets," you can time how long agents spend per ticket before and after, and you have a real number. If the feature is "make support smarter," you have no before-and-after to compare, because you never defined the after. We have seen teams spend months on the second kind of project and come out with nothing they can point to, not because the engineering was bad, but because the goal was never specific enough to reach.

One more reason narrow is better: it is the only way to build a useful evaluation set, which is the thing that tells you if a change made the feature better or worse. A broad, vague feature cannot be evaluated because there is no fixed target to score against. A narrow one can, from the start.

Isolate the AI code so a bad model call cannot cause your product to fail

Once you know what you are building, the next decision is architectural, and it matters more than which model you pick. The AI code has to be kept behind a separate interface the rest of your product does not need to know about.

In practice that means two things. First, the model call is kept behind a single function or service with a clear interface: give it an input, get back an output or a defined failure. Nothing else in your codebase should call the model directly, the same way nothing else should query your database directly instead of going through your data layer. This is ordinary software design, and it is more important here than almost anywhere else, because the thing behind that interface is the least predictable dependency your product has ever had. A database query fails in a small number of well-understood ways. A model call can fail, time out, return malformed output, or return a confident answer that is simply wrong, and your code needs a plan for each.

Second, put the feature behind a flag. Release it switched off, turn it on for a small share of traffic, watch it, and only widen the rollout once you trust it. A flag also gives you an instant off switch. If the feature misbehaves at 2 a.m., you turn it off in seconds instead of rolling back a deploy, and the product keeps running exactly as it did before you touched it. This is not a new idea. It is how any team releases a risky change safely. AI features are simply the case where skipping it costs the most, because the failure mode is not a crash you can see. It is a wrong answer shown to a user with total confidence.

The fallback deserves its own sentence, because it is the part teams design last, if at all. When the model call fails, what does the user see? The honest answer should be: whatever they saw before this feature existed. The plain text box instead of the AI-drafted reply. The unsorted list instead of the AI-ranked one. If your fallback is a loading icon that never stops or an error page, you have built a feature that can cause the rest of the product to fail, and that is the exact risk a team with an existing product cannot afford to take on. The whole point of adding one feature to something that already works is that the something that already works keeps working.

Hosted API or run it yourself: the real tradeoff

For a first feature, this decision is usually simpler than it looks. Call a hosted model API. The reasons add up quickly: no GPU capacity to set up, no model weights to keep updated, no on-call rotation for a service you would now own, and a cost structure that scales with usage instead of sitting there as a fixed bill whether you use it or not.

The tradeoff is real, though, and worth naming honestly rather than ignoring. A hosted API charges per request, and at high enough volume that per-request cost adds up to more than the fixed cost of running your own model would have been. It also means sending your data to a third party, which is a real consideration for some industries and some data, not a formality. And you are depending on someone else's uptime and someone else's pricing changes.

Self-hosting reverses each of those. You control the data path, and at high enough volume the fixed cost of GPU time can be lower than the per-request bill. But you take on the operational work: keeping a model server running, handling updates, capacity planning for spikes, and building the monitoring you would have gotten for less effort from a hosted provider. For a single first feature, that work is rarely worth it. It becomes worth revisiting only once you have a real number: your actual request volume, your actual per-request hosted cost, and the actual fixed cost of running the alternative. Decide with that number, not with a guess that self-hosting must be cheaper because you own the hardware.

Latency is the other tradeoff, and it changes how the feature feels to users more than people expect. A database read returns in milliseconds. A model call commonly takes one to several seconds, sometimes longer for a longer output. That is not a defect, it is the nature of the work, but it means the interface has to be designed around it: a visible loading state, streaming the answer token by token so the user sees progress instead of a frozen screen, or setting the expectation that this is a draft to review rather than an instant result. A feature that ignores this and just blocks the UI until the model responds will feel broken even when it is working exactly as intended.

The separate interface we described in the last section helps again here. If the model call is kept behind a clean interface, switching from a hosted API to a self-hosted model later, or from one provider to another, is a change inside that interface. The rest of your product does not need to change at all.

How to know if it is working after launch

Releasing the feature is not the end of the work. The question that matters is whether it is actually good, and that question needs an answer that is not a feeling.

Build a small evaluation set before launch: a set of real, representative inputs paired with the answer you would accept for each. This does not need to be large to be useful. Even a few dozen well-chosen cases, covering the normal case and the unusual cases you can already think of, gives you something to score new outputs against every time you change a prompt, swap a model, or adjust how much context you send. Without this, you cannot tell whether a change made the feature better or worse. You can only guess, and guessing is exactly the habit that turns a promising feature into one that stops improving.

After launch, watch two different things, because they answer two different questions. First, the outcome metric the feature exists to improve: time per ticket, reply acceptance rate, whatever the original narrow goal was. That tells you whether the feature is delivering value. Second, the operational signals: how often the fallback path is used, how often the model returns something your safety checks reject, how latency and cost move as real traffic arrives instead of your test traffic. That tells you whether the feature is working properly, which is a different question from whether it is valuable, and you need both answers.

Real users will hand the feature inputs you did not think of. That is not a failure of your planning, it is the reason evaluation has to continue after launch instead of stopping once the feature is released. The teams that keep watching this catch a feature that is slowly getting worse in a week. The teams that release and move to other work find out from a support ticket three months later, after users have already learned not to trust the feature.

Where Reveneau fits

Adding one dependable AI feature to a product that already works is one of the most common requests inside our AI development work, and it is a smaller, more contained job than a full AI product build. The scoping discipline, the isolation behind a clear interface, and the evaluation habits above are the same ones we bring to every AI engagement, just applied to a single feature instead of a whole system. If you want the longer argument for why a working demo and a production-ready feature are not the same thing, we wrote about that gap directly in an AI demo is not a product.

Thanks to the engineers and support teams who have shown us the messy real inputs that arrived after launch, the ones no demo ever received. The model gets you an impressive first answer. The integration is what turns that answer into a feature your users can depend on without your team worrying about it.

It is also the kind of job where generated code helps most and needs watching most. The integration itself is common work and is finished quickly. The parts specific to your product, where exactly the feature sits, what it must not touch, what happens when the model is unavailable, are the parts a model will fill in with something reasonable if your specification does not say. Those are the sentences worth writing carefully, and they are the ones we spend the first days on.

Common questions

What is LLM integration for an existing product?

It means adding one AI-powered feature, such as a summarizer, classifier, drafting assistant, or smart search, into software that already works, without rebuilding the product around the model. The model call is kept behind a defined interface in the existing codebase, with a feature flag and a fallback, so the rest of the product keeps working exactly as it did if that call fails or times out.

How do we pick the right first AI feature to add?

Pick the narrowest one you can grade, not the broadest one you can imagine. A feature like "summarize this ticket into three bullet points" has a clear right and wrong answer. A feature like "make the product smarter with AI" does not, and you will not be able to tell if it worked.

Will adding an LLM feature break our existing product?

Not if you isolate it correctly. Put the model call behind a feature flag and a clean interface, so the rest of the code calls a function and gets a result, with no awareness of what happens inside. If the model call fails or times out, the product should fall back to its prior behavior instead of erroring out.

What is a feature flag and why does it matter here?

A feature flag is a switch that turns a feature on or off without a new deploy. For an AI feature it means you can roll it out to five percent of users, watch what happens, and turn it off in seconds if something looks wrong, instead of rolling back a release.

Should we call a hosted model API or run a model ourselves?

For a first feature, almost always call a hosted API. Running your own model adds GPU costs, on-call work, and update work that is only worth it at a volume and cost sensitivity most teams have not reached yet. Revisit self-hosting later if a specific cost or latency number justifies it.

What are the real cost tradeoffs of a hosted model API versus self-hosting?

A hosted API charges per request with no fixed cost and near-zero operational work, which suits unpredictable or low-to-moderate volume. Self-hosting has a real fixed cost in GPU time and engineering attention that only pays for itself past a volume threshold you should calculate, not guess at, before committing.

How much latency does an LLM call add to a product?

Enough that it changes how the feature should feel. A model call commonly takes one to several seconds, versus milliseconds for a database read, so the interface needs to show progress or stream partial output rather than pretending the response is instant. A feature that ignores this and blocks the interface until the model responds will feel broken even when it is working exactly as intended.

How do we know if an AI feature is actually working after launch?

Build an evaluation set of real cases with the answer you would accept for each, and score new outputs against it before and after any change. Pair that with production monitoring of the metric the feature exists to improve, and a way to see how often the fallback path is used.

What happens if the model call fails in production?

The product should degrade to its behavior before the AI feature existed: the plain text box instead of the AI draft, the raw list instead of the AI-ranked one. That fallback only works if it was designed on purpose, not added after the first outage.

Do we need a data science team to add one AI feature to our product?

Most single-feature integrations are an engineering problem: calling a hosted API correctly, isolating the code behind a clear interface, measuring outputs against an evaluation set, and handling failure with a defined fallback. A data science team matters more for training your own models, which most first features do not require and which is rarely worth it until a specific volume threshold is reached.

How long does it typically take to add one AI feature to an existing product?

A well-scoped, single feature with a hosted model API is usually weeks, not quarters, if the team resists the urge to widen scope in the middle of the build. The scoping discipline matters more than the raw engineering time, since picking a narrow, gradable feature up front is what keeps the build small, while a vague goal like "make the product smarter" can continue for months with no useful result.

Does Reveneau help with LLM integration into products that already exist?

Yes. It is one of the most common requests in our AI development work: adding one dependable AI feature to a product you already run, isolated behind a clear interface so the rest of the product keeps working if the model call fails. The same scoping discipline, isolation, and evaluation habits apply whether the request is a summarizer, a classifier, or a smart search feature.