Taking AI agents from prototype to production / Plan the work
What it costs and how long it takes to release an AI agent
Any single figure quoted for building an AI agent should be treated with suspicion, because the honest answer depends entirely on the task and what it connects to. What is predictable is the pattern of the spending. The prototype is fast and cheap, which sets a false expectation, and the real cost sits in evaluation, integration, permissions, and the many rare, unusual inputs.
Published August 8, 2026. Updated September 30, 2026. Editorial.
Key takeaways
- The prototype is the cheap part. Treating it as most of the work is the most common budgeting error.
- Evaluation, integration, and strengthening against failure are where the time goes, and none of them improve the demo.
- Running cost scales with use, because a multi-step agent makes several model calls per request.
- Budget the prototype in weeks and the production system in months, with evaluation as a permanent line item.
We will not put a number on "an AI agent" here, because the number would be invented. A narrow internal agent reading one system and drafting text for a person to approve is a different project from an agent taking actions across several production systems with money involved. What we can describe honestly is where the effort goes, which is more useful for planning than a figure that stops being accurate once your real requirements are known.
The prototype is the cheap part, and that misleads people
Getting an agent to work impressively is often a matter of days. Modern tooling makes it genuinely fast, and that speed is real.
The problem is what it implies. A prototype that took a week suggests the finished system might take a month. In practice the prototype is a small fraction of the work, because it was solving a fundamentally easier problem: succeed a few times on chosen inputs with a friendly person watching. Production has to succeed repeatedly on inputs nobody chose, for users who will not interpret a rough answer generously, with something at stake when it is wrong.
Every estimate that goes wrong tends to go wrong here, in the assumption that the remaining work is a continuation of the prototype rather than a different kind of work.
Where the time actually goes
Evaluation. Defining correctness, gathering real cases, building the scoring, and validating that the scoring agrees with human judgment. This is substantial work, it produces nothing visible, and skipping it is why projects stall indefinitely.
Integration. Connecting properly to real systems: authentication, scoped permissions, rate limits, pagination, timeouts, and the specific ways each dependency fails. This is ordinary engineering and takes ordinary engineering time, but it is absent from the prototype and therefore absent from the estimate that was based on it.
Strengthening against failure. Tool validation, structured outputs, retries, fallbacks, step and cost limits, and logging. Individually small, collectively significant.
The many rare inputs. After launch, real users produce inputs nobody anticipated, and each one needs a decision: fix, accept, or refuse. This never entirely finishes, because new rare inputs keep appearing. It slows down, and it does not stop.
Deciding autonomy. Often underestimated because it is organizational rather than technical. Working out what the agent may do alone, getting agreement from the people accountable for the consequences, and building whatever review step that requires.
Running cost behaves differently from ordinary software
Traditional software has costs that are largely fixed against usage. Agents do not, and the difference matters for the business case.
An agent that reasons through several steps makes several model calls per request, plus whatever the tools cost. Per-request cost therefore grows with volume rather than staying the same, and it varies by request, because a difficult case takes more steps than an easy one.
Two consequences follow. Model the economics with realistic volumes early, since a cost that is trivial in testing can be significant at production scale. And watch the interaction between quality and cost, because the natural fix for an accuracy problem is to add steps, verification, or retries, all of which increase per-request cost. That tradeoff should be a decision rather than a surprise.
There is also ongoing quality work, which is a permanent cost rather than a phase. Models change, APIs change, and input patterns drift, so an agent needs monitoring and periodic correction indefinitely.
What makes a project cost more
The main drivers are worth checking against your own plan.
Breadth of scope is the largest. An agent doing one narrow task is far cheaper than a general assistant, and not by a small margin, because scope drives evaluation difficulty, tool count, and the number of ways it can fail, all at once.
Number of integrations matters more than people expect. Each system the agent touches brings its own authentication, permissions, failure modes, and quirks.
The cost of being wrong sets the required rigour. An agent whose output a person reviews needs far less strengthening against failure than one acting alone on customer-facing decisions, and that difference shows up across the whole project.
Data readiness is a frequent hidden cost. If the information the agent needs is scattered, inconsistent, or hard to query, that has to be fixed first, and it is sometimes larger than the agent work itself.
A realistic timeline
Weeks for a prototype that establishes whether the task is feasible and produces a first evaluation set.
Months for a production system, covering evaluation, integration, permissions, hardening, and a limited release with a human in the loop.
Then ongoing, at a lower level, for monitoring, rare inputs, and periodic correction as the surrounding systems change.
How to keep it under control
Choose a narrow first task, because scope has the largest effect on cost.
Build the evaluation set early, since it is what prevents an unbounded improvement phase with no way to tell whether anything is getting better.
Release with a human in the loop and remove that step later when the data supports it, rather than trying to reach full autonomy before the first release.
Model the running cost before committing, using realistic volumes and a realistic step count.
How to judge a proposal from an outside team
If someone else is quoting for this work, the structure of the proposal tells you a lot before you look at the number.
Look for evaluation as a named, funded part of the plan rather than an afterthought. A proposal that goes from prototype straight to launch is either underestimating the work or intending to leave the hard part with you.
Look for specifics about integration. A team that has asked which systems the agent touches, how they authenticate, and how they fail is estimating from the real problem. A team that has not asked is estimating from the demo.
Look for a stated position on autonomy. Anyone experienced will raise what the agent may do without a person early, because it changes the architecture. Silence on that question usually means it has not been thought about.
Be careful with a fixed price for an open-ended agent scope. The many rare, unusual inputs are genuinely unpredictable, and a fixed price against them tends to produce either an inflated number or pressure to declare the work finished while failures remain.
Ask what happens after launch. An agent needs monitoring and periodic correction indefinitely, and a proposal that ends at delivery is describing half the commitment.
The projects that finish close to plan are almost always the ones that started narrow and expanded from something already working. The ones that overrun are almost always the ones that tried to build the impressive version first.
A worked example: two projects that looked similar and were not
Two teams each set out to build an agent that answers questions about a product catalog. On the surface the tasks looked comparable in size. The results were not.
The first team's catalog lived in one well-structured database with consistent fields, and the agent needed only to read it, never write. The prototype took a week, and the production system followed some weeks later, most of that time spent on evaluation and on handling questions that referenced products by names customers use rather than the exact names in the database. The cost of being wrong was low: an imprecise answer was mildly annoying rather than costly, so a person did not need to review every response before launch.
The second team's catalog data was spread across three internal systems, one of which used a different product identifier than the other two, and some records had not been updated in years. Before any agent work could start, the team had to build a reconciliation step just to know which records were current. The agent itself was no more complex than the first team's, but data readiness added weeks before the agent work could even begin, and the fact that the catalog fed into pricing meant a wrong answer had a real cost, so a review step stayed in place well past what the first team needed. Both projects were honestly described as "an agent that answers catalog questions." Their actual timelines differed by more than double, for reasons that had nothing to do with the model.
Distinguishing a fixed cost from a cost that scales with volume
Budgeting an agent project well means separating the two kinds of spending described earlier in this page, because they behave completely differently as the project grows.
Building the evaluation set, connecting the integrations, and hardening the system against failure are fixed costs: they are paid once, mostly before launch, and do not grow simply because usage grows later. Per-request model and tool costs are the opposite: they are paid continuously and grow directly with how much the agent is used. A project that is affordable at a hundred requests a day can become expensive at ten thousand, even though nothing about the system changed, because the second number is what the running cost actually multiplies against. Modeling both separately, rather than folding them into one estimate, avoids the common mistake of treating a cheap pilot as proof that the running cost at scale will also be cheap.
What changes the estimate mid-project, and how to plan for it
Even a careful initial estimate will move once real work starts, and naming the specific things that move it in advance makes the change less disruptive when it happens.
Discovering that the data the agent needs is less consistent than assumed is the most common mid-project surprise, and it usually surfaces during the integration phase rather than during initial scoping, because nobody reads every record closely until the system actually has to act on each one. A second common shift is the evaluation set revealing that the task has more distinct sub-cases than anyone described at the start, which is a sign the scoping conversation undercounted the task's real variety rather than a sign the team is doing something wrong. Budgeting a contingency against these two specific risks, rather than a generic buffer against unnamed uncertainty, gives a team a concrete reason to revisit the plan rather than simply absorbing the overrun silently.
Comparing an agent's cost to the cost of the alternative
A number that is easy to lose sight of once evaluation, integration, and hardening are all on the table is the cost of the current process the agent would replace, whether that is a person's time, a slower manual workflow, or the cost of the problem going unaddressed at all.
An honest comparison is not "how much does the agent cost" in isolation, it is "how does the agent's total cost, including its ongoing running cost, compare to what the task costs today." A narrow agent that takes months to build and carries a real running cost can still be the right decision if the manual alternative costs more every month it continues, and a broad, ambitious agent can be the wrong decision even at a lower sticker price if the task it replaces was cheap to begin with. Making this comparison explicit, rather than evaluating the agent's cost on its own, is what turns a budgeting exercise into an actual business decision.
Best for
- A first agent scoped narrowly enough that correctness can be defined and measured
- Projects where a human review step is acceptable at launch and removed later on evidence
- Teams that can model per-request running cost with realistic volumes before committing
Avoid if
- The data the agent needs is scattered or inconsistent, since fixing that comes first
- The plan assumes the production system is a short continuation of the prototype
Check before you decide
- Confirm the budget treats evaluation as a permanent line item, not a phase
- Confirm per-request cost has been modelled at production volume, not demo volume
- Confirm every system the agent integrates with has been listed with its failure modes
Common questions
How much does it cost to build an AI agent?
The cost is set by scope, integrations, and the cost of being wrong, so a single quoted figure for 'an AI agent' should be treated with suspicion. What is predictable is that the prototype is the cheap part, while evaluation, integration, permissions, and the many rare, unusual inputs are where the real effort goes.
Why is the production system so much slower than the prototype?
Because they answer different questions. A prototype has to succeed a few times on chosen inputs. Production has to behave acceptably on inputs nobody chose, with a defined response when it fails and a decided boundary on what it may do alone. That work is largely absent from the prototype.
Do AI agents cost more to run than normal software?
They scale differently. A multi-step agent makes several model calls per request, so per-request cost rises with usage rather than staying flat, and harder requests cost more than easy ones. Model this at realistic volumes early, and watch it when adding verification steps to improve accuracy.
What has the largest effect on agent project cost?
Scope has the largest effect on agent project cost. A narrow agent doing one well-defined task costs less to build than a general assistant, because scope sets how hard the agent is to evaluate, how many tools it needs, how many systems it connects to, and how many ways it can fail in production.
What is the difference between the prototype cost and the production cost of an AI agent?
The prototype cost covers getting an agent to work impressively a handful of times, which is often a matter of days. The production cost covers evaluation, integration with real systems, strengthening against failure, and handling the many rare, unusual inputs, none of which are needed for a demo but all of which are needed before real users can depend on the agent. The production system is usually the larger share of the work.
Is a fixed price quote safe for an open-ended AI agent project?
Be cautious with one. The many rare, unusual inputs an agent will meet in production are genuinely unpredictable, and a fixed price against that uncertainty tends to produce either an inflated number to cover the risk or pressure to call the work finished while real failures remain. A proposal that names evaluation and rare inputs explicitly is a better sign than a single flat number.
How do you know if a vendor proposal for an AI agent is realistic?
Check whether it funds evaluation as a named part of the plan rather than an afterthought, asks specific questions about which systems the agent integrates with and how they fail, and states a position on what the agent may do without a person before the build starts. A proposal that goes straight from prototype to launch, with no mention of these, is estimating from the demo rather than the real problem.
Does an AI agent cost anything after it launches?
Yes, ongoing monitoring and correction are a permanent line item, not a one-time cost. Models change, the APIs an agent calls change their response formats, and the mix of real user requests drifts over time, so an agent needs its evaluation set run on a schedule and periodic fixes indefinitely. A plan that ends at launch describes only part of the real commitment.
Related reading
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.
How to get a working AI prototype in weeks, not quarters
Most AI ideas end during the planning stage. Here is how to show something real to users fast enough to know if the idea is worth the full build.