AI agent in production
The demo works. The real project is making it keep working with real users.
An agent that answers questions in a test environment takes a weekend to build. An agent that handles real customer traffic needs retrieval that stays current, evaluation you can run on every change, safety checks that stop a bad answer before it is sent, and a backup response for when the model is wrong. Most agent projects get stuck on that last 20%, so we plan it first rather than last.
Our approach
We start by writing the evaluation set, before any agent code exists. It forces the hard conversation early: what counts as a correct answer here, and who decides. From there the agent is built against that set, so every change is measured instead of judged by opinion. By default, a person reviews each answer before it is sent, and we remove that step only once the numbers show it is safe. Timeline we plan for: 6 to 10 weeks to a version handling live traffic.
Common questions
Why does an AI agent demo well but fail with real customers?
A demo covers the questions the person building it already expects. Real users ask questions nobody planned for, and a demo has no evaluation set behind it to catch a wrong answer before it reaches a customer. The difference between the two is safety checks, a backup response, and measurement. Those only get built once someone treats real customer use as the actual requirement.
What is an evaluation set and why write it before the agent?
It is a set of real questions with agreed correct answers, written before any agent code exists. Writing it first forces the team to settle what counts as a correct answer and who decides, which is normally the argument that teams skip, and that then comes back after launch as a support complaint. Every later change to the agent is measured against that same set instead of judged by whether it feels right.
Does the agent send answers directly to customers from day one?
No. At launch, by default, a person checks each answer before it reaches a customer. We remove that review step only once the evaluation set shows the agent's answers are correct without it, which is a measured decision rather than a launch-date deadline.
What happens when the model gives a wrong answer in production?
We plan the backup response from the start, instead of adding it later after an incident. That means a wrong or low-confidence answer goes to a safe response or a human, rather than reaching the customer as if it were correct. Retrieval also has to stay current, since an agent working from out-of-date information will confidently state something that used to be true.
How long does it take to get an agent handling live traffic?
Six to ten weeks is the timeline we scope to, not a measured average, since this is engagement pattern guidance rather than a record of past projects. The factor that changes that number is how well-defined the correct answers already are: an agent replacing a documented process is built faster than one where the team is still arguing about what a right answer looks like.
What is the 'last 20%' that most agent projects get stuck on?
It is retrieval that stays current as source data changes, an evaluation set you can run on every code change, safety checks that stop a bad answer before it is sent, and a backup response for when the model gets it wrong. A working demo covers the other 80% in a weekend; those four pieces are what turn a demo into something that keeps working with real customer traffic.
Who decides what counts as a correct answer for the agent?
The people who are responsible for the result, rather than the engineers building the agent, and that conversation happens while the evaluation set is being written, before any agent code exists. Deferring it to after launch means the first real disagreement about a wrong answer happens in front of a customer instead of in a planning session.
Can this agent pattern work for an internal tool instead of a customer-facing one?
Yes, the same evaluation-first approach applies, and the safety checks and human review can usually be reduced sooner for internal use, since the cost of a wrong answer is normally lower than it is for a customer. The evaluation set still has to exist first: an internal agent with no measured standard for a correct answer fails in the same way, only for fewer people.