AI evaluation and guardrails for production: how to know your AI actually works

A demo does not need evaluation. It runs once, in front of a person who picked an input that was likely to work, and everyone watching agrees it looked great. A production system runs thousands of times a day in front of people who did not pick their input carefully at all, and it needs two things a demo never asked for: evaluation, which is running the system against real cases with known good answers and scoring it, and guardrails, meaning the safety checks that stop a bad output from reaching a user or a bad action from happening. Together they are what separates a system you hope is working from a system you know is working.
We bring this up on nearly every AI engagement, usually before we talk about which model to use. That order is deliberate.
Why "it seems to work" is not a standard
Every AI project starts the same way. Someone writes a prompt, tries it on a handful of examples, and it looks good. That feeling is real, and it is also close to useless as a way of running a product, because it is a judgment about a handful of inputs someone chose on purpose. The inputs a real user sends were not chosen on purpose. They are messy, half-finished, sarcastic, in a language nobody tested, or just weird in a way nobody thought to try.
The fix is not more careful manual testing. It is a number. An evaluation set is a collection of real inputs paired with a correct or acceptable answer for each one, scored automatically so you can run it in minutes instead of an afternoon of reading results by hand. Once that number exists, a question that used to be a feeling becomes a fact you can check: did the new prompt help or hurt? Did switching models change accuracy on the ten hardest cases you have? Is Tuesday's version better than Monday's?
We saw this happen clearly on a retrieval feature we worked on. The demo answer looked good. The evaluation set, built from actual questions a support team had received that quarter, showed the system getting roughly a third of the harder multi-part questions wrong, confidently, with no signal to the user that anything was wrong. Nobody could see that from the demo. Everybody could see it from the score.
What makes this harder for AI than for normal software
Ordinary software is deterministic. The same input produces the same output today, next month, and next year, unless someone changes the code. That property is what makes a passing test suite trustworthy: pass it once against a stable system and you have real evidence.
AI systems keep changing. Three things change at the same time, without any action from you.
The same input can produce a different output. Most model providers run with some randomness by design, so asking the same question twice can return two different phrasings, or in unusual cases, two different answers. A single manual test tells you what happened once, not what usually happens.
Providers update models without asking you first. A model you built against gets a version update with no notice, a safety tuning change, or a full replacement, and your prompt now runs against different underlying behavior. Nothing in your codebase changed. Your system's behavior did anyway.
A working prompt gets worse without warning. As usage grows, the inputs reaching your system become different from what you originally tested against. A prompt tuned for the first hundred users can be noticeably weaker by the time real usage looks like the ten-thousandth user, and nothing announces that moment. You only see it if you are measuring.
Put together, this means the question is never "does it work" as a one-time fact. It is "is it still working," asked continuously. That is a different discipline than releasing a feature once and moving to the next task, and it is the discipline evaluation and safety checks exist to support.
Evaluation, in practice
An evaluation set starts small and grows from real use, not from imagination. Pull actual user inputs, not hypothetical ones, and pair each with the answer a domain expert on your side agrees is correct or acceptable. Fifty to a hundred cases is a reasonable starting size for most features. It should include the common cases people ask every day and the known hard cases: the ones with ambiguous phrasing, missing context, or wording that is easy to misread.
Scoring depends on the type of answer. Some tasks have one clearly right answer, like extracting a date or a dollar amount from a document, and you can check those with exact string or value matching. Open-ended answers, like a drafted email or a support response, do not have one right string to match against, so teams often use a second model as a judge, scoring the first model's output against a rubric. This works, but the judge can also change over time: the judge model needs its own periodic check against real human judgment, because a judge that has started passing bad answers without anyone noticing is worse than no judge, since it gives you false confidence.
The real benefit of an evaluation set appears on the day you want to change something. Swap a model, tune a prompt, adjust how much context retrieval hands the model: run the evaluation set before and after, and you have a real answer to "did this help," rather than a guess based on the three examples someone happened to try after the change. Without that number, every change is also a risk: it may make things worse, and you will not notice.
Safety checks, in practice
Evaluation tells you how the system is doing on average. Safety checks decide what happens on the specific request where it is not doing well, and they run in production, on every request, as well as in a test run.
Checking answers against source data. For any system pulling from your own documents or records, the output should be checked against what it retrieved before it reaches the user. If the model's claim is not supported by the source it was given, that is a signal to hold the answer back, flag it, or ask a follow-up rather than show a confident sentence with no source behind it.
Blocking unsafe or off-topic output. A content filter sitting between the model and the user catches the responses that should never go out: unsafe instructions, off-brand claims, answers to questions the product was never meant to answer. This is a narrow, boring check, and it is one of the cheapest safety checks to run relative to the cost of a single model call.
Validating structured output. When a model is asked to return data your code will act on, like a JSON object that updates a database or starts an action in another system, that output needs to be validated against a schema before anything happens with it. A model that returns a malformed field, or a value outside the range your system expects, should trigger a rejection and a retry, not a write straight into your database.
Limiting which tools an agent can call. An agent that can only draft a message is a very different risk than one that can send it, and one that can send a message is a very different risk than one that can move money. The safety check here is a limit on what the agent can do: give the agent the narrowest set of tools the task actually requires, and require a human confirmation step before any action that is expensive to undo. We explain in more detail how to decide where that limit sits in can you trust an AI agent with real work yet, where the short version is that the question is never whether agents in general are ready, it is whether this specific action, on this specific task, is safe to give to an agent.
The design principle behind all four is the same. Make the safe, reversible cases fast, and make the risky, hard-to-undo cases slow on purpose. A safety check that treats every action the same way is either too strict to be useful or too loose to be safe. The value comes from how it separates the two.
Building this in from the start, not after an incident
Most teams that end up with real evaluation and safety checks did not plan it that way from the start. They built it after something went wrong: a wrong answer reached a customer, an agent took an action nobody meant to allow, and the fix that followed was more expensive than building the discipline in up front would have been. Adding a safety check after an incident means writing it under pressure, with a customer already affected, and usually discovering three more problems while fixing the first one.
The alternative costs very little extra at the start. Build the evaluation set alongside the first working prototype, using the same real cases you are already collecting to judge whether the prototype is any good. Add the safety checks as part of the feature, not as a separate step added the week before launch. This is the order we default to on AI engagements: define the problem and how you will measure a right answer, build a narrow version with the safety checks already in place, then widen scope once the evaluation set says it deserves that trust. It is also the pattern that separates AI projects which reach production from those that do not: the teams that release are rarely the ones with the best model, they are the ones honest about how their system fails before it fails on a real user.
How Reveneau approaches this
We treat evaluation and safety checks as part of the build, not an add-on we sell separately after something breaks. Our AI development work starts with the failure modes and the way we will measure quality, before we spend time on model selection, because a strong model behind a weak measurement system still produces a product nobody can trust. If you have an AI feature that works in a demo and you are not sure it is ready for real traffic, that gap between the demo and a measured system with safety checks is exactly where we spend our time.
Thanks to the teams who let us study their evaluation failures honestly enough to learn from them. The model is rarely the reason an AI project struggles in production. The measurement is.
Related guides: LLM evals: how to measure whether an AI product works, why AI evals matter, Taking AI agents from prototype to production, and for the same discipline applied to the code itself rather than to a model's answers, eval-driven development.
The same discipline runs on our own work, which is a fair thing to check us on. We generate one hundred percent of our code, so we are exactly the sort of team that could release a large amount of confident, plausible, unverified output. The evaluation habit is what stops that: a named engineer reads every change against the specification it came from before it reaches your branch. A team that sells you safety checks and does not run any internally is worth asking about.


