AI

AI evaluation and guardrails for production: how to know your AI actually works

Editorial · Reveneau · July 28, 2026

AI evaluation and guardrails for production: how to know your AI actually works

A demo does not need evaluation. It runs once, in front of a person who picked an input that was likely to work, and everyone watching agrees it looked great. A production system runs thousands of times a day in front of people who did not pick their input carefully at all, and it needs two things a demo never asked for: evaluation, which is running the system against real cases with known good answers and scoring it, and guardrails, meaning the safety checks that stop a bad output from reaching a user or a bad action from happening. Together they are what separates a system you hope is working from a system you know is working.

We bring this up on nearly every AI engagement, usually before we talk about which model to use. That order is deliberate.

Why "it seems to work" is not a standard

Every AI project starts the same way. Someone writes a prompt, tries it on a handful of examples, and it looks good. That feeling is real, and it is also close to useless as a way of running a product, because it is a judgment about a handful of inputs someone chose on purpose. The inputs a real user sends were not chosen on purpose. They are messy, half-finished, sarcastic, in a language nobody tested, or just weird in a way nobody thought to try.

The fix is not more careful manual testing. It is a number. An evaluation set is a collection of real inputs paired with a correct or acceptable answer for each one, scored automatically so you can run it in minutes instead of an afternoon of reading results by hand. Once that number exists, a question that used to be a feeling becomes a fact you can check: did the new prompt help or hurt? Did switching models change accuracy on the ten hardest cases you have? Is Tuesday's version better than Monday's?

We saw this happen clearly on a retrieval feature we worked on. The demo answer looked good. The evaluation set, built from actual questions a support team had received that quarter, showed the system getting roughly a third of the harder multi-part questions wrong, confidently, with no signal to the user that anything was wrong. Nobody could see that from the demo. Everybody could see it from the score.

What makes this harder for AI than for normal software

Ordinary software is deterministic. The same input produces the same output today, next month, and next year, unless someone changes the code. That property is what makes a passing test suite trustworthy: pass it once against a stable system and you have real evidence.

AI systems keep changing. Three things change at the same time, without any action from you.

The same input can produce a different output. Most model providers run with some randomness by design, so asking the same question twice can return two different phrasings, or in unusual cases, two different answers. A single manual test tells you what happened once, not what usually happens.

Providers update models without asking you first. A model you built against gets a version update with no notice, a safety tuning change, or a full replacement, and your prompt now runs against different underlying behavior. Nothing in your codebase changed. Your system's behavior did anyway.

A working prompt gets worse without warning. As usage grows, the inputs reaching your system become different from what you originally tested against. A prompt tuned for the first hundred users can be noticeably weaker by the time real usage looks like the ten-thousandth user, and nothing announces that moment. You only see it if you are measuring.

Put together, this means the question is never "does it work" as a one-time fact. It is "is it still working," asked continuously. That is a different discipline than releasing a feature once and moving to the next task, and it is the discipline evaluation and safety checks exist to support.

Evaluation, in practice

An evaluation set starts small and grows from real use, not from imagination. Pull actual user inputs, not hypothetical ones, and pair each with the answer a domain expert on your side agrees is correct or acceptable. Fifty to a hundred cases is a reasonable starting size for most features. It should include the common cases people ask every day and the known hard cases: the ones with ambiguous phrasing, missing context, or wording that is easy to misread.

Scoring depends on the type of answer. Some tasks have one clearly right answer, like extracting a date or a dollar amount from a document, and you can check those with exact string or value matching. Open-ended answers, like a drafted email or a support response, do not have one right string to match against, so teams often use a second model as a judge, scoring the first model's output against a rubric. This works, but the judge can also change over time: the judge model needs its own periodic check against real human judgment, because a judge that has started passing bad answers without anyone noticing is worse than no judge, since it gives you false confidence.

The real benefit of an evaluation set appears on the day you want to change something. Swap a model, tune a prompt, adjust how much context retrieval hands the model: run the evaluation set before and after, and you have a real answer to "did this help," rather than a guess based on the three examples someone happened to try after the change. Without that number, every change is also a risk: it may make things worse, and you will not notice.

Safety checks, in practice

Evaluation tells you how the system is doing on average. Safety checks decide what happens on the specific request where it is not doing well, and they run in production, on every request, as well as in a test run.

Checking answers against source data. For any system pulling from your own documents or records, the output should be checked against what it retrieved before it reaches the user. If the model's claim is not supported by the source it was given, that is a signal to hold the answer back, flag it, or ask a follow-up rather than show a confident sentence with no source behind it.

Blocking unsafe or off-topic output. A content filter sitting between the model and the user catches the responses that should never go out: unsafe instructions, off-brand claims, answers to questions the product was never meant to answer. This is a narrow, boring check, and it is one of the cheapest safety checks to run relative to the cost of a single model call.

Validating structured output. When a model is asked to return data your code will act on, like a JSON object that updates a database or starts an action in another system, that output needs to be validated against a schema before anything happens with it. A model that returns a malformed field, or a value outside the range your system expects, should trigger a rejection and a retry, not a write straight into your database.

Limiting which tools an agent can call. An agent that can only draft a message is a very different risk than one that can send it, and one that can send a message is a very different risk than one that can move money. The safety check here is a limit on what the agent can do: give the agent the narrowest set of tools the task actually requires, and require a human confirmation step before any action that is expensive to undo. We explain in more detail how to decide where that limit sits in can you trust an AI agent with real work yet, where the short version is that the question is never whether agents in general are ready, it is whether this specific action, on this specific task, is safe to give to an agent.

The design principle behind all four is the same. Make the safe, reversible cases fast, and make the risky, hard-to-undo cases slow on purpose. A safety check that treats every action the same way is either too strict to be useful or too loose to be safe. The value comes from how it separates the two.

Building this in from the start, not after an incident

Most teams that end up with real evaluation and safety checks did not plan it that way from the start. They built it after something went wrong: a wrong answer reached a customer, an agent took an action nobody meant to allow, and the fix that followed was more expensive than building the discipline in up front would have been. Adding a safety check after an incident means writing it under pressure, with a customer already affected, and usually discovering three more problems while fixing the first one.

The alternative costs very little extra at the start. Build the evaluation set alongside the first working prototype, using the same real cases you are already collecting to judge whether the prototype is any good. Add the safety checks as part of the feature, not as a separate step added the week before launch. This is the order we default to on AI engagements: define the problem and how you will measure a right answer, build a narrow version with the safety checks already in place, then widen scope once the evaluation set says it deserves that trust. It is also the pattern that separates AI projects which reach production from those that do not: the teams that release are rarely the ones with the best model, they are the ones honest about how their system fails before it fails on a real user.

How Reveneau approaches this

We treat evaluation and safety checks as part of the build, not an add-on we sell separately after something breaks. Our AI development work starts with the failure modes and the way we will measure quality, before we spend time on model selection, because a strong model behind a weak measurement system still produces a product nobody can trust. If you have an AI feature that works in a demo and you are not sure it is ready for real traffic, that gap between the demo and a measured system with safety checks is exactly where we spend our time.

Thanks to the teams who let us study their evaluation failures honestly enough to learn from them. The model is rarely the reason an AI project struggles in production. The measurement is.

Related guides: LLM evals: how to measure whether an AI product works, why AI evals matter, Taking AI agents from prototype to production, and for the same discipline applied to the code itself rather than to a model's answers, eval-driven development.

The same discipline runs on our own work, which is a fair thing to check us on. We generate one hundred percent of our code, so we are exactly the sort of team that could release a large amount of confident, plausible, unverified output. The evaluation habit is what stops that: a named engineer reads every change against the specification it came from before it reaches your branch. A team that sells you safety checks and does not run any internally is worth asking about.

Common questions

What is AI evaluation?

AI evaluation is running your AI system against a set of real cases with known good answers and scoring the results, so "it seems to work" turns into a number you can track from one release to the next. A good evaluation set starts small, often 50 to 100 cases covering common inputs and known difficult cases, and grows every time a user finds something the model gets wrong.

What are AI guardrails?

Guardrails are the checks that constrain what an AI system can output or do: checking answers against your source data, blocking unsafe or off-topic output, validating structured responses before they reach your database, and limiting which tools an agent is allowed to call. They run in production, on every real request, and their job is to make the safe, reversible actions fast while making the risky, hard-to-undo ones slow on purpose.

Why does AI need evaluation and safety checks when normal software does not need them the same way?

Normal software gives the same output for the same input every time, so a passing test stays passing. AI model behaviour changes over time, providers update models without warning, and the same prompt can return a different answer today than it did yesterday, so a system that passed evaluation last month can quietly get worse without any code change on your side.

When should a team build evaluation and safety checks into an AI project?

A team should build evaluation and safety checks from the first working prototype, not after launch. Building an evaluation set alongside the first version costs a few days. Building one after an incident costs the incident plus the fix plus the trust you lost, and it usually happens under far more pressure, with a customer already affected and three more problems appearing while the team fixes the first one.

How do you build an evaluation set for an AI feature?

Start by collecting real cases, actual user inputs and the correct answer for each, not invented examples. A working evaluation set can start small, often 50 to 100 cases covering common inputs and known difficult cases, and it should grow every time a user finds a case the model gets wrong.

What is the difference between an evaluation set and a safety check?

An evaluation set measures quality before and after a change, run offline against known cases, so you can tell if a new prompt or model made things better or worse. A safety check runs in production, on every real request, and stops a bad output from reaching the user in the first place. You need both: one tells you if the system is good, the other stops it from doing damage on the days it is not.

Can you use an LLM to evaluate another LLM's output?

Yes, this is called LLM-as-judge, and it is common for open-ended answers that do not have one exact correct string. It needs its own validation: check the judge model's scores against human judgment on a sample regularly, because a judge model can change over time or develop its own gaps in judgment just like the model it is grading.

What should a safety check do when an AI agent wants to take a risky action?

It should stop and require a human to confirm before the action happens, especially when the action is hard to undo, such as sending money, deleting a record, or emailing a customer. The safety check's job is to make the easy, reversible actions fast and the hard, irreversible ones slow on purpose.

How often should an AI system be re-evaluated once it is in production?

An AI system should be re-evaluated continuously, not once at launch. Any time you change a prompt, swap a model, adjust retrieval, or a provider releases a model update, you should re-run the evaluation set before and after, because the alternative is finding out from a user that quality dropped. Models change over time and providers update them without warning, so a system that passed evaluation last month can quietly get worse with no code change on your side.

What happens if a team skips evaluation and safety checks to release faster?

The system usually still works until the day it fails, at which point the failure shows up as a wrong answer already in front of a customer instead of a score in a dashboard. Adding evaluation and safety checks after an incident takes longer than building them in from the start, because now the team is also rebuilding trust with whoever got the bad answer.

Does adding safety checks slow an AI feature down?

A safety check that checks output before it returns to the user adds a small amount of latency, usually not noticeable next to the model call itself, and the checks that matter most, like validating structured output or checking a claim against source data, are cheap compared to a single model call.

How does Reveneau build evaluation and safety checks into AI projects?

We build the evaluation set alongside the first working version of any AI feature, not after it, and we treat safety checks as part of the feature rather than a separate reliability step added before launch. See our AI development work for how this fits into the wider build.