Decide with evals

Using evals to decide whether an AI feature is ready to release

An AI feature is ready to release when it meets a pass line that the company wrote down before the test was run. Set one line for the whole test set, a higher line for outputs that can cost money or cause harm, and a short list of failures that block release whatever the average is. The company makes the decision and records it, and the people who run the tests report the result without owning the decision. Every failure found after release becomes a permanent test case. The EU AI Act describes testing against thresholds defined in advance, NIST's Generative AI Profile describes minimum thresholds reviewed when deployment is approved, and both leave the number to you.

Published September 30, 2026. Editorial.

Key takeaways

  • Write the pass line, with a date, before the test is run. A line chosen after the result is seen will be set below the result, so it decides nothing.
  • Set a separate, higher line for outputs that state a price, a policy or a law, and for outputs that take an action. The overall average can pass while that group fails.
  • List the failures that block release on a single occurrence, such as an invented policy or one customer's data shown to another, and write a test case for each one.
  • The company owns the release decision and records the result, the lines, the date and the approving role. NIST's framework asks for documented roles and for assessors who did not build the system.
  • A failure found after release is added to the test set and stays there, so the same failure is checked on every later change.

Take an invented furniture shop that has built a support assistant. On Thursday the team runs its test set of 400 cases and 372 pass, which is 372 divided by 400, or 93 percent. The release is planned for Friday. If nobody wrote down before Thursday what result would be enough, the Friday meeting will decide that 93 percent is enough, because the release is already planned and the number sounds high.

This page describes how to make that decision so that the result can change it. It assumes you already have an eval, which is a repeatable test of an AI product, explained in what an AI eval is. Every number in the furniture shop example is an invented illustration.

Write the pass line before the test is run

A pass line is the lowest result at which you will release. It has to exist in writing, with a date, before the test is run. A line written afterwards is set by people who already know the result and already want the release, so it will be set below the result.

Two public texts describe this order. The EU AI Act says, in Article 9(8), that for the systems it calls high-risk, "Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system" [3]. NIST, the US government standards body, published a voluntary Generative AI Profile whose action GV-1.3-002 reads: "Establish minimum thresholds for performance or assurance criteria and review as part of deployment approval" [2]. The scope and the dates of both texts are in what regulators and standards bodies say about testing AI.

Both texts leave the number to you. NIST says of its own AI Risk Management Framework that "it does not prescribe risk tolerance" [1]. The furniture shop has to work out what a wrong answer costs it and write the line from that.

A pass line states three things

A pass line that someone else can check has three parts.

  1. The set it is measured on. Name the test set and its size. A pass rate on 40 cases is weaker evidence than the same rate on 400, and how to read an AI eval report shows the arithmetic.
  2. The number. For example, "at least 360 of the 400 cases pass", which is 90 percent.
  3. What counts as a pass for one case. A language model can give a different output each time it receives the same input. Anthropic's engineering guide gives this example for an agent, which is an AI system that works in several steps and takes actions: an agent that succeeds on 75 percent of single attempts passes three attempts out of three only 42 percent of the time, because 0.75 x 0.75 x 0.75 is 0.421875 [4]. So the line must say whether a case passes when one attempt succeeds or only when every attempt succeeds. For anything a customer relies on, we use the second meaning. The method is in agent reliability across repeated runs.

Set a higher line where a wrong output costs money or causes harm

One line for the whole set treats every wrong answer as equal, and wrong answers differ in cost. A reply with an awkward sentence is a small annoyance. A reply that states a refund rule your company does not have can cost you the refund. In Moffatt v. Air Canada, decided on February 14, 2024, a British Columbia tribunal ordered the airline to pay $812.02 after its website chatbot told a customer they could apply for a bereavement fare after travelling. The airline's own policy page said the opposite. The tribunal found that "Air Canada did not take reasonable care to ensure its chatbot was accurate" [8]. The full record is in public AI failures and the tests that target each one.

So sort the outputs by what a wrong one costs, and give each group its own test cases and its own line. The table shows one way to do it. The percentages are examples made up for the invented furniture shop. Set your own from what a wrong answer costs your business.

Type of output Example Example pass line, with the reason What blocks release
General product answers "Does the oak table come in a 180 cm length?" 90 percent. A wrong answer causes a return or a lost sale. A result below the line
Statements of price, policy or law "Can I get a refund after 30 days?" 95 percent, measured on this group alone. The company may have to honour what the assistant said. A result below the line, or one invented policy
Actions that move money or data Issuing a refund, changing an address Every case, on every repeated attempt. A wrong action has already happened by the time a person sees it. One failure
Things that must never happen Showing one customer another customer's order Zero occurrences across the whole set. One occurrence

Now apply the table to the furniture shop's Thursday result. Suppose 40 of the 400 cases are refund and delivery policy questions, and 9 of the 28 failures are among them. The policy group passes 31 of 40, which is 77.5 percent. The example line of 95 percent needs 38 of 40. The other 360 cases pass 341, which is 94.7 percent. The overall 93 percent is above a 90 percent line, and the release still stops, because the group where a wrong answer costs money is below its own line.

Name the failures that block release whatever the average is

Some failures should stop the release on one occurrence. List them, and write at least one test case for each item. A list for a support assistant might read: states a policy the company does not have, shows one customer's data to another, gives advice that breaks a law, takes an action without the confirmation step.

Each item needs a count limit in place of a percentage. Anthropic's documentation gives an example of a criterion written this way: "Less than 0.1% of outputs out of 10,000 trials flagged for toxicity by the content filter" [5]. A content filter is a program that marks harmful text. Copy the form: a named failure, a number of attempts, and a limit.

What a release gate is

A release gate is a check that a change must pass before it is released. In practice it is a rule in the release process: the eval set runs, the results are compared with the written lines, and a result below any line stops the change automatically, before anyone is asked for an opinion.

The gate should run on every change to the prompt (the written instructions given to the model), to the model version, and to the documents the product reads. OpenAI's documentation advises teams to "run evals on every change" and to "grow the eval set over time" [6].

A working gate uses two sets. The first holds the cases for new behaviour, where the line is the one you wrote for this release. The second holds everything that already worked. A regression is a behaviour that worked before a change and stopped working after it, and Anthropic's guide says the tests that check for regressions should pass almost every time [4]. A drop in that second set blocks release even when the new behaviour passes.

Who decides the release

The company decides. A role holds the authority, such as the product owner or a release committee, and the decision is recorded: the result, the lines, the date, and the role that approved. NIST's framework describes the step in its MANAGE 1.1 outcome: "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed" [1]. Its GOVERN 2.1 outcome asks that roles and responsibilities for measuring and managing AI risks "are documented and are clear to individuals and teams throughout the organization" [1].

Keep the people who test apart from the decision. NIST's MEASURE 1.3 outcome says regular assessments should involve internal experts who did not build the system, or independent assessors [1]. The UK government body now called the AI Security Institute, which tests the most advanced models, wrote the same separation into its own approach in February 2024: "AISI is ultimately not responsible for release decisions made by the parties whose systems we evaluate" [7]. A tester reports what the test found, and the company that releases the product is responsible for the release.

When a partner builds the product for you, ask who approves a release on the partner's side, and whether that approval belongs to the company or to one engineer. Budget and ownership are covered in what AI evals cost and who should own them.

What to do when the result is below the line

A result below the line leaves four responses.

  • Fix and re-run. Change the prompt, the documents or the model, then run the whole set again, including the set of behaviours that already worked.
  • Narrow the feature. Remove the type of request that failed. The furniture shop could release the assistant for product questions and send every policy question to a person.
  • Release to a small group first. NIST's Generative AI Profile lists a staged release as an approach to consider, in its action MG-1.3-001 [2]. It limits how many customers a failure can reach.
  • Stop. The same profile asks organisations, in action GV-1.3-007, to "Devise a plan to halt development or deployment of a GAI system that poses unacceptable negative risk" [2]. GAI is NIST's short form for generative AI. Write that plan before you need it.

Changing the line is a fifth possibility, and it needs its own rule: the change is written down with the reason, the date and the approving role, before the next run.

A failure found after release becomes a permanent test case

No test set contains every input a real customer will send, so some failures will be found after release. The response has a fixed order.

  1. Limit the harm: turn the feature off for that type of request, or send those requests to a person.
  2. Write the failing input into the test set as a new case, with the expected behaviour.
  3. Confirm that the new case fails on the current version. A case that passes is testing something else.
  4. Fix the product, run the full set, and release through the same gate.
  5. Keep the case. It runs on every later change.

NIST's MANAGE 4.1 outcome describes the plan this belongs to: monitoring after deployment that includes "mechanisms for capturing and evaluating input from users", along with incident response and change management [1]. Anthropic's guide makes a related point about where useful cases come from. It says that 20 to 50 simple tasks "drawn from real failures" is a good start, and it gives no measurement behind that range [4].

Our post on evaluation and guardrails for production covers guardrails, which are checks that stop a bad output before it reaches a customer. The guide to turning production traces into eval cases covers the method, where a trace is the recorded steps of one real use.

How Reveneau applies this

Reveneau is an AI software development consultancy, and all of our code is written by AI. Every change has to pass a large eval suite before release, and that suite is written from the specification before the code exists. For an AI feature, this means the pass lines on this page are written at the same time as the specification: the overall line, the higher lines for outputs that state a price or a policy or take an action, and the list of failures that block release on one occurrence. We ask every client for that list before the first prompt is written.

The checks that need judgment are graded by Jev, TypeSafe AI's decision model. On our own suite that run is ten times faster than with our previous language-model grader, and a faster run is one reason we can apply the gate to every change. Reveneau, as a company, takes responsibility for the whole project through production and after release, so a failure found after release becomes a permanent case in the suite.

If you want release decisions that a test result can change, see AI development at Reveneau or contact us. The wider argument for testing AI products is in why AI evals matter.

Best for

  • Teams about to release an AI feature that states prices, policies or facts to customers
  • Buyers who want a written release condition in a contract with a development partner
  • Products that change often, where every change needs the same check

Avoid if

  • You have no test set yet: build one first, then write the lines
  • The feature is an internal experiment that shows nothing to customers and takes no action on real data

Check before you decide

  • The pass lines have a date earlier than the test run
  • The price, policy and action groups have their own cases and their own result
  • The record names the role that approved the release and the result it approved
  • Every failure found after release appears as a case in the current set

Common questions

How do I know an AI feature is ready to release?

An AI feature is ready to release when its eval result meets every pass line that was written down before the test was run. That means the overall line, the higher line for each group of costly outputs, and zero occurrences of the failures you listed as blocking. NIST's Generative AI Profile describes this as setting minimum thresholds and reviewing them as part of deployment approval.

What pass rate should block a release?

The pass rate that should block a release comes from what a wrong output costs your business. No public standard gives one, and NIST says its framework does not prescribe risk tolerance. On this page the invented furniture shop uses 90 percent overall and 95 percent for policy answers, and those two figures are examples for the illustration only.

Who decides whether an AI feature is released?

The company decides whether an AI feature is released, through a role it has given that authority, and it records the decision. The people who run the tests report the result. NIST's MEASURE 1.3 outcome asks for assessors who did not build the system, and the UK's AI Security Institute states that it is not responsible for the release decisions of the companies whose systems it tests.

What is a release gate?

A release gate is a check that a change must pass before it is released. For an AI product, the gate runs the eval set, compares the results with the written pass lines, and stops the change automatically when any result is below its line. OpenAI's documentation advises running evals on every change, which is what a gate does.

What do I do with a failure found in production?

Turn a failure found in production into a permanent test case. First limit the harm, then write the failing input and the expected behaviour into the test set, confirm the new case fails on the current version, fix the product, and re-run the full set. Anthropic's guide says 20 to 50 simple tasks drawn from real failures is a good start for a set.

Why must the pass line be written before the test is run?

The pass line must be written before the test because people who already know the result, and already want the release, will set the line below it. The EU AI Act uses the same order for high-risk systems: Article 9(8) says testing is carried out against prior defined metrics and probabilistic thresholds. A line with a date earlier than the run can be checked by someone outside the team.

Should every type of AI output have the same pass line?

Different types of output should have different pass lines, because wrong outputs differ in cost. An awkward sentence annoys a customer, while a wrong statement of policy can cost money. In Moffatt v. Air Canada, a tribunal ordered the airline to pay $812.02 after its chatbot misstated a fare policy. Give costly groups their own test cases and a higher line.

What should we do if the eval result is below the line and the release date is fixed?

When the eval result is below the line, fix and re-run, narrow the feature to the requests that passed, release to a small group first, or stop. NIST's Generative AI Profile lists a staged release as an approach to consider and asks for a plan to halt. If the line itself is changed, record the reason, the date and the approving role before the next run.

What happens if someone wants to lower the pass line?

Lowering the pass line is allowed only as a recorded change made before the next test run. Write down the new line, the reason, the date and the role that approved it. A line that is lowered on release day, after the result is known, no longer decides anything. NIST's Generative AI Profile describes approval thresholds as something reviewed through policy and process.

Does one successful attempt mean a test case passed?

One successful attempt is weak evidence that a test case passed, because a language model can answer differently each time. Anthropic's guide gives the arithmetic: an agent that succeeds on 75 percent of single attempts passes three out of three only 42 percent of the time. Your pass line should say how many attempts are run and how many must succeed.

Do these release rules apply to a small internal AI tool?

The same release rules apply to a small internal AI tool in a shorter form. Write one pass line and a short list of blocking failures, such as exposing staff data. A tool that shows nothing to customers and takes no action on real data can use a short set. The step that stays the same at any size is writing the line before the test.

References