Start here

What is an AI eval? A plain-language explanation

An AI eval is a repeatable test of an AI product. It has three parts: an input, a written expectation of what a correct output must do, and a rule that scores the output as pass or fail. A language model can give a different answer to the same question on a different day, so each case is run several times, and the full set, called an eval suite, reports a pass rate. OpenAI's documentation says this variation makes traditional software testing methods insufficient for AI. People who know the business write the expectations, and engineers turn them into checks a computer can run.

Published September 30, 2026. Editorial.

Key takeaways

  • An AI eval is a repeatable test with three parts: an input, a written expectation of what a correct output must do, and a rule that scores the output as pass or fail.
  • A language model can give different output for the same input, so each case is run several times. Anthropic's engineering post calls each attempt a trial.
  • An eval suite is the full set of test cases for one product, and its pass rate is the share of cases or trials that passed.
  • People who know the business write the expectations, and engineers turn them into checks a computer can run. NIST asks that people who did not build the system take part in assessing it.
  • A pass rate describes only the cases in the suite. Ask how many cases the figure was measured on and which situations have no case.

Take a support assistant for an invented furniture shop. A customer writes: "I bought a sofa 40 days ago. Can I still return it?" The shop's written policy allows returns for 30 days. On Monday the assistant answers that the return period has ended and offers a repair visit. On Tuesday, asked the same question in the same words, it says returns are accepted for 90 days. Nobody changed the product between the two answers.

The Tuesday answer is the reason evals exist. OpenAI's documentation says: "Models sometimes produce different output from the same input, which makes traditional software testing methods insufficient for AI architectures" [1]. An AI eval is the test built for software that behaves this way. This is the first page of the guide on why AI evals matter.

An eval is a repeatable test with three parts

An AI eval (short for evaluation) is a repeatable test of an AI product. Each eval has three parts:

  1. An input. The question, document or request that the product receives, such as the customer's message about the sofa.
  2. A written expectation. A sentence or two, written before the test is run, stating what a correct output must do.
  3. A way to score the output. A rule that turns the output into a result, most often pass or fail.

The companies that sell AI models describe the same structure. OpenAI's documentation says that evals "test model outputs to ensure they meet style and content criteria that you specify" [2]. Anthropic's engineering team describes the simplest form in one sentence: "an agent processes a prompt, and a grader checks if the output matches expectations" [3]. A prompt is the text sent to the AI model, an agent is an AI program that takes steps on its own to finish a task, and a grader is whatever does the scoring.

The word "repeatable" matters most. The same inputs can be run again after any change to the product, and the two results can be compared case by case.

One eval case, written out

Here is one case for the furniture shop, with all three parts. The shop, the policy and the order are an invented illustration.

Part What it is In this example
Input The message the product receives "I bought a sofa 40 days ago. Can I still return it?"
Expected behaviour What a correct answer must do, written before the test States that the return period is 30 days and has ended. Promises no refund. Offers the repair service named in the policy.
How it is scored The rule that produces pass or fail Fail if the answer names any period other than 30 days, or promises a refund. Pass if it does all three expected things.

Copy two details from this table. The expectation describes behaviour instead of exact wording, because a correct answer can be phrased in many ways. The scoring rule names what counts as a failure, so two people scoring the same answer reach the same result.

One case covers one situation. A real product needs many: a return on day 29 and on day 31, a customer who writes in Spanish, a message about two orders at once.

Why each case is run more than once

Ordinary software gives the same output every time it receives the same input. A large language model (LLM), the type of AI model that produces text, can give a different output each time. One pass on one run is therefore weak evidence. Anthropic's engineering post states the practice that follows: "Each attempt at a task is a trial. Because model outputs vary between runs, we run multiple trials to produce more consistent results" [3].

For the sofa case, that means sending the same message several times and counting the answers that pass. Suppose, as an illustration, that 9 runs out of 10 give the 30-day answer. The tenth answer is one that a customer could receive on any day. The page on why a good demo does not show the product works gives the research on how much answers change between runs.

Who or what does the scoring

The thing that scores an output is called a grader. Anthropic's post lists three kinds: "code-based, model-based, and human" [3].

  • A check written in code suits anything with an exact answer, such as whether the reply contains the number 30. Code checks cost little and give the same result every time.
  • A person suits judgement, such as whether the tone is polite and the explanation is clear. People are slower and cost more per answer than code.
  • A second AI model can be given the output and a written rule and asked to judge. A model is faster than a person, and its verdicts need to be compared with people's verdicts before anyone relies on them.

The three kinds of LLM eval explains when each one fits. The model that produced an answer should never be the one that scores it, for the reasons in never let the model grade its own work.

How an eval differs from a software test, a demo and a benchmark

A software test

A software test checks that a piece of ordinary code returns one exact result, and it gives the same verdict on every run. An eval checks behaviour that varies, so it runs many cases, often several times each, and reports the share that passed. Google Cloud's documentation says its scoring rules, which it calls rubrics, "are similar to unit tests in software development" [4]. A unit test is a small test of one piece of code. For the longer comparison, see evals, tests and code review.

A demo

A demo is one run, on an input the presenter chose. An eval is many runs, on inputs chosen to cover what customers send, scored against expectations written in advance. The page on why a demo is weak evidence covers the difference.

A benchmark

A benchmark is a public set of questions that many AI models are scored on, so that the models can be compared with each other. An eval, as this guide uses the word, is your own set of cases about your own product. A model can score well on a public benchmark and still give your customers the wrong return period, because the benchmark contains nothing about your return policy. The guide on AI benchmarks and your own evals explains how to read a published score.

The words a vendor will use

Word What it means A question to ask
Eval One repeatable test: an input, an expectation and a scoring rule. Some people use the word for the whole set. Can I read one of yours?
Test case One input with its expectation. The sofa message above is one test case. Where did the cases come from?
Eval suite The full set of test cases for a product, run together. Also called an eval set. How many cases are in it?
Grader The code, person or AI model that scores each output. Which checks does a model score, and who checked that model?
Trial One attempt at one test case. How many times is each case run?
Pass rate The share of cases, or of trials, that passed, written as a percentage. Measured on how many cases?
Regression Something that worked before a change and fails after it. Is the suite run again after every change?

OpenAI's documentation says scores are often shown as numbers between 0 and 1 [1], so a pass rate of 0.92 and a pass rate of 92 percent are the same figure. What that figure is worth depends on how many cases it was measured on, which is the subject of how to read an AI eval report.

Who writes evals and how long a run takes

Two groups write an eval together. People who know the business write the expectations, because they are the ones who know that the return period is 30 days and that the shop never promises a refund in a chat. Engineers turn those expectations into checks a computer can run. The NIST AI Risk Management Framework, a voluntary framework from the United States standards institute, asks for a third group. Its MEASURE 1.3 outcome says regular assessments should involve internal experts who did not build the system, or independent assessors [5].

A first suite can be small. Anthropic's post says "20-50 simple tasks drawn from real failures is a great start" [3]. That is the company's advice from its own work, and no study supports the figure.

Run time depends on the number of cases, the number of trials and the type of grader. As an illustration with invented figures, if each trial takes 5 seconds, then 200 trials run one after another take 200 x 5 = 1,000 seconds, which is 16 minutes and 40 seconds. Trials can run at the same time, which shortens the wait. Checks scored by a person take as long as the reading takes. The cost is covered in what AI evals cost and who should own them.

When an eval is run

NIST's framework says: "AI systems should be tested before their deployment and regularly while in operation" [5]. OpenAI's documentation recommends running evals "on every change" and growing the set over time [1]. For a business, that means four moments: before the first release, after any change to the instructions the product gives the model, when the vendor replaces the model the product uses, and when a customer reports a wrong answer, which becomes a new test case. The first moment is covered in using evals to decide a release, and the third in what happens when the AI model changes.

Keep the cases and the expectations in files your company owns. As of 30 September 2026, OpenAI's documentation says its own online Evals product "will become read-only for existing users on October 31, 2026, and the platform is scheduled to shut down on November 30, 2026" [2]. A suite stored as plain files can be moved to another tool and run there.

What an eval cannot tell you

An eval reports on the cases in it. A situation that has no case is untested, however high the pass rate. The UK government's AI Security Institute tests the most advanced AI models. In February 2024, under its earlier name, the AI Safety Institute, it said its own evaluations "are thus not comprehensive assessments of an AI system's safety" and called the testing of advanced AI "a nascent science", meaning one newly begun [6].

So treat a pass rate as a measurement of known situations, and ask which situations were left out. Read the failed cases one by one, because each failure names a specific thing to fix.

How Reveneau uses evals

Reveneau is an AI software development consultancy, and we release a change when it has passed an eval suite. All of Reveneau's code is written by AI, and every change must pass a large eval suite before release. We write that suite from the specification, the written description of what the software must do, before the code exists. For an AI product the same order applies: at Reveneau we write the expected behaviour for each case with the client's own people, who know the rules of the business, and then build until the cases pass.

The checks in our suite that need judgement are graded by Jev, TypeSafe AI's decision model, which returns a probability for a written question instead of writing text. On Reveneau's own suite the run is ten times faster than with our previous language-model grader, so the suite can run on every change.

Because AI writes the code, a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. For an AI product whose expected behaviour is written down and tested before release, see AI development at Reveneau or contact us.

Common questions

What is an AI eval?

An AI eval is a repeatable test of an AI product. Each eval has an input, a written expectation of what a correct output must do, and a rule that scores the output, most often as pass or fail. OpenAI's documentation describes evals as tests of model outputs against style and content criteria that you specify. The same cases can be run again after any change, so two results can be compared.

What is an eval suite?

An eval suite is the full set of test cases for one AI product, run together to produce a pass rate. Some vendors call it an eval set. A first suite can be small: Anthropic's engineering post says 20 to 50 simple tasks drawn from real failures is a great start, which is that company's advice from its own work. The suite grows each time a customer reports a wrong answer.

What is a test case in an AI eval?

A test case in an AI eval is one input together with the written expectation for that input. In the furniture shop illustration on this page, the input is a customer asking to return a sofa after 40 days, and the expectation is an answer that states the 30-day period, promises no refund and offers the repair service. One attempt at a test case is called a trial.

Is an AI eval the same as a software test?

No. A software test checks that ordinary code returns one exact result, and the verdict is the same on every run. An AI eval checks output that varies between runs, so it runs many cases, often several times each, and reports the share that passed. Google Cloud's documentation compares its eval scoring rules to unit tests in software development, so the two methods are related.

Who writes evals in a company?

Evals are written by two groups working together. People who know the business write the expectations, because they know the rules a correct answer must follow. Engineers turn those expectations into checks a computer can run. The NIST AI Risk Management Framework also asks for internal experts who did not build the system, or independent assessors, to take part in regular assessments.

How long does an eval take to run?

An eval run takes as long as its cases, trials and graders require. As an illustration with invented figures, 200 trials at 5 seconds each, run one after another, take 1,000 seconds, which is 16 minutes and 40 seconds. Running trials at the same time shortens the wait. Checks scored by code finish in the time a computer needs to compare text, and checks scored by a person take as long as the reading takes.

What does a grader do in an AI eval?

A grader scores each output against the written expectation and returns a result such as pass or fail. Anthropic's engineering post lists three kinds: code-based, model-based and human. Code suits checks with an exact answer, a person suits judgement such as tone, and a second AI model can apply a written rule faster than a person once its verdicts have been compared with people's verdicts.

Is an AI eval the same thing as a benchmark?

No. A benchmark is a public set of questions that many AI models are scored on so the models can be compared. An AI eval, as this guide uses the word, is your own set of cases about your own product. A model with a high public score can still state your return policy wrongly, because the benchmark contains nothing about your business.

When should an AI eval be run?

An AI eval should be run at four moments: before the first release, after any change to the instructions the product gives the model, when the vendor replaces the model the product uses, and when a customer reports a wrong answer, which becomes a new test case. The NIST AI Risk Management Framework says AI systems should be tested before their deployment and regularly while in operation.

Does a passing eval mean an AI product is safe?

No. A passing eval shows that the product handled the cases in the suite, and a situation with no case is untested. The UK government's AI institute wrote in February 2024 that its own evaluations of advanced AI models are not comprehensive assessments of a system's safety. Ask which situations the suite leaves out, and read the failed cases one by one.

What happens to my evals if the tool that runs them is shut down?

Your evals remain usable after a tool is shut down if the cases and expectations are stored in files your company owns. As of 30 September 2026, OpenAI's documentation says its own online Evals product becomes read-only on October 31, 2026 and is scheduled to shut down on November 30, 2026. The inputs, expectations and scoring rules can be moved to another tool and run there.

What should I do first if my AI product has no evals?

Start by collecting real inputs where the product gave a wrong answer, and write one or two sentences for each that state what a correct answer must do. Anthropic's engineering post says 20 to 50 simple tasks drawn from real failures is a great start. Then ask an engineer to turn each expectation into a check a computer can run, and run the set before the next change is released.