Guide

Why AI evals matter: the test that shows whether an AI product works

An AI product can give different output on different days and for different inputs, so a demo, a vendor's promise and a public benchmark score are weak evidence that it works for your customers. An AI eval is the stronger evidence: a written, repeatable test with an input, an expectation and a scoring rule. This guide is written for the people who approve budgets and releases. It explains what an eval is, what public records show about wrong AI answers and advertising claims, how to read a pass rate and use it to decide a release, what the work costs and who owns it, and what the texts of regulators say about testing.

Published September 30, 2026. Editorial.

Key takeaways

  • An AI eval is a repeatable test: an input, a written expectation of a correct output, and a rule that scores the output. The full set for one product is an eval suite.
  • A demo is one run on a chosen input. A 2024 study of five language models found accuracy differences of up to 15% between runs, under settings meant to make output repeatable.
  • On February 14, 2024, a British Columbia tribunal ordered Air Canada to pay $812.02 after finding it did not take reasonable care to ensure its chatbot was accurate.
  • The AI model your product uses can change. On 1,000 prime-number questions, GPT-4's accuracy was 84.0% in March 2023 and 51.1% in June 2023, as measured by Chen, Zaharia and Zou.
  • A 92% pass rate on 50 cases has a 95% range of 84.5% to 99.5%. On 1,000 cases the range is 90.3% to 93.7%. Always ask for the number of cases.
  • Write the pass line before the test runs. NIST's generative AI profile suggests minimum thresholds reviewed as part of deployment approval, and NIST leaves the number to each organisation.
  • As of 30 September 2026, the EU AI Act's testing rules for high-risk systems apply from 2 December 2027 or 2 August 2028, depending on the type of system. A lawyer should confirm what applies to you.

On February 14, 2024, the British Columbia Civil Resolution Tribunal ordered Air Canada to pay a customer, Jake Moffatt, a total of $812.02 in Canadian dollars. A chatbot on the airline's website had told Moffatt they could apply after travelling for a bereavement fare, a reduced fare for people who travel because of a death in the family. A page on the same website said the opposite. The tribunal member wrote: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate" [1].

The amount is small, and a small claims decision binds only the two parties. The reasoning is what a business owner should note. The tribunal treated the chatbot as part of the company's website and wrote: "It makes no difference whether the information comes from a static page or a chatbot" [1].

This guide is written for the person who approves the budget or the release of an AI product and who has never written a test. An AI product gives different output on different days and on different inputs. A demo, a vendor's promise and a public benchmark score (a model's result on a shared public test) are therefore weak evidence that the product works for your customers, and a written, repeatable test, called an eval, is the evidence to ask for.

An eval is a written test that can be run again

An AI eval is a repeatable test of an AI product. It has three parts: an input (the question or document the product receives), a written expectation of what a correct output must do, and a rule that scores the output, most often as pass or fail. The full set of these cases for one product is called an eval suite, and the share of cases that pass is the pass rate.

OpenAI's documentation gives the reason ordinary software tests are insufficient here: "Generative AI is variable. Models sometimes produce different output from the same input, which makes traditional software testing methods insufficient for AI architectures" [2]. An ordinary program gives the same output for the same input every time, so one check is enough. A language model, the type of AI model that produces text, can answer the same question correctly on Monday and wrongly on Tuesday. Each case is therefore run several times, across many cases, and the result is a rate.

The expectations come from people who know the business. Take an invented furniture shop whose policy allows returns for 30 days. The shop knows that rule and the AI model has to be told it, so the shop is the one who can say that an answer promising 90 days is a failure. Engineers then turn each expectation into a check a computer can run. The page what is an AI eval shows one case written out in a table and explains the words a vendor will use.

A good demo is one run on a chosen input

A demo shows that the product can give a good answer once, on an input the presenter selected and rehearsed. It gives no information about how often the product is wrong, because each input ran one time.

Research posted in August 2024 by Berk Atil and 12 other authors shows how much one run can mislead. They took five large language models, configured them with the settings meant to give the same output every time, and ran eight common tasks 10 times each. They report "accuracy variations up to 15% across naturally occurring runs" [3]. They also report that "none of the LLMs consistently delivers repeatable accuracy across all tasks" [3]. LLM is short for large language model. That is the largest difference they observed, in a paper posted before formal review. If accuracy can move that much between runs under settings chosen to prevent it, a single run in a meeting is a weak basis for a purchase.

NIST, the United States standards institute, gives the same advice in its 2024 profile for generative AI: "Avoid extrapolating GAI system performance or capabilities from narrow, non-systematic, and anecdotal assessments" [4]. GAI is NIST's short form for generative AI.

A buyer may also ask what share of AI pilots reach real use by customers. This guide gives no percentage, because no measured figure could be verified. After a demo, ask in writing for the eval set, the pass rate with the number of cases it was measured on, and the failed cases. The page why a good AI demo does not show the product works has the full argument and a table comparing what a demo and an eval each show.

Public records show what a wrong answer or claim has cost

Tribunals and regulators have already acted on wrong answers from AI products and on advertising claims about them. In the Air Canada decision above, the order of $812.02 was made up of $650.88 in damages, $36.14 in interest and $125 in tribunal fees [1].

In the United States, the Federal Trade Commission (FTC) finalized an order against DoNotPay on February 11, 2025. The FTC's release says the company had promoted its service as "the world's first robot lawyer". The order requires it to pay $193,000 and prohibits it from advertising that the service "performs like a real lawyer" unless it has sufficient evidence for the claim [5]. It is a consent order, which means the company settled and no court ruled on the facts.

A second FTC case shows the difference between a claimed figure and a tested one. According to the FTC's complaint, Workado promoted its AI content detector as "98 percent" accurate, and independent testing showed an accuracy rate of 53 percent on general-purpose content [6]. That case was also settled, and the FTC's release leaves out who ran the test and how many texts it used.

Each of these records concerns the accuracy of a statement: an answer the product gave, or a claim the company made about the product. A test written in advance, with its results kept, is how a company produces evidence of accuracy. The page public AI failures and the tests that target each one describes each record and names the type of test case that targets each type of failure.

The AI model your product uses changes on the vendor's schedule

Most AI products are built on a model that another company owns, updates and retires. Each of those changes reaches the product.

Lingjiao Chen, Matei Zaharia and James Zou measured one such change. They gave the March 2023 and June 2023 versions of GPT-4 the same 1,000 questions asking whether a number is prime. They report: "GPT-4's accuracy dropped from 84.0% in March to 51.1% in June", while GPT-3.5 on the same questions rose from 49.6% to 76.2% [7]. The changes went in both directions, and the figures describe 2023 models only. What still applies is the authors' conclusion: a service that keeps the same name can behave differently a few months later, so it needs continuous monitoring.

Vendors also retire models. As of 30 September 2026, Anthropic's documentation says the company gives "at least 60 days' notice before model retirement for publicly released models", and that "Requests to retired models will fail" [8]. That is a company's statement of its own policy, which it can change.

Without an eval suite, the first sign of a change for the worse is a customer complaint. With one, a model change means running the same cases again and comparing the two results before customers see the new model. The page what happens to your product when the AI model changes lists what to run again and when.

A pass rate means little without the number of cases it was measured on

An eval report gives a pass rate. The first question to ask is how many cases it was measured on, because a small set leaves a wide range of possible true rates.

Anthropic's explanation of its statistics paper gives the method. A 95% confidence interval, the range that is likely to contain the true rate, "can be calculated from the SEM by adding and subtracting 1.96 × SEM from the mean score" [9]. SEM is the standard error of the mean, a number that says how far the score could move if a different set of questions had been drawn. For a pass rate p measured on n cases, the standard formula is the square root of p x (1 - p) / n. Here is the calculation for a pass rate of 92%, as an illustration.

Cases Standard error 1.96 x standard error 95% range
50 square root of (0.92 x 0.08 / 50) = 0.0384 0.0752 84.5% to 99.5%
1,000 square root of (0.92 x 0.08 / 1,000) = 0.0086 0.0168 90.3% to 93.7%

On 50 cases, a reported 92% could be a true rate as low as 84.5%. On 1,000 cases, the lower end of the range is 90.3%.

An average can also hide a group that fails. In the 2018 Gender Shades study, Joy Buolamwini and Timnit Gebru tested 3 commercial systems that classify gender from a photograph of a face. The error rate was up to 34.7% for darker-skinned women, and the maximum for lighter-skinned men was 0.8% [10]. The study is about face analysis software, and it shows one overall score hiding a group that fails. Ask for results split by the groups of customers and types of request that matter to you. The page how to read an AI eval report is a printable checklist for this.

Set the pass line before the test is run

An eval result is useful when it decides something. The pass line is the lowest pass rate at which you will release. NIST's generative AI profile suggests this action: "Establish minimum thresholds for performance or assurance criteria and review as part of deployment approval" [4]. A threshold is a minimum value, and deployment is NIST's word for release. NIST's AI Risk Management Framework describes the decision itself: "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed" [11]. NIST leaves the number to you: the framework says it "does not prescribe risk tolerance" [11]. Both documents are voluntary.

Write the pass line down before the test runs, because a line chosen after the result is known can be set to match the result. Set a higher line for outputs that can cost a customer money or cause harm, and name the failures that block a release whatever the average is. For the invented furniture shop, one such failure is any answer that promises a refund outside the policy. And make the release a decision of the company, recorded together with the result it was based on. A check that a change must pass before release is called a release gate.

When a failure is found after release, add it to the suite as a permanent case, so every later change is tested for it. The page using evals to decide whether an AI feature is ready to release gives a table by type of output.

What evals cost and who owns them

No source we could verify measures what companies spend on evals, in money or in staff time, so this guide gives no typical budget. The cost has four parts: writing the expectations, collecting and labelling cases, running the tests, and keeping the set current.

One published figure shows the running cost of a large research project. In October 2025, Sayash Kapoor, Benedikt Stroebl and 29 other authors reported 21,730 agent attempts across 9 models and 9 public benchmarks, at a total cost that the authors round to $40,000 [12]. An agent is an AI program that takes steps on its own. That cost belongs to one study of public models, and a company's suite for one product is a different size.

On where to start, Anthropic's engineering post says "20-50 simple tasks drawn from real failures is a great start" and warns that "Evals get harder to build the longer you wait" [13]. Both statements are the company's advice from its own work.

Ownership has three parts. The people responsible for the product own what a correct answer is. Engineering owns running the tests. The company owns the release decision. NIST's framework adds, in its MEASURE 1.3 outcome, that regular assessments should involve internal experts who did not build the system, or independent assessors [11]. If a development partner builds the product, agree in the contract that the eval set, with its cases and expectations, belongs to you when the contract ends. The page what AI evals cost and who should own them sets out each cost part and who does it.

What regulators and standards bodies say about testing

In the European Union's AI Act, Article 9(8) says of high-risk AI systems: "Testing shall be carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system" [14]. That means deciding what you will measure, and the minimum result, before you test. The article applies to systems the Act classes as high-risk, a legal category that most business software is outside.

The dates changed in 2026. As of 30 September 2026, Regulation (EU) 2026/1744, published on 24 July 2026, makes the high-risk rules apply from 2 December 2027 for systems classed as high-risk under Annex III of the Act, and from 2 August 2028 for those classed under Annex I. The date it replaced was 2 August 2026 [15].

In the United States, the NIST AI Risk Management Framework is voluntary. It says: "AI systems should be tested before their deployment and regularly while in operation" [11].

This section describes what the texts say and is general information. A lawyer should confirm what applies to a specific product. The page what regulators and standards bodies say about testing AI covers each text with its article numbers and dates, and the EU AI Act and your product explains the Act's risk classes.

What to read after this guide

If you read one more page, read the one on eval reports and take its checklist to your next vendor meeting. For the questions to ask about the suite itself, see how to evaluate a vendor's eval suite. Teams that will build the tests can continue with the method guides: LLM evals, AI agent evals and AI benchmarks and your own evals. Evals for code that AI writes are covered in eval-driven development.

Where Reveneau fits

Reveneau is an AI software development consultancy, and this guide describes how we work. All of Reveneau's code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. The specification is the written description of what the software must do. For an AI product, we write the expected behaviour for each case with the client, set the pass line before the test runs, and show the client the cases, the pass rate and the failures.

The checks in that suite that need judgement are graded by Jev, TypeSafe AI's decision model, which returns a probability for a written question instead of writing text. On Reveneau's own suite the run is ten times faster than with our previous language-model grader, so the suite can run on every change.

We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To plan an AI product with its tests written first, see AI development at Reveneau or contact us.

Explore the guide

Common questions

Why do AI evals matter to a business?

AI evals matter to a business because an AI product can answer the same question correctly one day and wrongly the next, and customers act on the answers. In February 2024 a British Columbia tribunal ordered Air Canada to pay a customer $812.02 after finding the airline did not take reasonable care to ensure its chatbot was accurate. An eval is a written, repeatable test of what the product tells customers.

Do I need evals if I buy an AI product instead of building one?

Yes. A buyer needs the vendor's eval results for the same reason a builder needs evals: a demo is one run on a chosen input. Ask the vendor in writing for the eval set, the pass rate with the number of cases it was measured on, and the failed cases. NIST's 2024 generative AI profile advises against extrapolating performance from narrow, anecdotal assessments.

When should a company start writing evals?

A company should start writing evals before the first release of an AI feature, and where possible before the feature is built. Anthropic's engineering post says evals get harder to build the longer you wait, and that 20 to 50 simple tasks drawn from real failures is a great start. Both statements are that company's advice from its own work. Writing the expectation first also makes the team agree on what the feature is for.

Can a public benchmark score replace my own evals?

No. A public benchmark score compares AI models on a shared set of questions, and it contains nothing about your products, policies or customers. Your own evals test the product on your own cases. A model with a high benchmark score can still state your return policy wrongly. A chatbot stating a company policy wrongly was the failure in the Air Canada tribunal decision of February 2024.

What is the first thing to ask an AI vendor about testing?

Ask the vendor how many test cases the pass rate was measured on and where those cases came from. The count matters because a 92% pass rate on 50 cases has a 95% range of 84.5% to 99.5%, and the same rate on 1,000 cases has a range of 90.3% to 93.7%, using the confidence interval method described by Anthropic. Then ask to read the failed cases.

Are AI evals only for large companies?

No. AI evals suit any company whose AI product gives answers to customers, whatever its size. A first suite can be small: Anthropic's engineering post says 20 to 50 simple tasks drawn from real failures is a great start. The first task at that size is writing: people who know the business state what a correct answer must do for each case.

What happens if a company releases an AI feature without evals?

A company that releases an AI feature without evals learns about wrong answers from customer complaints, and has no record of what it checked. Regulators have also acted on advertising claims about AI products. The US Federal Trade Commission's final order of February 11, 2025, a settlement, requires DoNotPay to pay $193,000 and prohibits it from advertising that its service performs like a real lawyer unless it has sufficient evidence.

Do evals stop once the product is released?

No. Evals continue after release, because the product and the model it uses keep changing. The NIST AI Risk Management Framework says AI systems should be tested before their deployment and regularly while in operation. Run the suite again after every change to the product, whenever the vendor replaces or retires the model, and each time a customer reports a wrong answer, which becomes a new case.

Can a non-engineer take part in writing evals?

Yes. A non-engineer writes the part of an eval that the rest depends on, the expectation: one or two sentences stating what a correct answer must do for a given input. That knowledge comes from the business, such as a return policy or a pricing rule. Engineers then turn each expectation into a check a computer can run. NIST's framework also asks for people who did not build the system to take part in assessments.

Do evals make an AI product free of errors?

No. Evals measure how often a product passes the cases in the suite, and a situation with no case stays untested. In a 2024 study of five language models by Berk Atil and co-authors, none consistently delivered repeatable accuracy across all tasks. Evals show the error rate on known situations, so that a company can set a pass line, block specific failures and decide a release with evidence.

How does Reveneau use evals in its own work?

Reveneau writes all of its code with AI, and every change must pass a large eval suite, written from the specification before the code, before release. The checks that need judgement are graded by Jev, TypeSafe AI's decision model, and on Reveneau's own suite the run is ten times faster than with its previous language-model grader. Reveneau as a company takes responsibility for the project through production and after release.

References