Why AI evals matter: the test that shows whether an AI product works / Budget and rules
What AI evals cost and who should own them
No public source measures what a typical company spends on AI evals, so this page gives you the parts to price yourself. The cost has four parts: writing down what a correct output is, collecting and labelling test cases, paying for the model usage each test run consumes, and keeping the set current as the product and the model change. Three of those are people's time. The running cost is arithmetic on a vendor's published price. Ownership follows the same split: the product side owns what correct means, engineering owns the running, and the company owns the release decision. When a partner builds the product, the contract should say that you keep the test set.
Published September 30, 2026. Editorial.
Key takeaways
- No public source measures the typical spend on AI evals. The one published cost in our sources is a research study: 21,730 test runs for a total that its authors round to $40,000.
- The running cost is arithmetic: cases, times attempts per case, times the tokens (pieces of text) each attempt uses, times the vendor's price on the day. An invented 500-case set at $2 and $10 per million tokens costs $13.50 per run.
- Evals recur. The set is re-run on every change, grows with every failure found in real use, and must be re-run whenever a vendor retires the model the product uses.
- Product owns the definition of a correct output, engineering owns the running, and the company owns the release decision. NIST's framework asks for these roles to be documented.
- If a partner builds the product, put in the contract that the test cases, the expectations and the past results are yours, in a format you can open without the partner's tools.
The only published cost of a large AI test run in the sources for this page comes from a research project. In October 2025 a research paper by Sayash Kapoor, Benedikt Stroebl and 29 other authors reported 21,730 test runs across 9 models and 9 public benchmarks, at a total cost that they round to $40,000 [1]. A benchmark is a public test set that many models are scored on. Dividing the rounded total by the number of runs gives $1.84 per run. That division is ours, and it describes that study only.
The figure does not apply to a business budget. A company testing its own support assistant has a different test set, a different model and a different number of runs. No public source that we could verify measures what companies spend on evals, in money or in staff hours. This page gives the parts of the cost, who does each part, and the arithmetic for the one part that has a published price.
The four parts of the cost
An eval is a repeatable test of an AI product: a set of inputs, a written expectation for each one, and a way to score the output. The terms are explained in what an AI eval is. Building one and keeping it costs four things.
| Cost part | Who does it | When it recurs |
|---|---|---|
| Writing down what a correct output is | People who know the business: product, support, legal | When the product or a policy changes |
| Collecting and labelling test cases | Product and support staff, with engineering | At the start, then after each failure found in real use |
| Running the tests | Engineering sets it up; the model vendor bills for usage | On every change to the prompt, the model or the documents |
| Keeping the set current | Product and engineering together | All the time, and at every model retirement |
Three of the four parts are people's hours, which you price from your own salaries and your own estimate of the time. Running the tests has a published unit price, and a later section works through it.
Writing down what a correct output is
The first cost is the time of people who know what the business wants. An engineer can build the test. The person who can say whether a refund answer is right usually works in product, support or legal. Anthropic's documentation puts this step first: building a successful application "starts with clearly defining your success criteria and then designing evaluations to measure performance against them" [3].
Budget this as writing time for the product owner and for whoever knows the rules the product must follow. The output is a short document: the types of request the product must handle, what a correct response contains for each type, and what must never appear. The page on using evals to decide a release describes the pass lines that go with it.
Collecting and labelling test cases
The second cost is building the set: gathering real or realistic inputs and writing the expected behaviour for each. Labelling means that a person marks what the right output is, or marks whether a given output was right.
Start small. Anthropic's engineering guide says that 20 to 50 simple tasks "drawn from real failures" is a good start, and it gives no measurement behind that range [2]. The same company's documentation shows example sets of 100 customer inquiries, 200 articles and 1,000 tweets. Those are illustrations, with no claim that they are minimum sizes [3]. How many cases you need before you can trust a result is a statistics question, covered in how many test cases an LLM eval needs. LLM stands for large language model, the type of AI model that produces text.
The same guide warns about delay: "Evals get harder to build the longer you wait" [2]. That is a vendor's advice with no data attached. The reasoning behind it: the longer a product runs without recorded cases, the more of its expected behaviour exists only in people's memories.
Running the tests, shown as arithmetic
Running an eval means sending every test input to the model and scoring what comes back. Model vendors bill by the token. A token is a piece of text, often a short word or part of a word. Prices are quoted per million tokens, with one price for the text you send (input) and a higher price for the text the model writes (output).
The cost of one run is: cases x attempts per case x ((input tokens x input price) + (output tokens x output price)).
As of 30 September 2026, OpenAI's price page lists its gpt-6.1-sol model at $2.00 per million input tokens and $10.00 per million output tokens [4]. Anthropic's price page lists Claude Sonnet 5.5 at $2 and $10 [5]. Both are the vendors' own published prices and both can change, so read the current page before you budget.
Here is an invented illustration at those prices. A set has 500 cases. Each case is run 3 times, because a model can answer differently each time. Each attempt sends 2,000 tokens and receives 500.
| Step | Working | Result |
|---|---|---|
| Input cost per attempt | 2,000 x $2.00 / 1,000,000 | $0.004 |
| Output cost per attempt | 500 x $10.00 / 1,000,000 | $0.005 |
| Cost per attempt | $0.004 + $0.005 | $0.009 |
| Attempts per run | 500 x 3 | 1,500 |
| Cost per run | 1,500 x $0.009 | $13.50 |
| 20 runs in a month | 20 x $13.50 | $270.00 |
Three things change that result. If a second model grades each answer, every attempt needs a second call, which is one more request to the model. Assuming a call of the same size, the run costs $27.00. A more expensive model raises the cost in proportion: OpenAI's page lists gpt-6-astra at $10.00 and $50.00, which makes the same run $67.50 [4]. And OpenAI's page lists a batch mode, where results come back later, at half the standard price, which makes it $6.75 [4].
Agents cost more per case. An agent is an AI system that works in many steps, so one test case is many model calls in a row. The $1.84 per run in the research study above is an example [1]. Our guide to cost, time and step limits for agents covers how to limit it, and the real price of an LLM judge covers the cost of the grading call.
Keeping the set current
An eval set is a recurring cost. Three events add work.
- The product changes. Each new feature needs new cases, and each changed policy means editing the expectations that mention it.
- A failure is found in real use. That input becomes a new case and stays in the set.
- The vendor retires the model. OpenAI's model retirement page states a minimum notice of 6 months for its generally available models, meaning those released for all customers to use, and says the notice gives customers time to "test application behavior" before the model is switched off [6]. Each retirement means a full re-run on the replacement model and a review of every case whose result changed. The detail is in what happens to your product when the AI model changes.
Anthropic's guide describes the timing of the whole cost: the "costs are visible upfront while benefits accumulate later" [2]. The spending comes in the first weeks, and the return comes at every later change.
What a failure costs, from the records that exist
A budget should also count the cost of a failure that reaches a customer. Here the public record is short as well. No source that we could verify measures the average cost of an AI product giving a wrong answer. What exists is the amount in individual decisions.
- $812.02. Ordered by a British Columbia tribunal in Moffatt v. Air Canada on February 14, 2024, after the airline's website chatbot misstated its bereavement fare policy [8].
- $5,000. A penalty set by a US federal court in Mata v. Avianca on June 22, 2023, on two lawyers and their firm, who filed court papers that cited judicial opinions made up by ChatGPT [9].
- $193,000. The amount in the US Federal Trade Commission's final order against DoNotPay, announced on February 11, 2025, over claims about a service the company promoted as "the world's first robot lawyer". The company settled, so no court ruled on the facts [10].
These amounts leave out legal fees, staff time and lost customers, which the records do not state. The cases are described in public AI failures and the tests that target each one.
Leave one figure out of your budget argument. IBM's 2025 Cost of a Data Breach report, sponsored by IBM and conducted by the Ponemon Institute on 600 organisations, puts the global average cost of a data breach at $4.44 million [11]. That number measures security breaches. A wrong answer from an AI product is a different event, and it has no published average.
Who owns what inside the company
Ownership follows the parts of the cost.
Product owns what correct means. The product owner, with support and legal where needed, writes and approves the expectations. Anthropic's guide describes one setup in which the criteria are "defined by the product team", with people checking the automated grader from time to time [2].
Engineering owns the running. Engineers build the program that runs the tests, connect it to the release process and report the results.
The company owns the decision. Whether to release on a given result is a business decision, made by the role the company has given that authority, and recorded.
NIST's AI Risk Management Framework, a voluntary framework from the US government standards body, asks for this clarity. Its GOVERN 2.1 outcome says that roles and responsibilities for measuring and managing AI risks "are documented and are clear to individuals and teams throughout the organization" [7]. It also says: "Ideally, AI actors carrying out verification and validation tasks are distinct from those who perform test and evaluation actions" [7]. In plain words, the people who confirm the system is fit for use should be different from the people who ran the tests.
A dedicated eval team is optional. A product with one AI feature needs the three roles assigned and the hours set aside, and the people in those roles can hold other jobs.
Who keeps the eval set when a partner builds the product
If an outside company builds your AI product, the eval set is part of what you are paying for. It is the written record of what the product is supposed to do, and whoever holds it can check any future change. Ask these questions before you sign.
- Do we own the test cases, the written expectations and the grading instructions at the end of the contract?
- In what format are they delivered, and can we open and run them without your tools?
- Do we receive the results of past runs, with dates?
- Who writes the expectations, and who on our side approves them?
- Who pays for the model usage of each run, and is that inside the price?
- What happens to the set when the model is retired?
Question 2 has a current example. OpenAI's documentation says, as of 30 September 2026, that its own online Evals platform becomes read-only on October 31, 2026 and is scheduled to shut down on November 30, 2026 [6]. A test set that is stored only inside one tool depends on that tool staying open. Keep the cases in plain files that you hold.
Two related pages give more detail: how to evaluate a vendor's eval suite and the eval suite as a diligence artifact.
Where Reveneau fits
Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. The eval suite is therefore a normal part of what we build.
Two parts of our practice apply to cost. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. We grade the checks that need judgment with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader.
On ownership, we ask every client to name who approves the expectations on their side before the work starts, and to settle in the contract who holds the eval set at the end. Reveneau, as a company, takes responsibility for the whole project through production and after release, so keeping the set current after release is part of that responsibility.
To discuss an eval budget for your product, see AI development at Reveneau or contact us. The argument for testing AI products is in why AI evals matter.
Common questions
How much do AI evals cost?
AI evals cost four things: the time to write down what a correct output is, the time to collect and label cases, the model usage of each run, and the time to keep the set current. No public source measures a typical total. The one published cost in our sources is a research study that reported 21,730 runs for a total its authors round to $40,000.
Who should own evals in a company?
Ownership of evals in a company splits three ways. Product owns the definition of a correct output, engineering owns building and running the tests, and the company owns the release decision through a role it has given that authority. NIST's AI Risk Management Framework asks, in its GOVERN 2.1 outcome, that roles for measuring and managing AI risks are documented and clear across the organisation.
Do I need a dedicated eval team?
A dedicated eval team is optional for a company with one AI feature. What you need is the three roles assigned and the hours set aside: someone in product who approves the expectations, an engineer who runs the tests, and a role that decides releases. NIST's framework adds that the people who validate a system should ideally be different from the people who ran the tests.
Who owns the eval set when a vendor builds the product?
The contract decides who owns the eval set when a vendor builds the product, so settle it before you sign. Ask for the test cases, the written expectations, the grading instructions and the dated results of past runs, in a format you can open without the vendor's tools. OpenAI's own online Evals platform is scheduled to shut down on November 30, 2026, which shows why plain files matter.
Are evals a one-time cost?
Evals are a recurring cost. The set is re-run on every change to the product, a new case is added for every failure found in real use, and a full re-run is needed when the vendor retires a model. OpenAI's model retirement page states a minimum notice of 6 months for models released to all customers, so plan for a retirement during the life of the product.
How do I work out the cost of one eval run?
Work out the cost of one eval run by multiplying cases, attempts per case, and the token cost of one attempt, where a token is a piece of text that the vendor bills for. In this page's invented illustration, 500 cases run 3 times each, at 2,000 input and 500 output tokens, cost $13.50 at $2 and $10 per million tokens. OpenAI and Anthropic each listed one model at those prices on 30 September 2026.
What does an AI failure cost a company?
The cost of an AI failure has no published average that we could verify. The public record gives amounts in single decisions: $812.02 in Moffatt v. Air Canada, a $5,000 court penalty in Mata v. Avianca, and $193,000 in the Federal Trade Commission's order against DoNotPay. Those amounts leave out legal fees, staff time and lost customers, which the records do not state.
Can I use IBM's data breach figures to justify an eval budget?
IBM's data breach figures measure a different event, so leave them out of an eval budget. IBM's 2025 report, conducted by the Ponemon Institute on 600 organisations, puts the global average cost of a data breach at $4.44 million. A breach is a security event. An AI product giving a wrong answer is a quality failure, and the report gives no figure for it.
How small can a first eval set be?
A first eval set can be a few dozen cases. Anthropic's engineering guide says 20 to 50 simple tasks drawn from real failures is a good start, and it gives no measurement behind that range. The cost of a set that size is people's time, which you price from your own hours. Whether a set that size supports a release decision is a separate question about sample size.
Why does testing an AI agent cost more than testing a single answer?
Testing an AI agent costs more because one test case is many model calls in a row, and each call is billed. A 2025 research study by Kapoor, Stroebl and 29 other authors reported 21,730 agent runs for a total its authors round to $40,000, which is $1.84 per run by our division. A single question and answer in this page's illustration costs $0.009 per attempt.
Does the eval budget matter if we buy an AI product instead of building one?
The eval budget still matters when you buy an AI product, because someone has to check that it works on your own cases. Your cost is then the first two parts: writing down what a correct output is for your business, and collecting cases from your own use. Ask the seller who pays for running those cases and whether you keep the results.
References
- [1] Kapoor, Stroebl and 29 other authors, Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation, arXiv preprint (13 October 2025): 21,730 agent runs across 9 models and 9 benchmarks at a total cost the authors round to $40,000. A research study, so the cost is that study's only.
- [2] Anthropic, Demystifying evals for AI agents (9 January 2026): 20 to 50 simple tasks drawn from real failures as a start, evals get harder to build the longer you wait, costs are visible upfront while benefits accumulate later, and criteria defined by the product team. A vendor's own engineering advice with no measurement behind it.
- [3] Anthropic, Define success criteria and build evaluations, Claude Platform Docs (read 30 September 2026): start from written success criteria; example set sizes of 100, 200 and 1,000 items. Vendor documentation.
- [4] OpenAI, Pricing, API documentation (read 30 September 2026): gpt-6.1-sol at $2.00 input and $10.00 output per million tokens, gpt-6-astra at $10.00 and $50.00, and batch mode at half the standard price. The vendor's own prices, which can change.
- [5] Anthropic, Pricing, Claude Platform Docs (read 30 September 2026): Claude Sonnet 5.5 at $2 input and $10 output per million tokens. The vendor's own prices, which can change.
- [6] OpenAI, Deprecations, API documentation (read 30 September 2026): at least 6 months' notice for generally available models, the stated purpose of the notice, and the Evals platform becoming read-only on October 31, 2026 and shutting down on November 30, 2026. The vendor's own policy.
- [7] NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023): GOVERN 2.1 on documented roles, and the sentence on keeping verification and validation separate from test and evaluation.
- [8] British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, 2024 BCCRT 149 (14 February 2024): the order to pay $812.02 in total.
- [9] US District Court, Southern District of New York, Mata v. Avianca, Inc., Opinion and Order on Sanctions (22 June 2023): the $5,000 penalty on two lawyers and their firm for filing non-existent judicial opinions created by ChatGPT.
- [10] US Federal Trade Commission, FTC Finalizes Order with DoNotPay (11 February 2025): the final order requiring $193,000 in monetary relief. A consent order, so no court ruled on the facts.
- [11] IBM, press release for the 2025 Cost of a Data Breach Report (30 July 2025): research by the Ponemon Institute on 600 organisations, global average cost of a data breach $4.44 million. A vendor-sponsored study of security breaches.
Related reading
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
How to negotiate a software contract you can actually verify
Most build contracts describe effort, timeline, and payment, and leave the one hard question unanswered: on what evidence do you agree the thing is finished?
The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.