AI benchmarks vs your own evals: how to read a model score / Decide with your own data
How to choose an AI model with your own evals
To choose an AI model, run each candidate on the same set of your own real cases, with the same instructions, and compare three numbers: the share of cases that pass your rule, the cost per task, and the response time. Then read the failures before you decide. Public rankings cannot make this choice: as of March 2026 the AI Index showed the top models of six companies within 79 rating points on Arena. Prices do differ. On 30 September 2026, an example task of 3,000 input tokens and 500 output tokens (a token is a piece of a word) cost $55.00 or $5,500.00 per 100,000 requests on two models from OpenAI's own price list.
Published September 30, 2026. Editorial.
Key takeaways
- Write the task, the rule for a correct answer and the limits on cost and response time before you look at any model's output.
- With 100 cases and an 80% pass rate the margin of error is 7.84 points, so compare two models case by case on the same set instead of by their totals.
- Cost per request is input tokens times the input price plus output tokens times the output price. Use token counts measured in your own run and prices dated to the day you read them.
- On OpenAI's price page on 30 September 2026, one example task cost 100 times more on gpt-6-astra than on gpt-6-luna, so always include a smaller model among the candidates.
- Keep the case set and run it again when a vendor retires a model, releases a new one or changes a price. Anthropic's documentation says it gives at least 60 days of notice before retirement.
As of March 2026, the top models from six companies were within 79 points of each other on Arena, a public site where people ask two unnamed models a question and vote for the answer they prefer. Stanford's AI Index 2026 lists the ratings: Anthropic at 1,503, xAI at 1,495, Google at 1,494, OpenAI at 1,481, Alibaba at 1,449 and DeepSeek at 1,424 [1]. The first two are 8 points apart. The report says gaps this narrow are "shifting competitive pressure toward cost, reliability, and domain-specific performance" [1].
A public table shows none of those three things for your product. On 30 September 2026, the price of one task on two models from the same vendor's price list differed by 100 times, as the arithmetic below shows. This page gives a seven-step method for choosing a model with your own evals, meaning tests built from your product's real cases. It is part of our guide to AI benchmarks vs your own evals. The wider decision, including data handling and contracts, is in how to choose an AI model.
What the model vendors tell their own customers
The three vendors whose documentation is cited here give the same advice. Google Cloud says its evaluation service shows "how a model performs on your specific tasks and against your unique criteria providing valuable insights which cannot be derived from public leaderboards and general benchmarks" [2]. A leaderboard is a public table that ranks models by score. That page describes a paid Google service, so the advice also promotes a product. OpenAI's documentation says "Design task-specific evals: Make tests reflect model capability in real-world distributions" [3]. Anthropic's says "Be task-specific: Design evals that mirror your real-world task distribution" [4]. In plain words, both say to build the tests from the mix of inputs your product receives in real use.
Each statement is a company describing how to use its own product, and all three tell the customer to decide on the customer's own cases.
The seven steps
| Step | What you have at the end | Time it takes |
|---|---|---|
| 1. Write the task and what correct means | One page: the task, the pass rule, the limits on cost and response time | One working session with the person who owns the result |
| 2. Collect real cases | A set of real inputs, each with the expected answer or the rule it must meet | The longest step, because people must agree on the expected answers |
| 3. Run every candidate the same way | Each model's outputs, with words used and seconds taken per case | Machine time, shorter than step 2 |
| 4. Score quality, cost and response time | One row per model with three numbers | Short when a program marks the answers, longer when people do |
| 5. Read the failures | A list of what each model got wrong and how bad each error is | A reading session for two people |
| 6. Decide and write the reason | A dated decision that names the model and the evidence | One meeting |
| 7. Keep the set and run it again | A new row each time a model or a price changes | Repeats for as long as the product runs |
Step 1: write down the task and what correct means
Anthropic's documentation says that building an LLM application (LLM stands for large language model) "starts with clearly defining your success criteria and then designing evaluations to measure performance against them" [4]. The same page asks two questions that belong in this step: "What is your budget for running the model?" and "What is the acceptable response time for the model?" [4].
Take a support assistant for an invented bicycle shop. As an illustration, a correct answer states the delivery fee and the return period exactly as the shop's policy document does, offers nothing the policy does not allow, and hands any question about an injury to a member of staff. The illustration's limits are 2 cents per answer and 5 seconds of waiting. Write rules like these before you look at any model, so that the model's output cannot shape the rule.
Step 2: collect real cases
Use inputs your product has received or will receive. OpenAI's documentation advises teams to log during development so that real inputs can become eval cases, and it names a mistake to avoid: "Creating eval datasets that don't faithfully reproduce production traffic patterns" [3]. In plain words, a test set that looks unlike your real traffic will mislead you. How to build an LLM eval dataset from real usage gives the full method.
Your cases have never been published, so no model can have seen them in the text it was trained on. Benchmark contamination explains why that matters.
How many cases you need to compare models
A pass rate measured on a sample of cases has a margin of error. Evan Miller's 2024 paper on statistics for evals, written at Anthropic, recommends computing a standard error for every eval score [5]. A standard error says how far a score would be expected to move on a different sample of cases. To get the standard error of a pass rate, multiply the rate by one minus the rate, divide by the number of cases, and take the square root. Multiplying by 1.96 gives the margin that covers 95 of 100 repeated samples. The table is our own arithmetic for a model that passes 80% of cases. For 100 cases, 0.8 times 0.2 divided by 100 is 0.0016, its square root is 0.04, or 4.00 points, and 1.96 times that is 7.84 points.
| Cases | Standard error | Margin of error (95%) |
|---|---|---|
| 50 | 5.66 points | 11.09 points |
| 100 | 4.00 points | 7.84 points |
| 200 | 2.83 points | 5.54 points |
| 400 | 2.00 points | 3.92 points |
With 100 cases, a model measured at 80% and a model measured at 84% are inside each other's margins of 7.84 points, so the two pass rates alone cannot rank them. Miller's paper recommends a more precise comparison for two models: analyse the differences case by case instead of comparing the two totals [5]. In practice, run both models on the same cases and list every case where one passed and the other failed. Start with 100 to 200 cases. How many test cases an LLM eval needs covers sample size in full.
Step 3: run every candidate the same way
Give every model the same cases, the same instructions and the same supporting documents. Record the exact model name and the date. Run each case more than once, because one model can give different answers to the same input. Save the count of tokens in and out for every answer. A token is a piece of a word, and it is the unit vendors bill in. Save the seconds each answer took as well.
Step 4: score quality, cost per task and response time
For quality, apply the pass rule from step 1 to every output. Anthropic's documentation advises structuring questions so that a program can mark them [4], and the three kinds of LLM eval explains when a person or a second model has to mark instead.
For cost, multiply the tokens you measured by each vendor's published price. The table below uses an invented task of 3,000 input tokens and 500 output tokens, and the prices each vendor published on its own pricing page on 30 September 2026 [6] [7] [8]. The cost per request is 3,000 divided by 1,000,000, times the input price, plus 500 divided by 1,000,000, times the output price. For gpt-6.1-sol that is 0.003 times $2.00 plus 0.0005 times $10.00, which is $0.006 plus $0.005, or $0.011.
| Model, as named on the vendor's page | Input, per 1M tokens | Output, per 1M tokens | Cost per request | Cost per 100,000 requests |
|---|---|---|---|---|
| gpt-6-astra (OpenAI) | $10.00 | $50.00 | $0.055 | $5,500.00 |
| gpt-6.1-sol (OpenAI) | $2.00 | $10.00 | $0.011 | $1,100.00 |
| gpt-6-luna (OpenAI) | $0.10 | $0.50 | $0.00055 | $55.00 |
| Claude Fable 5.1 (Anthropic) | $10.00 | $50.00 | $0.055 | $5,500.00 |
| Claude Opus 5.5 (Anthropic) | $4.00 | $20.00 | $0.022 | $2,200.00 |
| Claude Sonnet 5.5 (Anthropic) | $2.00 | $10.00 | $0.011 | $1,100.00 |
| Claude Haiku 4.5 (Anthropic) | $1.00 | $5.00 | $0.0055 | $550.00 |
| Gemini 3.1 Pro Preview (Google) | $2.00 | $12.00 | $0.012 | $1,200.00 |
| Gemini 3.8 Flash (Google), to 31 Dec 2026 | $0.75 | $3.75 | $0.004125 | $412.50 |
| Gemini 3.8 Flash (Google), from 1 Jan 2027 | $1.50 | $7.50 | $0.00825 | $825.00 |
Read the table as arithmetic on published prices for one invented task. Four cautions apply.
- Prices change. Google's page states that the Gemini 3.8 Flash input price is $0.75 through 31 December 2026 and $1.50 from 1 January 2027 [8]. The same 100,000 requests go from $412.50 to $825.00.
- Models use different numbers of tokens for the same job. Google's page bills a model's internal reasoning text, which it calls thinking tokens, as output [8]. Use the token counts from your own run in place of the invented ones.
- The same model has several prices. OpenAI's page lists a mode it calls Batch at half the standard price, and a tier it calls Fast at double the standard price for gpt-6-astra [6].
- The spread inside one vendor's list is wide. On OpenAI's page the example task costs $55.00 per 100,000 requests on gpt-6-luna and $5,500.00 on gpt-6-astra, a ratio of 100 [6].
The Princeton paper AI Agents That Matter argues for the same habit: evaluation "should account for dollar costs" instead of indirect measures such as the size of the model [9].
For response time, use the seconds you recorded in step 3. Report the typical case and the slowest cases separately, and compare both with the limit from step 1.
Step 5: read the failures as well as the average
Two models with the same pass rate can fail in different ways. In the bicycle shop illustration, one model might ask the customer to repeat the question, and another might promise a refund the policy forbids. The first failure costs a moment. The second costs money. OpenAI's documentation says "There's more to evals than just scores. Combine metrics with human judgment to ensure you're answering the right questions" [3]. Sort each model's failures by type and by how much harm each type does. Error analysis before metrics describes the reading method.
Is a smaller, cheaper model good enough
Your run answers this question, so include one smaller, lower-priced model among the candidates. If it meets the pass rule on your cases and its failures are the harmless type, the price gap in the table is money you save. AI Agents That Matter reported in 2024 that simple approaches beat many complex AI agents (programs in which a model works in steps and uses tools) on one coding test, HumanEval, while costing much less, and it concluded that agent evaluations "must be cost-controlled" [9]. That finding is about one test in mid-2024, so treat it as a reason to test the smaller model on your own cases.
Steps 6 and 7: decide, then keep the set and re-test
Write the decision down with its date, the model version, the three numbers and the failures you accepted. Then keep the case set, because the model you chose will be withdrawn one day. The vendors publish their notice periods:
- OpenAI's documentation gives "At least 6 months" for generally available models, meaning models released for general use, and says preview models "may be retired with much shorter notice, such as 2 weeks" [10].
- Anthropic's documentation says it provides "at least 60 days' notice before model retirement for publicly released models", and adds "Requests to models past the retirement date will fail" [11].
Both pages tell customers to test the replacement before the date [10] [11]. Run your set again in four situations: a vendor announces a retirement, a new model is released, a price changes, or your product changes. What happens to your product when the AI model changes covers that event.
Add new cases over time, and keep part of the set unseen by the people who tune the instructions. Benchmark saturation and Goodhart's law explains how a fixed set stops measuring what it once measured.
Should you pick the model at the top of a leaderboard
Use the top of a public table to choose two to four candidates, and let your run choose among them. The AI Index figures at the start of this page show six companies inside 79 rating points [1], and the same report notes that "third-party evaluations have documented cases where models perform more poorly in independent testing compared to developer-reported results" [1]. How to read a model vendor's benchmark claims lists the questions to ask about any published score. For the same per-request arithmetic applied to the model that marks eval answers, see the real price of an LLM judge.
How Reveneau applies this
At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. We choose models the same way. We write the pass rule and collect the client's cases first, run every candidate model on those cases with the same instructions, and record quality, cost per task and response time. The case set stays in the project and runs again when a vendor releases or retires a model.
We grade the judged checks in our suite with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than with our previous language-model grader. We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau as a company takes responsibility for the whole project through production and after release, which includes the model choice. To compare models on your own cases with us, start at our AI development service.
Best for
- A team choosing between two to four models for one defined task
- A product where cost per task or response time has a fixed limit
- A team that must switch models because a vendor is retiring one
Avoid if
- You have no real cases yet: collect them first
- Nobody has agreed what a correct answer is for the task
Check before you decide
- Every model ran on the same cases with the same instructions
- Prices were read from each vendor's page and dated
- Token counts and seconds came from your own run
- Someone read the failed cases for each model
Common questions
How do I choose an LLM for my product?
Choose an LLM for your product by running two to four candidate models on the same set of your own real cases, with the same instructions, and comparing pass rate, cost per task and response time. Then read the failed cases before deciding. Google Cloud's documentation says results on your specific tasks give insights which cannot be derived from public ranking tables and general benchmarks.
How many test cases do I need to compare two models?
Start with 100 to 200 test cases to compare two models. By our arithmetic, a model that passes 80% of 100 cases has a margin of error of 7.84 points, and with 400 cases the margin is 3.92 points. Evan Miller's 2024 paper on statistics for evals recommends comparing two models case by case on the same questions, which shows smaller differences than comparing totals.
Should I pick the model at the top of a leaderboard?
Use the top of a leaderboard, a public table that ranks models by score, to choose which models to try, and let a test on your own cases make the final choice. Stanford's AI Index 2026 shows the top models of six companies between 1,424 and 1,503 rating points on Arena as of March 2026, with the first two 8 points apart. The report says competition is moving toward cost, reliability and domain-specific performance.
How do I compare the cost of two AI models?
Compare the cost of two AI models by multiplying the input and output tokens each one used on your cases by the vendor's published price per million tokens. For an example task of 3,000 input and 500 output tokens, OpenAI's price page on 30 September 2026 gives $0.011 per request for gpt-6.1-sol. Date every price, because vendors change them.
How often should I re-test AI models?
Re-test AI models whenever a vendor announces a retirement, a new model is released, a price changes or your product changes. OpenAI's documentation gives at least 6 months of notice for generally available models and says preview models may be retired with notice such as 2 weeks. Anthropic's documentation says it gives at least 60 days. A saved case set makes each re-test a repeat run.
Is a smaller, cheaper model good enough for my product?
A smaller, cheaper model is good enough when it meets your pass rule on your own cases and its failures are the harmless type. Only your run can show that. The price gap can be large: on OpenAI's price page on 30 September 2026, an example task cost $55.00 per 100,000 requests on gpt-6-luna and $5,500.00 on gpt-6-astra. Always include one smaller model among the candidates.
How long does it take to compare models with my own evals?
Comparing models with your own evals takes as long as collecting the cases and agreeing on the expected answers, which is the longest of the seven steps. Running the models is machine time, and marking is short when a program can do it. Anthropic's documentation advises structuring questions so that grading can be automated, which shortens every later run.
What does it cost to run the model comparison itself?
Running the model comparison itself costs little compared with running the product. Using the example task of 3,000 input and 500 output tokens, 200 cases on gpt-6-astra cost 200 times $0.055, which is $11.00 per run, at OpenAI's published price on 30 September 2026. The larger cost is the time people spend agreeing on expected answers and reading failures.
What goes wrong when teams compare AI models?
Three mistakes make an AI model comparison unreliable: giving each model different instructions, using test cases that look unlike real traffic, and deciding on the average without reading the failures. OpenAI's documentation names the second mistake directly, as creating eval datasets that do not faithfully reproduce production traffic patterns. Fix all three before you compare any numbers.
Do AI model vendors recommend testing on your own data?
Yes. OpenAI, Anthropic and Google all tell customers to test models on their own tasks. OpenAI's documentation says to design task-specific evals. Anthropic's says to design evals that mirror your real-world task distribution. Google Cloud's says results on your specific tasks give insights that public ranking tables cannot. Each statement is a vendor describing its own product, and all three agree on this point.
What should I do when a vendor retires the model I chose?
When a vendor retires the model you chose, run your saved case set on the recommended replacement and on one or two other candidates before the retirement date. Anthropic's documentation says requests to models past the retirement date will fail, and OpenAI's says its notice periods give customers time to evaluate replacement models, test application behaviour and finish moving to them.
Do I need my own evals if I only use one vendor's models?
Yes, your own evals are still needed with one vendor, because one vendor sells several models at different prices and replaces them over time. Anthropic's pricing page on 30 September 2026 lists four current models from $1 to $10 per million input tokens. Your cases show which of them meets your pass rule, and whether a newer model still does.
References
- [1] Stanford HAI, AI Index Report 2026, Chapter 2: Technical Performance: Arena ratings as of March 2026 (Anthropic 1,503, xAI 1,495, Google 1,494, OpenAI 1,481, Alibaba 1,449, DeepSeek 1,424), the statement on cost, reliability and domain-specific performance, and the note on developer-reported results.
- [2] Google Cloud documentation, Gen AI evaluation service overview (read 30 September 2026): the statement that insights on your specific tasks cannot be derived from public leaderboards and general benchmarks. A product page for a paid service.
- [3] OpenAI API documentation, Evaluation best practices (read 30 September 2026): design task-specific evals, log during development, the mistake of datasets that do not reproduce production traffic, and combining metrics with human judgment.
- [4] Anthropic, Claude Platform documentation, Define success criteria and build evaluations (read 30 September 2026): define success criteria first, be task-specific, automate grading where possible, and the budget and response time questions.
- [5] Evan Miller (Anthropic), Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations (arXiv, 1 November 2024): compute standard errors for eval scores and compare two models on question-level paired differences.
- [6] OpenAI API documentation, Pricing (read 30 September 2026, standard tier, short context, US dollars per 1M tokens): gpt-6-astra $10.00 input and $50.00 output, gpt-6.1-sol $2.00 and $10.00, gpt-6-luna $0.10 and $0.50; batch at half price; Fast tier for gpt-6-astra at $20.00 and $100.00.
- [7] Anthropic, Claude Platform documentation, Pricing (read 30 September 2026, US dollars per 1M tokens): Claude Fable 5.1 $10 input and $50 output, Claude Opus 5.5 $4 and $20, Claude Sonnet 5.5 $2 and $10, Claude Haiku 4.5 $1 and $5.
- [8] Google, Gemini Developer API pricing (read 30 September 2026, paid tier, US dollars per 1M tokens): Gemini 3.8 Flash $0.75 input and $3.75 output through 31 December 2026, then $1.50 and $7.50; Gemini 3.1 Pro Preview $2.00 and $12.00 for prompts up to 200k tokens; output prices include thinking tokens.
- [9] Kapoor, Stroebl, Siegel, Nadgir and Narayanan, AI Agents That Matter (arXiv, 1 July 2024, Princeton University): agent evaluations must be cost-controlled, should account for dollar costs, and simple baselines beat many complex agents on HumanEval at lower cost.
- [10] OpenAI API documentation, Deprecations (read 30 September 2026): at least 6 months of notice for generally available models, preview models may be retired with notice such as 2 weeks, and the purpose of the notice period.
- [11] Anthropic, Claude Platform documentation, Model deprecations (read 30 September 2026): at least 60 days of notice before retirement, requests after the retirement date fail, and the advice to test replacement models before that date.
Related reading
The real price of an LLM judge is per decision, and it multiplies
A grader's price per call looks small. Multiply it by every check on every change, or every message through a guardrail, and the public figures from LangChain, Openlayer and TypeSafe show where a cheaper decision model pays for its own migration and where it does not.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
How to tell if an AI feature idea is worth building
Most AI feature ideas look good in a demo and fail during the work of making them reliable. Here are the four questions we ask to tell the ones worth building from the ones that just look good in a demo.