LLM evals: how to measure whether an AI product works / Run it over time
Offline evals vs online evals: before release and after it
An offline eval runs a fixed set of test cases before release and tells you whether a change made a known case better or worse. An online eval measures real traffic after release, through user feedback, sampled human review, automatic checks and controlled experiments, and finds the failures nobody wrote a case for. Each finds what the other cannot, so a product needs both. The link between them is one habit: a failure seen online becomes a new offline case. Run the offline set on every change and on a fixed schedule, because a model reached over the internet can change under the same name: one 2023 study measured GPT-4 at 84 percent in March and 51 percent in June on the same questions.
Published September 30, 2026. Editorial.
Key takeaways
- An offline eval runs a fixed set of cases with confirmed answers before release. It finds a change that made a known case worse and misses requests unlike anything in the set.
- An online eval measures live traffic after release, using user feedback, sampled human review, automatic checks and controlled experiments. It finds new failure types, after users have seen them.
- Hamel Husain and Shreya Shankar suggest reviewing 100 or more fresh traces in each review cycle, with typical cycles of 2 to 4 weeks.
- Chen, Zaharia and Zou measured GPT-4 at 84 percent in March 2023 and 51 percent in June 2023 on the same prime number questions, so run the offline set on a fixed schedule as well as on every change.
- OpenAI's documentation, read on 30 September 2026, says its hosted Evals platform becomes read-only for existing users on 31 October 2026 and shuts down on 30 November 2026.
In March 2023, GPT-4 answered 84 percent of a set of questions about prime numbers correctly. In June 2023, the model sold under the same name answered 51 percent of the same questions correctly. Lingjiao Chen, Matei Zaharia and James Zou measured both figures and published them that year [6]. A team that tested its product in March and stopped there had no record of the change.
Two different measurements would have found it. Running the same fixed test set again in June is an offline eval. Watching the quality of real answers after release is an online eval. This page, part of the guide to LLM evals for AI products, explains what each one finds, what each one misses, how they connect, and which open-source tools exist as of 30 September 2026.
What an offline eval is
An offline eval runs a fixed set of test cases through the product before a change is released. The LangSmith documentation defines it this way: "Offline evaluations target examples from datasets: curated test cases with reference outputs that define what "good" looks like" [1]. A reference output is the answer a person has confirmed as correct for that case.
Because the cases stay the same, the result can be compared from one run to the next. That makes an offline eval the tool for one question: did this change make a known case better or worse? The LangSmith documentation lists this use as "Regression testing: Ensure new versions don't degrade quality" [1]. A regression is something that used to work and now fails.
Hamel Husain and Shreya Shankar describe the usual form. The set runs automatically on every change, a practice programmers call continuous integration, or CI. "Test datasets for CI are small (in many cases 100+ examples) and purpose-built", and the authors say to "Favor assertions or other deterministic checks over LLM-as-judge evaluators" [2]. In plain words: prefer checks written in code, which give the same result on every run, over a model that grades. How to build the set is in how to build an LLM eval dataset from real usage.
What an online eval is
An online eval measures the product on live traffic after release. The LangSmith documentation again: "Online evaluations target runs and threads from tracing: real production traces without reference outputs" [1]. A trace is the stored record of one request: the input, each step the system took, and the output.
Real requests arrive with no confirmed answer attached, so the measures are different. Four are in common use.
- User feedback. A rating button, an edited answer, a repeated question, a request to speak to a person.
- Sampled human review. A person reads a sample of traces and labels them. Husain and Shankar suggest a goal of 100 or more fresh traces in each review cycle, and describe typical cycles of 2 to 4 weeks [2].
- Automatic checks on a sample. Code checks, and model graders that need no reference answer. Husain and Shankar note that these graders cost more to run, and advise tracking a confidence interval, the range the true rate plausibly lies in, for every production figure [2].
- Controlled experiments. Part of the traffic gets the new version. The next section covers these.
How to make the human review reliable is in human review and agreement between reviewers.
What an A/B test is for an AI feature
A controlled experiment splits live users at random between two versions and compares one agreed measure. Ron Kohavi and three co-authors, then at Microsoft, give the standard definition in a 2009 survey: "In the simplest controlled experiment, often referred to as an A/B test, users are randomly exposed to one of two variants: Control (A), or Treatment (B)" [4]. They call such experiments "the best scientific design for establishing a causal relationship between changes and their influence on user-observable behavior" [4].
For an AI feature, version A is the current prompt (the written instruction given to the model) or the current model and version B is the new one. Three points from the survey apply to an AI feature [4].
- Choose the measure before the test starts. The survey names it the Overall Evaluation Criterion: "A quantitative measure of the experiment's objective". For a support assistant it might be the share of conversations closed without a person.
- Assign users at random, so that any difference in the result comes from the versions.
- Test the measuring system first by giving both groups the same version, which the survey calls an A/A test. At a 95 percent confidence level, such a test should report a difference in 5 percent of runs.
Two limits apply. The survey was written about web pages, before LLM products existed, and it says nothing about the quality of model output. And an experiment is slow: Anthropic's guide lists the weakness of A/B testing as "Slow; days or weeks to reach significance and requires sufficient traffic" [3]. An A/B test reports which version did better on the chosen measure. Finding the individual wrong answers still takes sampled review.
What each one finds and what each one misses
| Offline eval | Online eval | |
|---|---|---|
| When it runs | Before release, on every change | After release, on live traffic |
| What it runs on | A fixed set of cases with confirmed answers | Real requests with no confirmed answer |
| What it can check | Correctness against the expected answer | Quality patterns, safety and real behaviour [1] |
| What it finds | A change that made a known case worse | Failures nobody wrote a case for, and how often each happens |
| What it misses | Requests unlike anything in the set | A problem before users see it |
| Cost | Low for each run once the set exists | Review time in every cycle; days or weeks of traffic for an experiment |
The LangSmith documentation puts the first difference in one sentence: "offline evaluations can check correctness against expected answers, while online evaluations focus on quality patterns, safety, and real-world behavior" [1]. Husain and Shankar put the second: "CI evals protect against known regressions before deployment", while online monitoring finds failures in production traffic and estimates how often they occur [2].
Anthropic's guide lists two weaknesses of production monitoring that explain the "misses" row: "Reactive; problems reach users before you know about them" and "Lacks ground truth for grading" [3]. Ground truth means a confirmed correct answer. Husain's ordering of cost is that an A/B test costs more than human and model review, which costs more than checks in code [7].
Our position follows from the table. A team with only offline evals learns nothing about requests it never imagined. A team with only online evals learns about each failure from a user. Run both.
A failure seen online becomes an offline case
The two are connected by one habit. The LangSmith documentation describes it: problems found by online evaluations become offline test cases, offline evaluations confirm the fixes, and online evaluations confirm the improvement in production [1]. Husain and Shankar give the same instruction: "when production monitoring reveals new failure patterns through error analysis and evals, add representative examples to your CI dataset" [2].
In steps:
- Sample traces from live traffic and label them.
- Group the failures into types, as error analysis describes.
- For each new type, copy representative requests into the offline set, remove personal data, and record the expected result.
- Repair the fault and run the offline set. The new cases should now pass and the old ones should still pass.
- After release, watch the online rate of that failure type.
OpenAI's guide adds the precondition: record every request and response during development, so that the records can supply test cases later [5]. The same loop for multi-step systems is in turning production traces into eval cases.
How often each should run
Run the offline set on every change to the prompt, the model, the search step or the code. OpenAI's guide says to "run evals on every change" and to "grow the eval set over time" [5]. Anthropic's guide says automated evals should run "on each agent change and model upgrade" [3].
Run it on a fixed schedule as well, even when you changed nothing. The opening figures are the reason: a model reached over the internet can change while its name stays the same. That study compared two dated versions of two OpenAI models, and we read only its summary [6]. It shows that such a change can happen and gives no rate for how often. The business side is in what happens when the AI model changes.
Husain's own schedule is a useful model: "I often run Level 1 evals on every code change, Level 2 on a set cadence and Level 3 only after significant product changes" [7]. Level 1 is checks in code, Level 2 is human and model review, Level 3 is A/B testing.
Open-source eval tools as of 30 September 2026
Here are six open-source eval tools. The licence and the owning account below were read from each project's public code store on GitHub, a code hosting site, on 30 September 2026. Each description comes from the project's own introduction file, the README, so it is the project describing itself.
| Tool | Owner on GitHub | Licence | Its own description |
|---|---|---|---|
| promptfoo [8] | promptfoo | MIT | A tool "for evaluating and red-teaming LLM-based apps" |
| Inspect AI [9] | UKGovernmentBEIS | MIT | "a framework for large language model evaluations created by the UK AI Security Institute" |
| OpenAI Evals [10] | openai | MIT for the code; each dataset has its own | "a framework for evaluating large language models (LLMs) or systems built using LLMs" |
| Ragas [11] | vibrantlabsai | Apache 2.0 | "Evaluate your LLM applications with precision using both LLM-based and traditional metrics." |
| DeepEval [12] | confident-ai | Apache 2.0 | "a simple-to-use, open-source LLM evaluation framework" |
| Langfuse [13] | langfuse | MIT outside three enterprise folders | "an open source LLM engineering platform" |
Red-teaming, in the promptfoo line, means attacking your own product on purpose to find its weaknesses.
Four notes on ownership and support, each from the project's or vendor's own pages as read on 30 September 2026. The promptfoo README says: "Promptfoo is now part of OpenAI. Promptfoo remains open source and MIT licensed" [8]. The Langfuse README says "since January 2026 we're part of ClickHouse" [13]. The DeepEval README promotes Confident AI, the paid product of the company that owns it [12]. And OpenAI's documentation says its hosted Evals platform will "become read-only for existing users on October 31, 2026" and "is scheduled to shut down on November 30, 2026" [5]. That notice concerns the hosted product, and it does not mention the open-source code [10].
We have not tested these tools against each other and we do not rank them. The test set and the labels are what a team should own. Keep them in plain files that any of these tools, or the next one, can read.
Where Reveneau fits
At Reveneau all code is written by AI, and every change must pass a large eval suite written from the specification before the code exists. That suite is our offline eval, and it runs on every change. We grade its judged checks with Jev, TypeSafe AI's decision model, and on our own suite the run is ten times faster than it was with our previous language-model grader. A faster run shortens the wait between a change and its result.
Reveneau, as a company, takes responsibility for the whole project through production and after release. That is the online half: we scope the review of live traces into the work, and a failure found in production becomes a new offline case before the repair is written. Because we use AI instead of hiring more engineers, the build takes a small team, and that saving goes into the client's price.
For the general argument, read our post on measuring an AI product before and after release. To plan both halves for your product, see AI development at Reveneau or contact us.
Best for
- Offline: deciding whether one change to a prompt, model or search step is safe to release
- Online: finding failure types nobody predicted and counting how often each happens
- Both together: any AI feature that real users depend on
Avoid if
- Offline alone: the product receives requests the team never imagined
- Online alone: a wrong answer reaching a user has a cost you cannot accept
- An A/B test: traffic is too low to give a result within weeks
Check before you decide
- The offline set runs on every change and on a fixed schedule
- A named person reviews a sample of live traces in every cycle
- Each failure found online was added to the offline set
- The test set and labels are stored in files the team owns
Common questions
What is the difference between offline evals and online evals?
Offline evals run a fixed set of test cases with confirmed answers before a change is released. Online evals measure real requests after release, where no confirmed answer exists. The LangSmith documentation says offline evaluations can check correctness against expected answers, while online evaluations focus on quality patterns, safety and real-world behaviour. Offline finds a known case that got worse, and online finds failures nobody wrote a case for.
How do I monitor LLM quality in production?
Monitor LLM quality in production with four measures: user feedback, human review of a sample of traces, automatic checks on a sample, and controlled experiments. Hamel Husain and Shreya Shankar suggest reviewing 100 or more fresh traces in each review cycle, with typical cycles of 2 to 4 weeks, and tracking a confidence interval for each production figure. Every new failure type found is then added to the offline test set.
What is an A/B test for an AI feature?
An A/B test for an AI feature sends live users at random to one of two versions, such as the current prompt and a new one, and compares one measure chosen in advance. Ron Kohavi and co-authors define it in a 2009 survey as users randomly exposed to a control or a treatment. Anthropic's guide notes that A/B testing takes days or weeks and needs sufficient traffic.
Which open-source eval tools exist?
Six open-source eval tools are promptfoo, Inspect AI, OpenAI Evals, Ragas, DeepEval and Langfuse. As read from each project's public code on 30 September 2026, promptfoo and Inspect AI use the MIT licence, Ragas and DeepEval use the Apache License 2.0, OpenAI Evals uses MIT for its code, and Langfuse uses MIT outside its enterprise folders. Reveneau has not tested them against each other.
How often should evals run?
Offline evals should run on every change to the prompt, the model, the search step or the code, and also on a fixed schedule when nothing was changed. Online review runs in cycles. Hamel Husain writes that they run checks in code on every code change, human and model review on a set schedule, and A/B tests only after large product changes. Husain and Shankar describe review cycles of 2 to 4 weeks.
Why run the offline set again if I changed nothing?
Run the offline set on a fixed schedule because a model reached over the internet can change while its name stays the same. Chen, Zaharia and Zou measured GPT-4 at 84 percent accuracy on prime number questions in March 2023 and 51 percent on the same questions in June 2023. That study covers two dated versions of two models, so it shows the change can happen and gives no rate.
Should a new product start with offline or online evals?
A product with no users yet starts with offline evals, because there is no live traffic to measure. Build a small fixed set and run it on every change. During development, record every request and response, which OpenAI's guide advises so that the records can supply test cases later. Online review starts with the first real users, and its findings are added to the offline set.
What does an online eval cost compared with an offline one?
An offline eval costs little for each run once the test set exists, because most of its checks are code. An online eval costs review time in every cycle, model calls for graders that work without a reference answer, and days or weeks of traffic for an experiment. Hamel Husain orders the cost as A/B testing above human and model review, and that above checks in code.
What goes wrong when a team relies on offline evals alone?
A team that relies on offline evals alone learns nothing about requests it never imagined. The fixed set only contains cases someone thought of, so a new failure type passes unseen until a user reports it. Husain and Shankar describe the split: evals that run on every change protect against known regressions, and online monitoring finds failures in production traffic and estimates how often they occur.
What goes wrong when a team relies on online evals alone?
A team that relies on online evals alone learns about each failure from a user. Anthropic's guide lists the weaknesses of production monitoring as reactive, because problems reach users before the team knows about them, and as lacking confirmed correct answers for grading. With no fixed set to run before release, the team also cannot tell whether a repair broke a case that used to work.
What is happening to OpenAI's hosted Evals platform?
OpenAI's documentation, read on 30 September 2026, says the company is retiring its hosted Evals platform. According to that notice, Evals becomes read-only for existing users on 31 October 2026 and the platform is scheduled to shut down on 30 November 2026. The notice concerns the hosted product. The open-source OpenAI Evals code on GitHub is a separate thing, and its licence file was MIT when read on 30 September 2026.
How does a failure found in production become a test case?
A failure found in production becomes a test case in five steps: label a sample of live traces, group the failures into types, copy representative requests into the offline set with personal data removed and the expected result recorded, repair the fault and run the set, then watch the online rate of that failure type. The LangSmith documentation describes the same loop between online and offline evaluation.
References
- [1] LangChain, LangSmith documentation, Evaluation concepts (read 30 September 2026): definitions of offline and online evaluation; what each can check; regression testing; how online findings become offline test cases.
- [2] Hamel Husain and Shreya Shankar, AI Evals: Everything You Need to Know (page dated 18 September 2026): test sets that run on every change are small, in many cases 100 or more examples; prefer deterministic checks; online monitoring finds failures and estimates how often they occur; 100 or more fresh traces each review cycle; cycles of 2 to 4 weeks; confidence intervals for production figures.
- [3] Anthropic, Demystifying evals for AI agents (9 January 2026): automated evals run on each agent change and model upgrade; production monitoring is reactive and lacks confirmed answers; A/B testing takes days or weeks and needs sufficient traffic.
- [4] Kohavi, Longbotham, Sommerfield and Henne, Controlled experiments on the web: survey and practical guide, Data Mining and Knowledge Discovery 18, pages 140 to 181 (2009): the definition of an A/B test; random assignment; the Overall Evaluation Criterion; the A/A test; the 95 percent confidence level.
- [5] OpenAI, Evaluation best practices (API documentation, read 30 September 2026): log everything; run evals on every change and grow the eval set; the notice that the hosted Evals platform becomes read-only on 31 October 2026 and shuts down on 30 November 2026.
- [6] Chen, Zaharia and Zou, How is ChatGPT's behavior changing over time? (arXiv, submitted 18 July 2023; summary only was read): GPT-4 identified prime against composite numbers with 84 percent accuracy in March 2023 and 51 percent in June 2023.
- [7] Hamel Husain, Your AI Product Needs Evals (29 March 2024): three levels of eval, their order of cost, and how often the author runs each.
- [8] promptfoo repository on GitHub (read 30 September 2026): owner promptfoo; MIT licence; README description; README notice that Promptfoo is now part of OpenAI and remains MIT licensed.
- [9] Inspect AI repository on GitHub (read 30 September 2026): owner UKGovernmentBEIS; MIT licence, copyright UK AI Security Institute; README description.
- [10] OpenAI Evals repository on GitHub (read 30 September 2026): owner openai; MIT licence for the code, with each dataset under its own licence; README description.
- [11] Ragas repository on GitHub (read 30 September 2026): owner vibrantlabsai, to which the older address redirects; Apache License 2.0; README description.
- [12] DeepEval repository on GitHub (read 30 September 2026): owner confident-ai; Apache License 2.0; README description; the README promotes the Confident AI platform.
- [13] Langfuse repository on GitHub (read 30 September 2026): owner langfuse; MIT licence outside the ee folders, which have a separate licence; copyright ClickHouse, Inc.; README says the project has been part of ClickHouse since January 2026.
Related reading
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
Adding eval tests was the best decision we made
We generate every line of code we release. The single change that made that safe was writing the check before the code, and it turned out to change how we review, how we spec, and how fast we can work.
Never let the model grade its own work
When the same run writes the code and the tests, a passing build proves only that the code is consistent with itself. That is not verification, and it is the most common way an AI-built codebase becomes confidently wrong.