Why a good AI demo does not show the product works
A demo is one run of an AI product on an input the presenter chose, so it shows that the product can give a good answer once. Three things limit what it proves: the inputs were selected, a language model's output changes between runs, and real customers send inputs nobody planned for. In a 2024 study of five language models set up to give the same output every time, accuracy still varied by up to 15% between runs across eight tasks. After a demo, ask for three things in writing: the eval set, the pass rate with the number of cases it was measured on, and the failed cases.
Published September 30, 2026. Editorial.
Key takeaways
- A demo is a single run on chosen inputs. NIST's 2024 generative AI profile says to avoid extrapolating performance from narrow, non-systematic and anecdotal assessments.
- Output varies between runs. A 2024 study of five language models, eight tasks and 10 runs each found accuracy differences of up to 15% between runs, with settings meant to prevent variation.
- As an illustration, a product that is right 90% of the time on each try gives five correct demo answers in a row 59% of the time, because 0.9 to the power of 5 is 0.59049.
- Impressions can differ from measurement. In a 2025 trial with 16 developers and 246 tasks, the developers estimated AI saved 20% of their time while measured completion time rose 19%.
- After a demo, ask in writing for the eval set, the pass rate with the count of cases and runs it was measured on, and the failed cases with their outputs.
In a study published in July 2025, 16 experienced software developers completed 246 tasks, some with AI tools allowed and some without. Afterwards the developers estimated that AI had cut their completion time by 20%. The measurement showed that allowing AI had increased completion time by 19% [1]. The study is small and covers one type of work, so it supports no general claim about AI and speed. Its use here is narrow: people who used an AI tool on real tasks formed an impression of how well it worked, and the measurement showed the opposite.
A demo asks you to form that impression in one meeting. A demo is one run of an AI product on an input somebody chose, and it shows that the product can produce a good answer once. Three things limit what it proves: the input was selected, the output of a language model changes between runs, and real users send inputs nobody planned for. The stronger evidence is an eval, a repeatable test made of inputs, a written expectation for each, and a rule that scores each output. This page belongs to the guide on why AI evals matter, and the definition is in what an AI eval is.
A demo is one run on an input somebody chose
Every demo input is selected. The presenter picks questions the product handles well, in the wording it handles well, and rehearses them. The demo is therefore a sample of the product's best cases.
NIST, the United States standards institute, addresses this in its profile for generative AI, published in July 2024. One of its suggested actions reads: "Avoid extrapolating GAI system performance or capabilities from narrow, non-systematic, and anecdotal assessments" [2]. GAI is NIST's short form for generative AI. A demo is narrow because it has few inputs, non-systematic because nobody planned which situations to cover, and anecdotal because it reports one occasion. The profile is voluntary guidance.
The same document warns about a second popular piece of evidence, an AI model passing a test designed for people. It says that testing through "standardized tests designed for humans (e.g., intelligence tests, professional licensing exams) does not guarantee GAI system validity or reliability in those domains" [2]. A slide saying that the model passed a professional exam belongs in the same category as the demo.
OpenAI's documentation names the habit in its list of mistakes to avoid: using "it seems like it's working" as an evaluation strategy [3].
The same input gives different output on different runs
A large language model (LLM) is the type of AI model that produces text, and it can give a different output each time it receives the same input. Research posted in August 2024 by Berk Atil and 12 other authors tested this under the conditions most favourable to repeatability. They took five LLMs, configured them with the settings meant to make output the same every time, and ran eight common tasks 10 times each. They report "accuracy variations up to 15% across naturally occurring runs with a gap of best possible performance to worst possible performance up to 70%" [4]. They add that "none of the LLMs consistently delivers repeatable accuracy across all tasks" [4]. Those are the largest differences the authors observed, and the paper is a preprint, a research paper posted before formal review.
A demo samples that variation once. If the run you watched was one of the good ones, the demo gives you no way to know.
Simple arithmetic shows how much a short demo can hide. Take an invented product that gives a correct answer 90% of the time on any single try, with each try independent of the others. The chance that five demo questions in a row all come out correct is 0.9 x 0.9 x 0.9 x 0.9 x 0.9 = 0.59049, which is 59%. A product that fails one customer in ten therefore gives a faultless five-question demo in 59 of every 100 demos. The chance that 20 questions in a row all come out correct is 0.9 to the power of 20 = 0.1216, which is 12%. These figures are an illustration.
Anthropic's engineering team gives the same type of calculation for agents, which are AI programs that take several steps on their own: "If your agent has a 75% per-trial success rate and you run 3 trials, the probability of passing all three is (0.75)³ ≈ 42%" [5]. A trial is one attempt at one task. The method for measuring this is in agent reliability across repeated runs.
Real users send inputs nobody planned for
The third limit is the difference between a meeting room and real use. NIST states it in one sentence: "Measurement gaps can arise from mismatches between laboratory and real-world settings" [2]. Customers misspell words, write in other languages, paste in long documents, ask two things at once, and ask for things the product was never meant to do.
A test published in June 2024, called τ-bench (tau-bench), measured agents under conditions closer to real use. Its authors, Shunyu Yao and three co-authors, set up conversations between a simulated user and an AI agent that had been given tools and policy rules to follow. They report that agents built on leading models of that date "succeed on <50% of the tasks", and that in the retail setting fewer than 25% of tasks were completed correctly in all 8 of 8 attempts [6]. Those models are old now, the user in the test was simulated by a language model, and the figures give no information about current products. The method still applies: the authors ran the same task repeatedly and counted how often every attempt succeeded. A business with thousands of customers needs that number, and a demo cannot produce it.
The records of what wrong answers have cost companies are in public AI failures and the tests that target each one.
What the research on failed AI projects says
A buyer may ask what share of AI pilots reach production, meaning real use by customers. This page gives no percentage. The figures that are often repeated could not be traced to a measurement we could read, and the RAND Corporation's 2024 report on the subject repeats an estimate made by other people.
RAND's own research was a set of interviews. For the report published on 13 August 2024, James Ryseff, Brandon De Bruhl and Sydne Newberry interviewed 65 data scientists and engineers, 50 from industry and 15 from universities, each of whom had built AI models for at least five years [7]. Their findings are about causes:
- "Misunderstandings and miscommunications about the intent and purpose of the project are the most common reasons for AI project failure" [7].
- "Too often, trained AI models are deployed that have been optimized for the wrong metrics or do not fit into the overall business workflow and context" [7]. A metric is the number a team chooses to measure success.
- 14 of the 50 industry interviewees said senior leaders underestimated the time needed to train a model that solved the business problem [7].
- 30 of the 50 discussed persistent problems with data quality [7].
RAND states that most interviewees were engineers, so the findings may reflect an engineer's view more than a manager's. The first two findings describe questions that a demo cannot answer. A demo shows a capability. What the project is for, and which number counts as success, are the two things an eval makes a team write down.
What a demo shows and what an eval shows
| A buyer's question | What a demo shows | What an eval shows |
|---|---|---|
| Can the product do this task at all? | Yes, on the inputs shown | Yes, and on what share of the cases tested |
| How often is it wrong? | Nothing | A pass rate on a stated number of cases |
| Does it give the same answer next time? | Nothing, because each input ran once | The share of repeated attempts that passed |
| What does it do with inputs we did not plan? | Nothing, because the presenter chose the inputs | Results on cases taken from real customer messages |
| Which cases fail, and how? | Nothing, because failures are left out | The failed cases, readable one by one |
| Will it still work after a change? | Nothing | The same cases run again and compared with the last run |
A demo still has a use. It shows what the product is meant to do and whether the idea is worth testing. Treat it as the first step of the evaluation and make the buying decision after the later steps.
How a pilot differs from an eval
A pilot is a period in which a limited group of real users works with the product. It is better evidence than a demo, because the inputs are real. A pilot's inputs arrive once, so after a change to the product nobody can run last month's pilot again. And a pilot is usually judged by how the users felt, which the developer study at the top of this page shows can differ from the measurement.
Save the inputs from the pilot and write the expected behaviour for each. They become the first cases of an eval suite, the full set of test cases for a product. Anthropic's engineering post says "20-50 simple tasks drawn from real failures is a great start" [5], which is that company's advice from its own practice, with no study supporting the figure. The method is in how to build an LLM eval dataset from real usage.
What to ask for after a demo
Ask the vendor, or your own team, for three things in writing.
- The eval set. The list of inputs and the expected behaviour for each. Check who wrote the expectations and whether the inputs came from real customers. If no set exists, the product has been judged by impression alone.
- The pass rate, with the count it was measured on. A percentage, the number of cases it was measured on, and how many times each case was run. How to read an AI eval report explains why 92% of 50 cases is weaker evidence than 92% of 1,000.
- The failures. The cases that failed, with the actual outputs. They show what some customers will receive.
Then ask to add ten of your own inputs, taken from real customer messages, and to see the results. Ask also whether the demo inputs are in the eval set, and what happens to the product when the AI model it uses is replaced, which is covered in what happens when the AI model changes. Two articles give more detail: how to evaluate a vendor's eval suite and the work between an AI demo and production.
Where Reveneau fits
At Reveneau we treat a demo as an introduction and an eval suite as the evidence. All of Reveneau's code is written by AI, and every change must pass a large eval suite before release. We write that suite from the specification, the written description of what the software must do, before the code exists. For an AI feature, the expected behaviour for each case is written down with the client before the feature is built, and what we show a client is the suite's result: the cases, the pass rate and the failures.
The checks in that suite that need judgement are graded by Jev, TypeSafe AI's decision model, which returns a probability for a written question instead of writing text. On Reveneau's own suite the run is ten times faster than with our previous language-model grader, so the suite can run on every change.
We use AI instead of hiring more engineers, so a build takes a small team, and that saving goes into the client's price. Reveneau, as a company, takes responsibility for the whole project through production and after release. To discuss a product you are planning, see AI development at Reveneau or contact us.
Common questions
Why does my AI product work in the demo and fail later?
An AI product works in the demo and fails later because the demo was one run on inputs the presenter chose, and real use is thousands of runs on inputs nobody chose. A 2024 study by Berk Atil and co-authors found accuracy differences of up to 15% between runs on five language models set up to give repeatable output. Customers also send messages the demo never covered.
How many AI pilots reach production?
The share of AI pilots that reach production has no measured figure that could be verified for this page, so none is given. The RAND Corporation's 2024 report on failed AI projects repeats an estimate made by others and reports causes from 65 interviews instead of a rate. Its most common cause was misunderstanding about the intent and purpose of the project, which a written eval makes a team settle.
What should I ask for after an AI demo?
After an AI demo, ask for three things in writing: the eval set (the inputs and the expected behaviour for each), the pass rate with the number of cases and runs it was measured on, and the failed cases with their outputs. Then ask to add ten of your own real customer messages to the set. OpenAI's documentation lists judging by the impression that the product seems to be working as a mistake to avoid.
How is a pilot different from an eval?
A pilot is a period with a limited group of real users, and an eval is a fixed set of inputs with written expectations that can be run again after every change. A pilot gives real inputs once, and its result is usually how the users felt. Saving the pilot's inputs and writing the expected behaviour for each turns them into the first cases of an eval suite.
Why does an AI product give different answers to the same question?
An AI product built on a language model gives different answers to the same question because the model's output varies between runs. A 2024 study by Berk Atil and 12 other authors ran eight tasks 10 times each on five models configured for repeatable output, and reported that none of them consistently delivered repeatable accuracy across all tasks. A single demo run samples that variation once.
Is an AI demo useful for anything?
Yes. An AI demo shows what the product is meant to do and whether the idea is worth testing further. A demo shows nothing about how often the product is wrong, because each input ran once and the presenter chose it. Treat the demo as the first step of the evaluation, then ask for the eval set, the pass rate and the failed cases before any buying decision.
Does passing a professional exam show that an AI product works?
No. NIST's 2024 profile for generative AI says that testing through standardized tests designed for humans, such as professional licensing exams, does not guarantee validity or reliability in those domains. An exam result describes the model on exam questions. Your product needs a test on your own customers' questions, with expectations that come from your own business rules.
Can I judge an AI tool by how well it seemed to work for my team?
An impression of an AI tool is weaker evidence than a measurement, and the two can disagree. In a 2025 randomized trial by the research group METR, 16 experienced developers completed 246 tasks and afterwards estimated that AI had cut their completion time by 20%, while the measurement showed completion time rose by 19%. The study is small and covers one type of work.
How much can a short demo hide about error rates?
A short demo can hide an error rate of one answer in ten. As an illustration, a product that is correct 90% of the time on each independent try gives five correct answers in a row in 59% of demos, because 0.9 multiplied by itself five times is 0.59049. The same product gives 20 correct answers in a row 12% of the time, because 0.9 to the power of 20 is 0.1216.
What does the research say causes AI projects to fail?
RAND's 2024 report, based on interviews with 65 data scientists and engineers, says the most common cause of AI project failure is misunderstanding and miscommunication about the intent and purpose of the project. The report also says models are often deployed after being optimized for the wrong measure, and 30 of the 50 industry interviewees discussed persistent data quality problems. Most interviewees were engineers.
Does a longer demo give better evidence?
A longer demo gives more examples of the same type of evidence, because the presenter still chooses the inputs and each one still runs once. NIST's generative AI profile advises against extrapolating performance from narrow, non-systematic and anecdotal assessments. Better evidence comes from changing the method: inputs taken from real customers, each run several times, and scored against expectations written in advance.
References
- [1] Becker, Rush, Barnes and Rein (METR), Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv, 12 July 2025, revised 25 July 2025): 16 developers, 246 tasks; developers estimated a 20% time saving; measured completion time rose 19%.
- [2] NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024): action MS-2.5-001 on anecdotal assessments; tests designed for humans do not guarantee validity; measurement gaps between laboratory and real-world settings.
- [3] OpenAI, Evaluation best practices (API documentation, read 30 September 2026): lists using "it seems like it's working" as an evaluation strategy among mistakes to avoid.
- [4] Atil and 12 other authors, Non-Determinism of "Deterministic" LLM Settings (arXiv, 6 August 2024, revised 2 April 2025): five LLMs, eight tasks, 10 runs; accuracy variations up to 15% across runs; best to worst gap up to 70%.
- [5] Anthropic, Demystifying evals for AI agents (9 January 2026): a 75% per-trial success rate gives a 42% chance of passing three trials; 20-50 simple tasks drawn from real failures as a start.
- [6] Yao, Shinn, Razavi and Narasimhan, tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (arXiv, 17 June 2024): agents succeed on fewer than 50% of tasks; fewer than 25% of retail tasks pass in all 8 attempts.
- [7] Ryseff, De Bruhl and Newberry, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (RAND Corporation, 13 August 2024): 65 interviews; most common cause of failure; 14 of 50 and 30 of 50 industry interviewees.
Related reading
An AI demo is not a product
A convincing AI demo takes an afternoon. Turning it into something people trust in production takes most of the work, and most failures happen at that stage.
Ask an AI vendor for its time to production, and skip the demo
Each of Wonderful's customer stories starts with the time it took to reach production. That time is the better proof, because it measures the vendor, while a demo only measures the model.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.
How to tell if an AI feature idea is worth building
Most AI feature ideas look good in a demo and fail during the work of making them reliable. Here are the four questions we ask to tell the ones worth building from the ones that just look good in a demo.