Why AI evals matter: the test that shows whether an AI product works / What goes wrong without evals
Public AI failures and the tests that target each one
Six public records show four types of AI failure that a test can target. A tribunal ordered Air Canada to pay $812.02 after its chatbot stated a rule the airline's policy page contradicted. A court imposed $5,000 on lawyers who filed invented court opinions. The Markup reported a New York City chatbot giving answers that conflict with city law. Regulators challenged advertising claims by Workado, DoNotPay and Pieces. A written test, called an eval case, targets each type: policy questions scored against the policy text, a lookup of every cited source, legal questions with expectations written by a specialist, and a test set built from the content real customers send.
Published September 30, 2026. Editorial.
Key takeaways
- On 14 February 2024 the British Columbia Civil Resolution Tribunal ordered Air Canada to pay $812.02 and wrote that it makes no difference whether information comes from a static page or a chatbot.
- On 22 June 2023 a United States federal court imposed a $5,000 penalty on two lawyers and their firm for filing six court opinions that an AI tool had invented.
- The FTC's complaint against Workado alleged a product advertised as 98 percent accurate scored 53 percent on general content in an independent test. The final order is dated 28 August 2025.
- Four test types target these failures: policy consistency cases, a source existence check, legal questions written with a specialist, and a test set that matches the content customers send.
- A test finds only the failure types somebody wrote a case for, so each new failure seen in real use is added to the eval set as a permanent case.
On 14 February 2024 the British Columbia Civil Resolution Tribunal ordered Air Canada to pay a customer $812.02 in Canadian dollars. The chatbot on the airline's website had described a refund rule that the airline's own policy page contradicted [1]. On 22 June 2023 a federal court in New York imposed a $5,000 penalty on two lawyers and their law firm for filing a document that cited six court opinions which did not exist [2].
Each record describes a type of failure that a team can write down as a test before release. This page covers six cases: what the record says, the type of failure in plain words, and the eval case that targets that type. An eval is a repeatable test: a set of inputs, a written expectation for each one, and a way to score the output. The full definition is in what an AI eval is.
How to read the cases on this page
Every fact below comes from the document cited beside it, read on 30 September 2026. Three limits apply.
- No record here describes how the organisation tested its product, so this page makes no statement on that subject.
- Several records are settlements, where the facts are the regulator's allegations and no court ruled on them.
- Each test named here targets a type of failure. Whether it would have changed the outcome in the named case is unknown.
This page is general information. A lawyer should confirm what applies to your product.
The six cases in one table
| Case | Record | Failure type | Test that targets it |
|---|---|---|---|
| Moffatt v. Air Canada | Tribunal decision, 14 February 2024, $812.02 [1] | Stated a rule the company's policy page contradicted | Policy questions scored against the policy text |
| Mata v. Avianca | Court order, 22 June 2023, $5,000 [2] | Cited court opinions that did not exist | Every cited source looked up in a trusted database |
| New York City business chatbot | Report by The Markup, 29 March 2024 [3] | Gave answers that conflict with city law | Legal questions with expectations written by a specialist |
| Workado | FTC final order, 28 August 2025 [5] [6] | Advertised accuracy that, the FTC alleged, an independent test contradicted | A test set made of the content customers send |
| DoNotPay | FTC final order, 11 February 2025, $193,000 [4] | Made claims about a chatbot's abilities that the FTC called deceptive | Outputs compared with a qualified person's work |
| Pieces | Texas settlement, 18 September 2024 [7] | Advertised an error rate the investigation found likely inaccurate | An error rate traced to a documented test set |
FTC is the United States Federal Trade Commission, the federal consumer protection regulator.
A chatbot stated a policy that the company's own page contradicted
What the record says. In November 2022 a customer used the chatbot on Air Canada's website while booking flights after a death in their family. The decision quotes the chatbot telling them they could submit a ticket for a reduced bereavement rate "within 90 days of the date your ticket was issued". The airline's own web page said the bereavement policy "does not apply to requests for bereavement consideration after travel has been completed" [1].
The tribunal wrote: "It makes no difference whether the information comes from a static page or a chatbot." Its finding reads: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate." It held that the customer had proved their claim of "negligent misrepresentation" and was entitled to damages. The order was $650.88 in damages, $36.14 in interest and $125 in tribunal fees, which adds up to $812.02 [1].
A small claims tribunal decision binds only the two parties in it. The decision also states that Air Canada "did not provide any information about the nature of its chatbot", so the record does not show whether the chatbot used a language model.
The failure type. The product told a customer a rule that the company's written policy contradicted.
The test that targets it. A policy consistency set. Write one eval case for each rule in each policy document. The input is a customer question about that rule, in several wordings. The expectation is that the answer agrees with the policy text. The grader, which is the program or person that scores an output, fails any answer that grants what the policy refuses.
Take an invented furniture shop with a 30-day return rule. The input is "I bought the sofa 40 days ago, can I still send it back?" An answer that offers a return fails.
A tool invented court opinions and lawyers filed them
What the record says. The sanctions order in Mata v. Avianca, signed on 22 June 2023 by a federal court in New York, states that the lawyers "submitted non-existent judicial opinions with fake quotes and citations created by the artificial intelligence tool ChatGPT", and that they kept defending those opinions after the court's orders questioned whether the opinions existed [2]. The order names six such opinions. The penalty was $5,000, imposed jointly on two lawyers and their firm.
The same order says "there is nothing inherently improper about using a reliable artificial intelligence tool for assistance" [2]. The sanction was for the lawyers' conduct. The tool was a public chat product that nobody in the case had built.
The failure type. The output named a source that does not exist. The United States National Institute of Standards and Technology (NIST) calls this confabulation: "the production of confidently stated but erroneous or false content" [8]. The common name is hallucination, a statement the model made up.
The test that targets it. A source existence check. A program takes every court case, law, article number and web address out of each output and looks each one up in a database the business trusts. Any source that cannot be found fails the case. Ordinary code does the scoring, so the check can run on every output. NIST's generative AI profile lists this as suggested action MS-2.5-003: "Review and verify sources and citations in GAI system outputs during pre-deployment risk measurement and ongoing monitoring activities" [8]. GAI is NIST's short form for generative AI.
A city chatbot gave answers that conflict with city law
What the record says. On 29 March 2024 The Markup published the results of its reporter's own tests of a chatbot that New York City had announced in October 2023 to help business owners [3]. The article reports these answers:
- Asked whether landlords must accept tenants who pay with housing vouchers, "the bot said no, landlords do not need to accept these tenants. Except, in New York City, it's illegal for landlords to discriminate by source of income".
- The bot replied: "Yes, you can take a cut of your worker's tips." In plain words, it said an employer may keep part of a worker's tips.
- The bot said a restaurant could refuse cash, which the article calls "one wholly false response".
The Markup also wrote: "Similar errors appeared when the questions were asked in other languages". The city's reply, quoted in the article, was that the chatbot was a pilot programme that "has already provided thousands of people with timely, accurate answers" [3].
This record is journalism. The article gives no count of the questions asked and no share of wrong answers, so no error rate can be stated. It describes the chatbot in March 2024, and this page checked nothing about what happened to it afterwards.
The failure type. As The Markup reports it, the product gave advice that would lead a user who followed it to break the law.
The test that targets it. A set of legal questions where the lawful answer is settled, with each expectation written by a person who knows the law. Any output that tells the user they may do the prohibited thing fails. Mark these cases as blocking: one failure stops the release whatever the average pass rate is. The method is in using evals to decide a release.
Run the same cases in every language the product supports, and run each one several times, because a language model can give different output for the same input. Why a good AI demo does not show the product works explains that variation.
Three companies made advertising claims that regulators challenged
Workado. The FTC's release of 28 April 2025 says the company promoted its AI Content Detector as 98 percent accurate at telling whether a person or an AI wrote a text. The release continues: "independent testing showed the accuracy rate on general-purpose content was just 53 percent, according to the FTC's administrative complaint". The FTC alleged that the model had been trained to classify academic content only [5]. The final order, approved on 28 August 2025, requires the company to stop advertising accuracy unless it has "competent and reliable evidence showing those products are as accurate as claimed" [6]. The company settled. The release leaves out who ran the independent test and how many texts it used.
DoNotPay. The FTC's release of 11 February 2025 says the final order requires the company "to stop making deceptive claims about the abilities of its AI chatbot". The order requires a payment of $193,000 and prohibits the company from "advertising that its service performs like a real lawyer" unless it has sufficient evidence for the claim [4]. This is a consent order: the company settled and no court ruled on the facts.
Pieces. On 18 September 2024 the Office of the Attorney General of Texas announced a settlement with Pieces. The release says at least four Texas hospitals had been giving the company patient data so that its generative AI product could summarise patients' condition and treatment for hospital staff. It says the company advertised a "severe hallucination rate" of "<1 per 100,000". The investigation "found that these metrics were likely inaccurate and may have deceived hospitals about the accuracy and safety of the company's products" [7]. The release reports no court finding and no reply from the company.
Older lists of FTC cases about AI include a 2024 consent order involving Rytr LLC. The FTC reopened and set aside that order on 22 December 2025 [9], so this page leaves it out. The DoNotPay and Workado orders are recorded as final in the FTC releases read on 30 September 2026.
The failure type. In each regulator's account, a number or a claim in the company's advertising was inaccurate or deceptive, or was likely to be.
The tests that target it.
- A representative test set. Take the cases from the types of input real customers send, and report the pass rate for each type separately. The Workado allegation describes a model built for one type of text and sold for another.
- A comparison with a qualified person. Give the same cases to the product and to a professional, then have a second professional score both sets without knowing which is which.
- A documented error rate. Any published rate comes with the number of cases, their origin, and the written definition of an error. How to read an AI eval report shows what each of those tells a buyer.
Who is responsible for what the product says
The tribunal in Moffatt wrote that it "should be obvious to Air Canada that it is responsible for all the information on its website" [1]. That is one tribunal in one province. A practical reading for a product owner is to plan as if every answer the product gives is a statement by the company, and to ask a lawyer what the law says where you operate.
What a test cannot find
A test finds the failure types that somebody wrote a case for. A new type of failure appears first in real use. NIST's profile says "Measurement gaps can arise from mismatches between laboratory and real-world settings" [8].
Two practices reduce that gap. Read a sample of real conversations every week and sort the failures by type, the method described in error analysis before metrics. Then add each new failure to the eval set as a permanent case. Fixed limits on what the product may say or do, often called guardrails, are a separate protection, covered in AI evaluation and guardrails for production.
How to turn these cases into your first tests
- List every written policy, price and rule the product can speak about. Write one case per rule.
- List every type of source the product can name. Add an existence check for each type.
- Ask a person who knows the law which questions have a wrong answer that breaks a rule. Mark those cases as blocking.
- Write down each claim in your advertising. Beside each one, name the test set that supports it.
- Repeat the set in each language you support.
The wider argument for doing this before release is in why AI evals matter.
How Reveneau applies this
At Reveneau all code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before it is released. For an AI product, we write the four types of case on this page into that suite at the start of the work: policy consistency, source existence, prohibited advice, and a test set that matches what users send. We ask every client for their policy documents and for access to a person who knows the rules that apply to the product, because that person writes the expectations a model cannot guess.
The judged checks in the suite are graded with Jev, TypeSafe AI's decision model. On our own suite the run is ten times faster than with our previous language-model grader.
Reveneau, as a company, takes responsibility for the whole project through production and after release. A failure found after release becomes a permanent case in the suite. To discuss a product that speaks to your customers, see AI development at Reveneau or contact us.
Common questions
Is a company liable for what its chatbot says to a customer?
In Moffatt v. Air Canada, decided on 14 February 2024, the British Columbia Civil Resolution Tribunal held the airline responsible and wrote that it makes no difference whether information comes from a static page or a chatbot. That decision binds only the two parties and comes from one small claims tribunal in one province. A lawyer should confirm what the law says for your product and your country.
What was the Air Canada chatbot case about?
The Air Canada chatbot case, Moffatt v. Air Canada, concerned a customer who booked flights in November 2022 after a death in their family. The chatbot told them they could apply for a reduced bereavement rate within 90 days of the ticket being issued, and the airline's policy page said the opposite. The tribunal found negligent misrepresentation and ordered $812.02: $650.88 in damages, $36.14 in interest and $125 in fees.
Was the Air Canada chatbot built on a large language model?
The record does not say. The tribunal's decision of 14 February 2024 states that Air Canada did not provide any information about the nature of its chatbot, so nobody can say from the decision what technology it used. The test that targets the failure type is the same for any chatbot: ask questions about each written policy and score every answer against the policy text.
What happened in Mata v. Avianca?
In Mata v. Avianca a United States federal court in New York sanctioned two lawyers and their law firm on 22 June 2023. The order states that they submitted non-existent judicial opinions with fake quotes and citations created by ChatGPT, then kept defending those opinions after the court questioned whether they existed. The penalty was $5,000. The same order says there is nothing inherently improper about using a reliable artificial intelligence tool for assistance.
What did the New York City business chatbot get wrong?
The Markup reported on 29 March 2024 that the New York City business chatbot, in the reporter's own tests, said landlords need not accept tenants with housing vouchers, said an employer could keep part of workers' tips, and said a restaurant could refuse cash. The article describes these answers as conflicting with city law. The article gives no error rate, and it describes the chatbot as it was in March 2024.
What kinds of AI failure can a test find before release?
A test can find any failure type that somebody has written a case for. The six public records on this page show four such types: stating a rule the company's policy contradicts, naming a source that does not exist, giving advice that conflicts with the law, and advertising accuracy that a test on real content does not support. A failure type nobody has thought of appears first in real use and is then added to the set.
How do I test that a chatbot states our company policies correctly?
Write one eval case for each rule in each policy document. The input is a customer question about that rule in several wordings, and the expectation is that the answer agrees with the policy text. A grader, the program or person that scores an output, compares each answer with the policy paragraph and fails any answer that grants what the policy refuses. This policy consistency set targets the failure type recorded in Moffatt v. Air Canada.
How do I check that an AI product names only real sources?
Use a source existence check. A program takes every case name, law, article number and web address out of each output and looks it up in a database the business trusts, and any source that cannot be found fails. Ordinary code can do this on every output. NIST's generative AI profile lists it as suggested action MS-2.5-003, which says to review and verify sources and citations before deployment and during monitoring.
What did the FTC require of Workado and DoNotPay?
The FTC's final order against Workado, approved on 28 August 2025, requires the company to stop advertising the accuracy of its AI detection products unless it has competent and reliable evidence for the claim. The final order with DoNotPay, announced on 11 February 2025, requires a $193,000 payment and prohibits advertising that the service performs like a real lawyer without sufficient evidence. Both are consent orders, so the companies settled and no court ruled on the facts.
How much money did the failures on this page cost the organisations involved?
The records on this page state three amounts. The tribunal ordered Air Canada to pay $812.02 in Canadian dollars. The court in Mata v. Avianca imposed a $5,000 penalty. The FTC's final order required DoNotPay to pay $193,000. The Texas release on Pieces reports no penalty amount. No record on this page gives the cost of lost customers, of staff time or of legal advice, so the full cost of each failure is unknown.
Can testing guarantee that an AI product stays free of public failures?
No. Testing finds the failure types that somebody wrote a case for, and a new type of failure appears first in real use. NIST's generative AI profile says measurement gaps can arise from mismatches between laboratory and real-world settings. The working answer is to read a sample of real conversations every week, sort the failures by type, and add each new one to the eval set as a permanent case.
Which test should a small company write first after reading these cases?
A small company should start with the policy consistency set, because the material already exists. List every written policy, price and rule the product can speak about and write one case per rule. Then ask a person who knows the law which questions have a wrong answer that breaks a rule, and mark those cases as blocking, so one failure stops the release whatever the average pass rate is.
References
- [1] British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, 2024 BCCRT 149 (14 February 2024): the chatbot's wording, the policy page wording, the separate legal entity argument, the reasonable care finding, and the order of $650.88 damages, $36.14 interest and $125 fees, $812.02 in total.
- [2] United States District Court, Southern District of New York, Mata v. Avianca, Inc., Opinion and Order on Sanctions, case 1:22-cv-01461 (22 June 2023): the $5,000 penalty on two lawyers and their firm, the six non-existent opinions, and the court's statement on using a reliable AI tool.
- [3] The Markup, NYC's AI Chatbot Tells Businesses to Break the Law (29 March 2024): the reporter's test results on housing vouchers, tips and cash, the errors in other languages, the city's reply and the page's warning text.
- [4] US Federal Trade Commission, FTC Finalizes Order with DoNotPay That Prohibits Deceptive 'AI Lawyer' Claims (11 February 2025): the $193,000 payment, the notice to subscribers from 2021 to 2023, and the evidence requirement for claims that the service performs like a real lawyer.
- [5] US Federal Trade Commission, FTC Order Requires Workado to Back Up Artificial Intelligence Detection Claims (28 April 2025): the 98 percent claim, the 53 percent independent test result alleged in the complaint, and the allegation that the model was trained for academic content.
- [6] US Federal Trade Commission, FTC Approves Final Order against Workado, LLC (28 August 2025): final approval by a vote of 3 to 0 and the requirement for competent and reliable evidence behind accuracy claims.
- [7] Office of the Attorney General of Texas, Attorney General Ken Paxton Reaches Settlement in First-of-its-Kind Healthcare Generative AI Investigation (18 September 2024): the advertised error rate, the investigation's finding that the metrics were likely inaccurate, the four hospitals, and the settlement term on disclosing accuracy.
- [8] NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024): the definition of confabulation, suggested action MS-2.5-003 on verifying sources and citations, and the statement on gaps between laboratory and real-world settings.
- [9] US Federal Trade Commission, FTC Reopens and Sets Aside Rytr Final Order (22 December 2025): the 2024 consent order involving Rytr LLC was reopened and set aside.
Related reading
What the Air Canada chatbot ruling means for a company with an AI assistant
A British Columbia tribunal treated a chatbot's answer as the airline's own statement and ordered a payment of $812.02. This post sets out what the decision says, what it leaves out, and the written test that compares an assistant's answers with the policy text.
AI evaluation and guardrails for production: how to know your AI actually works
"It seems to work" is not a production standard. Here is how we turn that feeling into a number, and how we stop the system from doing damage on the days the number is bad.
How to test a feature you cannot fully specify
Some features have no single correct output. A summary, a ranking, a suggested reply. You cannot check those against one exact expected answer, and the usual conclusion, that they cannot be tested, is wrong.
The real cost of shipping unverified code
The cost of unverified code does not arrive as a bug report. It arrives as a codebase nobody will touch, a review queue that never empties, and a team that has stopped trusting its own pipeline.
How to evaluate a vendor's eval suite
"We test everything" means nothing until you know who wrote the tests, what they check, and who is allowed to grade them.