Strategy

What the Air Canada chatbot ruling means for a company with an AI assistant

Editorial · Reveneau · October 1, 2026

What the Air Canada chatbot ruling means for a company with an AI assistant

On 14 February 2024 the British Columbia Civil Resolution Tribunal ordered Air Canada to pay a customer $812.02 in Canadian dollars. The customer had asked the chatbot on the airline's website about bereavement fares, which are reduced fares for people who travel because of a death in the family. The chatbot described a rule that a page on the same website contradicted. The decision is Moffatt v. Air Canada, 2024 BCCRT 149.

The amount is small. The reasoning is the part that concerns a company whose AI assistant answers customers, because the tribunal treated the chatbot's answer as the airline's own statement. This post sets out what the decision says, what it leaves out, and the written test that targets this type of failure. It is general information and gives no legal advice. A lawyer should confirm what applies to your company.

What the tribunal decided

According to the decision, the customer booked flights with Air Canada in November 2022 after their grandmother died. While researching flights they used the chatbot on the airline's website. The decision quotes the chatbot's answer: "If you need to travel immediately or have already travelled and would like to submit your ticket for a reduced bereavement rate, kindly do so within 90 days of the date your ticket was issued by completing our Ticket Refund Application form."

The airline's own web page, also quoted in the decision, said that "the bereavement policy does not apply to requests for bereavement consideration after travel has been completed". The same website gave two answers to one question. The decision records that an Air Canada representative later admitted the chatbot had provided "misleading words". The customer claimed $880, which they said was the difference in price between the regular fare and the bereavement fare.

The tribunal member wrote: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate." The tribunal held that the customer had proved a claim the decision calls negligent misrepresentation, and that they were entitled to damages. The order has three parts.

Part of the order Amount
Damages $650.88
Interest up to the decision (pre-judgment interest) $36.14
Tribunal fees $125.00
Total $812.02

What the decision leaves out

Read these three limits before you repeat this case in a meeting.

First, this is a small claims decision from a tribunal in one Canadian province. It binds only its two parties, the customer and the airline. What the law says where your company operates is a question for a lawyer.

Second, the decision records no information about the chatbot's technology. The tribunal wrote that "Air Canada did not provide any information about the nature of its chatbot". Nobody can say from the decision whether the chatbot used a large language model, the type of AI model that produces text, or something older and simpler.

Third, nothing is known about how the airline tested the chatbot. The decision describes no testing, so this post makes no statement about what the airline checked before or after release. It also makes no claim that a specific test would have changed the outcome. The record shows a type of failure, and a team can write a test for it.

The tribunal treated the chatbot's answer as the company's statement

The tribunal described the airline's argument in these words: "In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions." The tribunal member called that "a remarkable submission" and rejected it. A chatbot, the decision says, "is still just a part of Air Canada's website". The same paragraph says: "It makes no difference whether the information comes from a static page or a chatbot."

The practical reading for an executive is this. A policy page is written once, and somebody reads it before it is published. An AI assistant produces a new answer each time a customer asks, and the customer is the first person to read it. If a tribunal can treat the page and the answer as the same type of statement, then both need a check before a customer sees them. The page gets its check when it is edited. The assistant's answers can be checked in advance in one way: with a set of the questions customers ask, and the correct answer written beside each one.

A second public record shows the same type of failure in a product that its source does describe as AI. On 29 March 2024 The Markup published its reporter's own tests of a chatbot that New York City had announced in October 2023 to help business owners. The article says the chatbot was "powered by Microsoft's Azure AI services". In those tests the chatbot said an employer could keep part of a worker's tips, and said a restaurant could refuse cash, which the article calls "one wholly false response". The Markup gives no count of the questions it asked, so no error rate can be stated. The city replied, as the article quotes, that the chatbot "has already provided thousands of people with timely, accurate answers".

The test that compares each answer with the policy text

The test that targets this type of failure is a policy comparison test. Engineers call a repeatable test of an AI product an eval: an input, a written expectation of what a correct output must do, and a rule that scores the output. Our guide explains what an AI eval is.

A policy comparison test is built in four steps, and the first three need no engineer.

  1. List every rule in every policy the assistant can speak about: refunds, fares, delivery, warranties, prices.
  2. For each rule, write the questions a customer would ask about it, in several wordings. Include a customer who does not qualify, such as a person asking for a refund after the deadline.
  3. Beside each question, write what a correct answer must say, and copy the policy paragraph it comes from. The person who owns the policy writes this, because they know what the company will honour.
  4. An engineer turns the list into a program that sends each question to the assistant several times and scores each answer against the policy paragraph. Any answer that grants what the policy refuses is a failure.

Here is an invented illustration with made-up numbers. Take a company that rents bicycles and has three policy documents containing 24 rules. The team writes 5 wordings for each rule, which is 24 x 5 = 120 questions. Each question is sent 3 times, because an AI model can give different answers to the same question. That makes 120 x 3 = 360 answers. Suppose 354 of them agree with the policy. 354 divided by 360 is 0.983, a pass rate of 98.3 percent.

Now read the 6 answers that failed. All 6 concern one rule: the deposit is returned when the bicycle comes back within the rental period. In each failed answer the assistant told a late customer that the deposit would be returned in full. One rule has 5 x 3 = 15 answers, so the assistant stated that rule wrongly in 6 of 15 answers, which is 40 percent. An overall rate of 98.3 percent contains a rule that failed 4 times in 10 in this run.

So a policy test needs a rule of its own, written before the test runs: one answer that grants what the policy refuses stops the release, whatever the average is. The guide to public AI failures and the tests that target each one covers five other public records.

Run the test before release and after every change

A policy test that ran once before release describes the assistant as it was on that day. NIST, the United States standards institute, says in its voluntary AI Risk Management Framework, published in January 2023: "AI systems should be tested before their deployment and regularly while in operation." Deployment is NIST's word for release.

Three types of change should each cause a full run.

The policy changes. When the policy page is edited, the expected answers in the test are edited on the same day. Otherwise the test checks the assistant against the old rule.

The instructions change. An assistant works from written instructions, called a prompt, and an edit to them can change its answers. Run every policy question again after each edit.

The AI model changes. Most assistants are built on a model that another company owns and updates. Lingjiao Chen, Matei Zaharia and James Zou measured one such update. They gave the March 2023 and June 2023 versions of GPT-4 the same 1,000 questions asking whether a number is prime. They report that accuracy "dropped from 84.0% in March to 51.1% in June", while GPT-3.5 on the same questions went from 49.6% to 76.2%. The changes went in both directions, and the figures describe 2023 models only. A service with the same name gave different answers three months later.

How to turn a result into a yes or a no is covered in using evals to decide whether an AI feature is ready to release.

What the test cannot find, and four questions for your team

A test finds the failures that somebody wrote a question for. A rule nobody listed is untested. Two practices cover part of that. When a wrong answer is found after release, add its question to the test and keep it there for every later run. And set fixed limits on what the assistant may promise, often called guardrails, which our post on AI evaluation and guardrails for production describes.

If your company has an assistant that answers customers today, ask your team four questions this week.

  1. Which policies can the assistant speak about, and where is the list?
  2. Who wrote the expected answer for each rule, and when was it last compared with the policy page?
  3. When did the policy test last run, and what has changed since then?
  4. Which failed answers stop a release, and which role decides?

If a vendor built the assistant, ask the vendor the same four, and read how to evaluate a vendor's eval suite before the meeting. A lawyer can tell you what standard of care applies where you operate. The test gives you a dated record of what was checked. The wider argument is in our guide to why AI evals matter.

Where Reveneau fits

Reveneau is an AI software development consultancy. All of our code is written by AI, and every change must pass a large eval suite, written from the specification before the code, before release. For an assistant that answers customers, we write the policy comparison cases into that suite at the start of the work. We ask every client for the policy documents and for a person who can say what each rule means, because that person writes the expected answers.

The checks that need judgment are graded by Jev, TypeSafe AI's decision model. On our own suite the run is ten times faster than with our previous language-model grader, so the suite can run on every change. Reveneau, as a company, takes responsibility for the whole project through production and after release. To plan an assistant with its policy tests written first, see AI development at Reveneau.

Sources

Common questions

What did the tribunal decide in Moffatt v. Air Canada?

In Moffatt v. Air Canada, issued on 14 February 2024, the British Columbia Civil Resolution Tribunal held the airline responsible for an answer its website chatbot gave about bereavement fares. The tribunal wrote that it makes no difference whether information comes from a static page or a chatbot, found that Air Canada did not take reasonable care to ensure its chatbot was accurate, and ordered a payment of $812.02.

Does the Air Canada chatbot decision apply to my company?

The Air Canada chatbot decision binds only its two parties, because it is a small claims decision from the Civil Resolution Tribunal in British Columbia. It shows how one tribunal reasoned about a chatbot's answer on 14 February 2024. Whether the same reasoning applies to your company depends on the law where you operate, so ask a lawyer. A practical step that needs no legal opinion is to test what your assistant says about each policy.

Was the Air Canada chatbot built on a large language model?

The decision in Moffatt v. Air Canada leaves that question open. The tribunal wrote that Air Canada did not provide any information about the nature of its chatbot, so the record cannot show what technology produced the answer. The test that targets this type of failure is the same for any assistant that answers customers, which is to ask questions about each written policy and compare every answer with the policy text.

Does the decision say how Air Canada tested its chatbot?

No. The decision in Moffatt v. Air Canada records nothing about how the airline tested its chatbot before or after release, so nobody can say from the record what testing took place. The tribunal's finding is that Air Canada did not take reasonable care to ensure its chatbot was accurate. This post therefore makes no claim that a specific test would have changed the outcome of the case.

How much did Air Canada have to pay in the chatbot case?

The tribunal ordered Air Canada to pay $812.02 in total. The order lists $650.88 in damages, $36.14 in pre-judgment interest and $125 in tribunal fees, and those three amounts add up to $812.02. The customer had claimed $880. The amounts are in Canadian dollars. The decision gives no figure for the airline's staff time or legal advice, so the full cost of the case to the airline is unknown.

What test checks that an AI assistant states company policy correctly?

A policy comparison test checks that an AI assistant states company policy correctly. List every rule in every policy the assistant can speak about, write the questions a customer would ask about each rule in several wordings, and write the correct answer beside each question with the policy paragraph it comes from. A program then sends every question to the assistant and fails any answer that grants what the policy refuses. This targets the type of failure recorded in Moffatt v. Air Canada.

When should a company run the policy comparison test?

Run the policy comparison test before release and again after every change to the policy text, the assistant's written instructions or the AI model. The voluntary AI Risk Management Framework from NIST, the United States standards institute, was published in January 2023 and says AI systems should be tested before their deployment and regularly while in operation. A result from before release describes the assistant as it was on that day.

Why does a change of AI model mean the policy test has to run again?

A change of AI model can change the assistant's answers while the product keeps the same name. Chen, Zaharia and Zou gave the March 2023 and June 2023 versions of GPT-4 the same 1,000 questions about prime numbers and measured accuracy of 84.0 percent in March and 51.1 percent in June. Those figures describe 2023 models. What still applies is to run the policy test again after each model change.

Who should write the expected answers in a policy comparison test?

The person who owns each policy should write the expected answers in a policy comparison test, because they know what the company will honour. An engineer then turns the list into a program that scores the assistant's answers. If the expected answer comes from someone who has to guess the rule, the test can pass an answer the company would refuse to honour. Review the expected answers whenever the policy page changes.

Is a high overall pass rate enough to release an AI assistant?

A high overall pass rate is insufficient when the failures all concern one policy rule. In this post's invented illustration, 354 of 360 answers pass, which is 98.3 percent, and all 6 failures concern one rule that the assistant states wrongly in 6 of 15 answers. Write the rule down in advance that any answer granting what the policy refuses blocks the release, whatever the average is.