Air Canada Argued Its Chatbot Was a Separate Legal Entity
A tribunal called it "a remarkable submission" — and in eight hundred dollars' worth of small claims, set the standard every hotel chatbot now has to meet.
In February 2024 the British Columbia Civil Resolution Tribunal decided a case worth CAD $812.02. A passenger had asked Air Canada's website chatbot about bereavement fares. The chatbot told him something that was not the airline's policy. He relied on it.
Air Canada's defence is the part worth reading. The airline argued it could not be held liable for information provided by its chatbot. The tribunal member's response, at paragraph 27:
Air Canada argues it cannot be held liable for information provided by one of its agents, servants, or representatives — including a chatbot. It does not explain why it believes that is the case. In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission.
Moffatt v. Air Canada, 2024 BCCRT 149
And then, plainly: "It makes no difference whether the information comes from a static page or a chatbot."
The standard the tribunal applied is not novel, which is exactly what makes it significant. It is ordinary negligence: "the applicable standard of care requires a company to take reasonable care to ensure their representations are accurate and not misleading." The finding: "I find Air Canada did not take reasonable care to ensure its chatbot was accurate."
The award was trivial. The doctrine is that a chatbot has no separate legal existence from the business running it, and its accuracy is governed by the same duty as everything else you publish.
This is a small-claims tribunal decision, not binding appellate precedent, and anyone citing it should say so. Its weight is doctrinal, not financial. But if you run a hotel, it answers a question you may not have realised you had an answer to: when your chatbot quotes a rate, you quoted that rate.
Which brings us to the only question hotel guests actually ask
Guests ask a great many things. Overwhelmingly, they ask one: what does it cost.
A room rate is not general knowledge. It exists in one place, changes daily, and depends on dates, occupancy, child ages, minimum-stay rules and which rate plan applies. It has no correct answer outside the reservation system holding it.
Ask a language model for it without access to that system and it will still produce a number, because producing plausible text is what it does. Not from malice or a bug — that is the machine working as designed.
Hallucination is not a defect that better models remove
The common assumption is that this was an early-model problem, now largely fixed. The current measurements say otherwise.
Stanford's AI Index Report 2026 reports hallucination rates across 26 current models ranging from 22 to 94 per cent on the AA-Omniscience benchmark — 6,000 questions across six domains. Not 2021 models. Current ones.
There is also a formal result. Kalai and Vempala, at the ACM Symposium on Theory of Computing in 2024, showed "an inherent statistical lower-bound on the rate that pretrained language models hallucinate certain types of facts, having nothing to do with the transformer LM architecture or data quality."
The limit on that result matters and is usually dropped when it gets quoted. The bound applies to facts appearing once or rarely in training data. The authors are explicit that there is no statistical reason for hallucination on facts that appear repeatedly.
So it does not mean models always hallucinate. It means they must hallucinate on exactly the class of fact a hotel cares about: a rate that exists in one database, for one property, on one night.
The architectural difference, which is the whole thing
Two systems, same model, same guest question — "family room, three nights in July, two adults and a child of nine, how much?"
The first has been trained on the hotel's website and FAQs. It reads the question correctly, then generates a number that looks like what a rate for a hotel like this would plausibly be. Nothing checks it, because there is nothing to check it against. The guest receives a price the hotel never set — along with invented minimum-stay rules and an invented child-age bracket.
The second calls the reservation system. An authenticated query returns live availability, the applicable rate plan, the minimum stay and the child-age policy. The answer is constructed from what came back. If the record does not cover the case, the correct output is a clarifying question — not a number. And the booking link encodes the same parameters as the quote, so the two cannot diverge.
The research literature has a name for this distinction. Lewis and colleagues, introducing retrieval-augmented generation at NeurIPS 2020, framed it precisely: language models' "ability to access and precisely manipulate knowledge is still limited," and "providing provenance for their decisions and updating their world knowledge remain open research problems." Rates change nightly. A model relying on stored parameters is stale by construction and cannot say where its answer came from.
Shuster and colleagues tested the effect directly in a paper whose title is the finding: retrieval augmentation reduces hallucination in conversation. No percentage appears in that paper, and we do not quote one.
Grounding is not a complete answer either. Models can hallucinate the API call, not just the prose, and the Berkeley Function Calling Leaderboard's conclusion at ICML 2025 is that while models "excel at single-turn calls, memory, dynamic decision-making, and long-horizon reasoning remain open challenges." A booking conversation is multi-turn and stateful. That is the hard case, not the easy one.
Why hotels want any of this: the arithmetic nobody disputes
Set the technology aside. A week contains 168 hours. A full-time post covers roughly 40. To have one person available at every hour of every day therefore requires 4.2 full-time posts — before annual leave, before sickness, before breaks, and before any allowance for two guests writing at once.
Eurostat puts 2025 hourly labour cost in accommodation and food service activities at €20.9 across the EU-27, €19.0 in Greece, €17.6 in Spain and €12.9 in Portugal. Applied to the 728 coverage hours in an average month, round-the-clock single-person cover costs between €9,391 and €15,215 a month.
Notice what that argument does not require: any assumption about how many messages a person can handle per hour. We looked for a credible independent benchmark for that and could not find one — the figure in general circulation comes from a chat-software vendor and implies fewer than two conversations an hour. So our report makes no headcount-equivalence claim anywhere. The coverage arithmetic needs none.
The statistic your competitors are misquoting
Nearly every article on hotel enquiry response times cites a 2011 Harvard Business Review study for the claim that replying within an hour makes you "seven times more likely to convert."
That is not what it says. The actual finding: firms contacting a lead within an hour were nearly seven times as likely to qualify that lead as firms contacting it one hour later — and more than sixty times as likely as firms waiting twenty-four hours or more.
The seven-fold comparison is one hour against two. The outcome is qualification, not a sale. And the sixty-fold figure — the genuinely dramatic one, and the one that matters for an enquiry arriving overnight — is the one nobody quotes.
The same audit found something more useful than either. Of 2,241 companies sent a test enquiry, 37 per cent responded within an hour, 24 per cent took more than a day, and 23 per cent never responded at all. Among those that did respond, the average was 42 hours.
The opportunity in that data is not being fast. It is being present at all.
We should add the honest caveat our own report carries: we searched for peer-reviewed research linking hotel enquiry response time to booking conversion, and found none. Every result was a hospitality-technology vendor recycling the 2011 figures. The transfer from cross-industry sales to a guest asking about a room rate on WhatsApp is a reasonable assumption. It is not a demonstrated result.
The evidence that cuts against us
We sell this category of software, so here is the research that argues against it.
Wüst and Bremser ran a preregistered experiment with 467 German participants on hotel booking support. Booking intention was significantly lower with an AI chatbot than with a human agent, and fell more steeply for chatbots than humans when the booking situation turned negative. Their recommendation: complaints "should offer the option to be connected to a person." Germany is Türkiye's second-largest inbound market at 6.75 million visitors in 2025. This is not a distant finding.
And disclosure has a measured cost. From 2 August 2026, Article 50 of the EU AI Act requires that people be informed they are interacting with an AI system. A field experiment across more than 6,200 customers, published in Marketing Science, found that disclosing chatbot identity before the conversation "reduces purchase rates by more than 79.7 per cent" — because customers perceived the disclosed bot as less knowledgeable and less empathetic.
That study is from 2019, covers outbound sales calls in China, and predates current models. The magnitude will not replicate. But the mechanism is the point: the penalty ran through perceived lack of knowledge. Once disclosure is compulsory, being demonstrably knowledgeable is the only lever left.
What we measured in our own deployment
Across two hotel properties under common ownership, running continuously for around seven months, our system handled roughly 424,000 guest messages. In a representative 30-day window it handled 60,611 — about 2,020 a day, or 84 an hour around the clock — resolving 99 per cent without human intervention, ending 64.8 per cent of conversations with a guest-specific booking link, at an average first response of 7.9 seconds. Every price quoted was checked against the hotels' live reservation system before being sent.
What that does not establish. No control group, so no causal claim about bookings. No revenue attribution, so no return-on-investment figure appears anywhere in our report. No guest-satisfaction instrument. No independent audit. Two properties under one ownership is not a sample. Read it as an existence proof that the architecture works at volume — not as evidence of what it is worth.
Set the 7.9 seconds against the 42-hour average above if you like. But the comparison that matters at three in the morning is not fast against slow. It is answered against unanswered — and 23 per cent of the companies in that audit were in the second category.
Full reportThe Grounded Front Desk: Guest Communication, Response Latency and Factual Reliability in Hotel Operations, 2026 — 22 pages, 10 figures, 53 cited sources including peer-reviewed work from Marketing Science, Journal of Consumer Research, Manufacturing & Service Operations Management, ACM Computing Surveys and the proceedings of NeurIPS, ICLR, ACL and EMNLP, alongside Eurostat, the US Bureau of Labor Statistics, Stanford HAI, NIST, EUR-Lex, HOTREC and the American Hotel & Lodging Association. Read the full report →
The report ends with a ten-point specification a hotel can put straight into a procurement document. Every requirement traces to a cited finding, and three of them exist because of evidence that cuts against AI guest service rather than for it.
RelatedHow this works for hotels · Grounding, escalation and refusal behaviour · Channels: WhatsApp, Instagram, Messenger and others · Published pricing
Corrections are welcome and will be published. If you can supply peer-reviewed research linking hotel enquiry response time to booking conversion, we will update the report and credit you.
Try it on your own business
Set up Vera.Support and start answering every customer — on every channel, in every language — without growing your team.
Get Started