Disclosure: DROZlegal publishes this guide as part of its free Lawyer AI Academy and builds a practice-automation product for Canadian law firms. The hallucination-rate and court-sanctions statistics below are drawn directly from Stanford RegLab's peer-reviewed research and Damien Charlotin's public AI Hallucination Cases Database, fetched and verified for this article, not ours.
What a large language model actually does
Strip away the marketing and a large language model does one thing: given the text so far, it predicts the most statistically likely next word, then the next, then the next — using patterns learned from an enormous body of training text. It does this so fluently that the output reads like reasoned analysis. It is not performing legal reasoning the way you do; it is pattern-completing text that resembles what it saw during training.
That's the core difference from the deterministic legal-tech automation your firm already runs. A document-assembly template fills in a client's name and produces the same clause every time. A calendaring rule always adds the same number of days to a limitation date. A large language model doesn't work that way — ask it the same legal question twice and you can get two differently worded, occasionally differently reasoned, answers, because it's sampling from a probability distribution over plausible next words, not executing a fixed rule.
That single fact — probable next word, not verified fact — explains most of what follows in this lesson.
Three failure modes every lawyer needs to recognize
The Law Society of Ontario's white paper on generative AI names hallucinations and inaccurate information directly as a key risk of the technology, alongside the unanticipated spread of confidential information. Three specific patterns show up often enough to be worth naming individually.
- Hallucination. The model states something false with the same fluent confidence as something true — a case that doesn't exist, a quote a court never wrote, a statute section that says something different from what's cited. A peer-reviewed Stanford RegLab study (Magesh et al., Journal of Empirical Legal Studies, 2025) tested leading legal AI research tools and found Lexis+ AI hallucinated on 17% of queries and Westlaw AI-Assisted Research on roughly 33% — with a general-purpose model (GPT-4, no legal-specific retrieval) hallucinating on 43% of the same query set.
- Staleness. A model only knows what was in its training data up to a fixed cutoff date. Anything after that — a new decision, a statutory amendment, a rule change — simply isn't there unless the tool is connected to a live search or database, and even retrieval-augmented tools are only as current as what they've indexed.
- Confident wrongness. These models don't hedge the way a careful colleague would when unsure. Wrong output reads exactly as fluent and certain as right output, which is precisely why courts keep catching it after the fact rather than lawyers catching it before filing.
Always independently verify any information produced by generative AI that you intend to rely on. The verification process should be completed by a human being, not the AI system itself. Law Society of Ontario, Generative AI: Your Professional Obligations, April 2024
That verification duty isn't theoretical. Damien Charlotin's public AI Hallucination Cases Database, which tracks documented court filings containing AI-fabricated citations worldwide, grew from roughly 100 cases in mid-2025 to nearly 1,900 by August 2026 — a trend line moving the wrong way even as judges have issued repeated public warnings.
Where AI is reliable — and where it needs a skeptical human every time
Not every task carries the same risk profile. The table below maps the categories this Academy curriculum keeps coming back to.
| Task category | Reliability today | Why |
|---|---|---|
| Summarizing a document already in front of you (a contract, a discovery production, a transcript) | Reliable with a quick read-through | Grounded in text you supplied; far less room to invent facts than open-ended generation |
| First-draft plain-language explanations or correspondence | Reliable as a starting point | Low stakes, and errors are easy for you to catch before anything sends |
| Legal research and case-law citation | Needs a skeptical human every time | Stanford RegLab found 17–33% hallucination rates in dedicated legal research tools tested |
| Predicting outcomes or answering "will I win" | Needs a skeptical human every time | No verified retrieval behind the prediction; confident wrongness is the default failure mode |
| Anything that reaches a client or a filing without your review | Needs a skeptical human every time | Matches the LSO's competence (Rule 3.1-2) and supervision (Rule 6.1-1) duties directly |
The next lesson in this curriculum — Module 2, on your professional obligations — walks through what these failure modes mean for competence, confidentiality, and supervision duties under Ontario's rules. If you'd rather not wait for it to publish on its own schedule, join the newsletter and it lands in your inbox the day it's live.
Why this isn't just a technology question
Every failure mode above traces back to the same root cause: a language model predicts plausible text, it doesn't verify facts, and nothing in how it was built stops it from being wrong with total confidence. That's a mechanical property of the technology, not a bug a future model version necessarily fixes away — which is exactly why the LSO's guidance puts the verification duty on you, not on the tool.
Because verification is a human job, not a system feature, the tools worth adopting are the ones that make that human check easy rather than optional. DROZlegal's agents, for example, are built as named, observable jobs that finish with a readable audit trail rather than a black-box chat response — not because the underlying model has stopped hallucinating, but because the workflow around it still assumes it can.
Once these fundamentals are in place, our four-question safety checklist for evaluating any AI vendor turns them into a practical audit before you sign anything. For a category-by-category look at what's actually on the market this year, see our job-by-job buyer's guide to AI tools for lawyers in 2026. And if generative AI specifically — drafting, summarizing, first-pass research — is the use case you're weighing, our plain-English guide to generative AI for lawyers covers where it earns its keep and where it doesn't. If the compliance question on your mind is more basic than any of that — whether using AI at all risks crossing into practising law without a licence — our short, direct answer on AI and unauthorized practice of law in Ontario covers exactly that.
Frequently asked questions
How can lawyers use AI without getting into trouble? Treat every AI output as a first draft from a fast, occasionally confident-wrong research assistant, not a verified answer, and read everything it produces before it reaches a client or a filing. That is the independent-verification duty the Law Society of Ontario's April 2024 guidance describes explicitly, and it applies regardless of which tool you use.
What exactly is a hallucination in AI? A hallucination is fabricated information — a case that doesn't exist, a quote a court never wrote, a statute section that says something different from what's cited — presented with the same fluent confidence as accurate output. A peer-reviewed Stanford RegLab study found this happened in 17% to 33% of queries even to legal-specific AI research tools built with retrieval, and more often in general-purpose models with no legal grounding at all.
Does staleness mean AI is always out of date? Not always, but a model only knows what was in its training data up to a fixed cutoff, and any legal development after that point simply isn't there unless the tool is connected to a live search or database. Even retrieval-augmented legal tools are only as current as what they've indexed, so confirming a citation's currency is still worth a human's five minutes.
Will AI get accurate enough that I can stop double-checking it? Not on the evidence so far. The public database tracking AI-fabricated citations in court filings grew from roughly 100 documented cases in mid-2025 to nearly 1,900 by August 2026, even as courts explicitly warned lawyers about the risk — independent verification is a standing professional duty, not a stopgap until the technology matures.
Read more from the Academy: Lawyer AI Academy hub, or continue to Module 2: Your Professional Obligations.
Not ready to subscribe? Join the DROZlegal waitlist instead.