Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Hallucinates a Legal Citation in Court
Transcript
- Lucas: A federal judge in New York sanctions an attorney for citing cases that don't exist. The lawyer admits he didn't read the cases — he had an AI research tool generate them. This is not a hypothetical. It happened in March of this year. Luna: And this isn't the first time. There was a big case in 2023 where an attorney used ChatGPT and cited fake cases. But it's happening again, which is worrying — especially with so many firms now rushing to deploy legal AI tools. Lucas: Right, the 2023 case where a lawyer submitted a brief with citations to cases like 'Varghese v. China Southern Airlines' that simply never existed. That was a lesson many thought the profession had learned. But here we are, three years later, with a fresh sanction order. And the judge wasn't lenient — she ordered the attorney to pay the opposing party's legal fees and attend continuing education on legal research. Luna: That's a tangible consequence. What exactly happened in this March case? Lucas: So the attorney was representing a client in a securities fraud case. He used an AI tool — not naming the vendor, but it's marketed specifically for legal research — to draft a motion to dismiss. The tool returned a half-dozen case citations that looked perfectly real. They had docket numbers, judge names, even parenthetical explanations. But when opposing counsel tried to pull the cases, they weren't in any database. They were entirely fabricated by the model. Lucas: This is a classic example of what AI researchers call 'confabulation' — the model's not lying, it's just generating text that looks like a citation because that's what it's been trained to do. But in a legal context, the effect is the same as perjury. Luna: And that's the key distinction you're drawing — that calling it 'hallucination' almost makes it sound like a glitch, when really it's a predictable output of how these models work. They don't know what a 'true' citation is. They just know what a citation looks like probabilistically. Lucas: Exactly. The term 'hallucination' is misleading because it implies the model is seeing something that isn't there. A better frame is that the model is generating text that fits a pattern, without any connection to ground truth. A big language model has no internal database of real cases. It's predicting the next word based on billions of examples from the internet — many of which are themselves inaccurate. Lucas: And this is why the 'garbage in, garbage out' maxim is only part of the problem. Even with high-quality training data, the model can invent things. It's inherent to the architecture. Luna: So if the model's design makes this inevitable, the responsibility falls on the lawyer who uses the tool without verifying. But how much of that responsibility is being codified? Are bar associations stepping in? Lucas: They are. The American Bar Association's AI task force issued draft recommendations in April — right after this sanction came down. One key proposal is a mandatory disclosure rule: any filing that contains ai generated text must be flagged, and the attorney must certify that they have independently verified all citations. Another proposal is a continuing legal education requirement specifically on AI literacy. Lucas: But here's the tension — law firms are under enormous pressure to adopt AI for efficiency. Every major Am Law 100 firm has a pilot program. Associates are told to use these tools to draft briefs, memos, even contracts. The promise is that AI cuts research time from hours to minutes. The risk is that speed replaces rigor. Luna: And that's not just a law firm problem. We're seeing similar issues in journalism, in academic research, in medical writing. Any domain where citing sources is essential, ai generated text introduces a new category of error. But in court, the stakes are uniquely high — you can lose a case, you can be sanctioned, you can even face malpractice claims. Lucas: And we already have a real consequence from the UK. In February, a tribunal in London struck out a party's entire submission because the solicitor used an AI tool to draft a witness statement, and the AI had inserted a fake quote from a nonexistent parliamentary debate. The tribunal ruled that the submission was 'fundamentally unreliable' and awarded costs against the firm. Luna: That's significant because it's not just about citations — it's about fabricated evidence. And the UK's approach seems even stricter than the US. The tribunal didn't allow the firm to re-submit; they just lost the argument entirely. Lucas: Right. And the UK's Solicitors Regulation Authority has since issued a warning notice stating that using AI to generate legal documents without human review may amount to 'dishonest conduct' — which is a very serious finding under the professional code. Lucas: This brings us to a deeper question: what does 'human review' actually mean in practice? Is it enough to skim the output and say 'looks right'? Or do you need to independently verify every source? The ABA's draft says 'independent verification' — which means checking each citation against a trusted legal database like Westlaw or LexisNexis. That's not quick. It might take almost as long as doing the research from scratch. Luna: So the efficiency gain disappears if you do proper due diligence. That's a real problem for the business case of these tools. Firms aren't investing in AI to save twenty percent of time — they want dramatic productivity leaps. But if the ethical obligation requires full re-verification, the net time savings might be negligible. Lucas: Exactly. And that's the tension we're living through. The technology is advancing fast — some newer models specifically designed for legal use claim to reduce hallucination rates to below one percent. But 'below one percent' is not zero. In a profession where a single bad citation can lose a case or trigger malpractice, even one percent is too high. Luna: And if today's tech conversation gave you something usable — a new perspective, a practical takeaway — that's exactly why we keep these episodes free. We deliberately don't run ads on this show. If you want to support that choice, the link is buy me a coffee dot com slash fexingo. Lucas: Yeah, and listener support is the only reason we can keep doing deep dives like this without pitching products. It's a small way to keep the content independent. Luna: Exactly. So back to the legal world — what about the vendors themselves? Are they doing anything to address the hallucination problem? Lucas: Some are. The major legal research platforms — the ones that already have proprietary databases of case law — are building models that retrieve citations from their own databases rather than generating them from scratch. That's called retrieval-augmented generation, or RAG. Instead of the model inventing a citation, it searches a trusted database and pulls the real citation into the text. That's much more reliable. Lucas: But not every firm uses those platforms. Many use general-purpose models like GPT-4 or Claude, which don't have that retrieval functionality built in. And even RAG systems can fail if the retrieval step returns an irrelevant or incorrect result. Luna: So the burden remains on the human. But there's another angle here: what about the client? If an attorney uses AI without telling the client, and the AI makes a mistake, is that a breach of fiduciary duty? Should clients be informed that their legal work is being done partly by a machine? Lucas: That's an open question. The ABA's model rules currently require informed consent for outsourcing legal work to third parties — but does an AI tool count as a third party? Some ethics opinions say yes, because the AI is performing a function that would otherwise be done by a human paralegal or associate. Others argue that AI is just a tool, like a word processor, and no disclosure is needed. Lucas: The New York State Bar Association issued an opinion in 2025 saying that lawyers must disclose the use of generative AI to clients if the AI is used to draft substantive legal documents. That's a leading standard, but it's not universal. Most states haven't ruled yet. Luna: That inconsistency creates a patchwork of obligations. A firm in New York has a disclosure duty; a firm in Texas might not. But the cases cross state lines. If you're a client in a state without a rule, you might never know your brief was ai generated. Lucas: And that's exactly what consumer advocates are pushing to change. They want a federal rule — maybe through the Federal Rules of Civil Procedure — requiring disclosure of AI use in any filing in federal court. A bill was introduced in Congress in May, the 'AI in Court Accountability Act,' but it's early stages. Luna: So where does that leave a practicing attorney today? If you're a lawyer listening to this, what's the practical takeaway? Lucas: I think the safest approach is threefold. One: never trust an ai generated citation without verifying it in a primary source. Two: disclose your use of AI to your client and, where required, to the court. Three: document your verification process — keep a record that you checked each case. That way, even if the AI hallucinates, you can show you fulfilled your professional duty. Lucas: And I'd add a fourth: if your firm is deploying an AI tool, make sure it has a 'retrieval-augmented generation' architecture that pulls from a verified database. Don't just buy the cheapest general-purpose model. Luna: That sounds like a reasonable set of guardrails. But I wonder if they'll persist. As models improve, the hallucination rate will drop. At some point, might the profession decide that verification is not worth the cost? Lucas: That's the billion-dollar question. And it's not just a legal question — it's a question for every profession that relies on verified facts. If we accept a one percent error rate because the ninety-nine percent is so much faster, we're making a value judgment about accuracy versus speed. And in law, accuracy isn't a nice to have. It's the entire product. Luna: So maybe the real ethical line is this: AI can draft, but only a human can attest. And attestation means having actually read the source. Lucas: That's exactly the principle the ABA is moving toward. The human must be the final verifier. And that's not anti-technology — it's pro-responsibility.