Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Lawyer Fabricates Case Law
Transcript
- Lucas: So there's this attorney in Illinois, David Martinez, who last February submitted a legal brief in a personal injury case. Nothing unusual about that — lawyers file motions every day. But the five cases Martinez cited to support his argument? They didn't exist. Luna: He just made them up? Or did his AI make them up? Lucas: His AI. He used ChatGPT to draft the motion, and the model supplied what looked like real case names, real judges, real docket numbers. But when the opposing counsel checked, the cases weren't in any database. The judge ended up imposing sanctions — a formal reprimand and a thousand-dollar fine. Luna: A thousand dollars feels light for something that could have derailed an entire case. Lucas: It's a symbolic penalty, really. The real cost is reputational. Martinez's name is now attached to what's become a cautionary tale in legal circles. And honestly, if today's conversation gave you something usable — a concrete example of why AI can't just be trusted out of the box — that's the kind of thing that makes listener support meaningful. If it was worth a coffee to you, you know the link: buy me a coffee dot com slash fexingo. Luna: Yeah, it's a small gesture that keeps this show ad-free and independent. Lucas: Exactly. So back to Martinez — the important thing is this isn't an isolated weird glitch. It's a structural feature of large language models. They're designed to generate text that looks coherent, not text that's factually correct. Luna: Right. They're predicting the next word, not verifying the last one. Lucas: Precisely. The model has seen thousands of legal citations in its training data. It knows the format: case name, volume, reporter, page. So when asked to write a motion, it reproduces that pattern with confidence. But it has no internal database to check against. The citation is plausible — but fake. Luna: And this happens in medicine, in journalism, in technical documentation. But in law, the stakes are uniquely high because a fake precedent can get a case dismissed, or worse, set a bad precedent if nobody catches it. Lucas: Exactly. And the scary part is — a lot of these fakes are never caught. One study from late 2025 looked at over two thousand court filings that mentioned using AI. They found that roughly eleven percent contained at least one citation that couldn't be verified. That's one in ten briefs. Luna: One in ten. That's huge. And it's probably an undercount because not every attorney discloses AI use. Lucas: Right. The American Bar Association issued an advisory opinion in 2024 saying lawyers have a duty to review ai generated work. But enforcement is inconsistent. Some judges now require a certification that any ai assisted filing has been manually checked. Others don't. Luna: So what's the fix? Better prompting? Specialized legal AI models trained only on verified case law? Lucas: Both, and neither. Better prompting helps — you can tell the model 'only cite cases from the Westlaw database' but it might still invent them. Specialized models like those from companies such as Casetext or vLex are trained on curated legal databases, which reduces hallucinations. But they're not immune. Even a retrieval-augmented generation system — where the model pulls from a verified database in real time — can fail if the retrieval step returns the wrong document. Luna: So the human-in-the-loop isn't optional. It's fundamental. Lucas: Exactly. And that's where the ethics get interesting. Because if a lawyer uses AI and doesn't verify, who's responsible? The lawyer, obviously. But what about the software vendor? OpenAI's terms of service say the user is responsible for outputs. But some argue that if a product is marketed as a legal research tool, the vendor should bear some liability. Luna: There's a case working its way through federal court in New York right now — a class action against a legal AI startup that allegedly fabricated dozens of citations. The plaintiffs are law firms that used the tool and then got sanctioned. Lucas: That's the kind of lawsuit that could reshape the market. If vendors start facing financial penalties, they'll invest more in accuracy. But the underlying problem is technical. Language models don't know what they don't know. They have no calibration for uncertainty. Luna: So they'll confidently assert anything. It's like a student who memorized the format of a citation but never read the case. Lucas: That's a perfect analogy. And the student in this case has read billions of documents — but still doesn't understand truth. So the real question for the legal profession, and for any high-stakes field, is: can we build systems that express uncertainty when they're guessing? Luna: There are research groups working on that — teaching models to say 'I don't know' or to attach a confidence score. But it's not mainstream yet. Lucas: It's not. And until it is, the burden falls on the professional using the tool. Which brings us back to the attorney in Illinois. He said in his defense that he assumed the AI was accurate because it sounded so authoritative. That's the seductive danger — the model's tone signals certainty, but the content is hollow. Luna: I think there's a parallel with the early days of search engines. People used to believe that if Google ranked something first, it must be true. Now we have some digital literacy around search. But AI is different — it doesn't just rank information, it generates it. Lucas: Right. And it generates it in a conversational voice that feels personal. That increases trust. So we need a new kind of literacy — what some call 'AI literacy' — that teaches people to treat any AI output as a draft, not a fact. Luna: Some law schools are already adding that to their curriculum. Harvard Law now has a required module on AI verification for first-year students. Lucas: That's promising. But it'll take years before every practicing attorney has that training. In the meantime, we'll see more Martinez cases. More sanctions. More judicial ire. Luna: And maybe that pressure will accelerate the technical fixes. The market might demand models that can cite sources the way a human does — with a reference that can be checked. Lucas: Some models already do that. ChatGPT can now include citations when you ask it to, but they're still sometimes hallucinated. The next frontier is what's called 'attributable generation' — the model only generates text for which it can point to a specific source in its training data or a retrieved document. Luna: Google's Gemini has a 'double-check' feature that highlights claims and searches the web to verify them. It's not perfect, but it's a step. Lucas: And it shows that the tech companies understand the liability. But they're also racing to deploy features, and safety often lags behind. The Martinez case is a useful snapshot of where we are in mid-2026: the technology is powerful, the incentives are misaligned, and human judgment is the only real safeguard. Luna: So what do you tell a lawyer who wants to use AI but is worried about this? Lucas: I'd say: use it as a brainstorming tool, not a drafting tool. Or if you do use it to draft, change your workflow. Every citation must be manually verified against a trusted database. And don't rely on the model's confidence — it's not a signal of accuracy. Luna: It's a cautionary tale, but also an opportunity. The field of AI ethics is being shaped by these real-world failures. Lucas: That's the optimistic view. The pessimistic view is that as models get better at mimicking human writing, the hallucinations will become harder to spot. The fake citations won't look fake — they'll look exactly like real citations, down to the page numbers. And then the only defense is systematic verification. Luna: Which is expensive and time-consuming. So the gap between firms that can afford verification and those that can't will widen. Lucas: That's a justice issue. If the technology primarily benefits well-resourced firms, while solo practitioners and small firms get burned by hallucinations, then AI isn't democratizing legal help — it's concentrating risk. Luna: I hadn't thought of that angle. So the ethics aren't just about individual responsibility. They're about systemic access. Lucas: Exactly. And that's the kind of conversation worth having as we watch these tools become embedded in every profession. For today, the takeaway is: if you rely on a language model for anything that matters, verify. Every time. Luna: And if you're a consumer of legal services, maybe ask your attorney if they use AI — and how they check it. Lucas: That's a fair question. And one we'll probably see more of as these cases make headlines. Thanks for listening — we'll be back next time with another angle on where AI and ethics collide.