Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Suggests You Break the Law
Transcript
- Lucas: So you ask a chatbot a question about tax law, it gives you a confident answer with citations — and those citations are completely made up. This is not hypothetical. It happened in a federal court in New York in 2023, and it's happening thousands of times a day right now. Luna: The Matsa case — or Mata, rather. I remember that. A lawyer used ChatGPT to draft a brief and the AI invented six whole court decisions that didn't exist. Lucas: Exactly — Mata v. Avianca. The lawyer, Steven Schwartz, was representing a client in an aviation injury claim. He asked ChatGPT for precedent cases, and it generated six rulings with real-sounding names, docket numbers, and quotes. The only problem? None of them had ever been filed. The opposing counsel's team couldn't find them. The judge ultimately sanctioned Schwartz and his firm. Luna: And the lawyer's defense was essentially, 'I didn't know the AI could do that.' Which raises the question: should a reasonable person expect a language model to invent legal citations? Lucas: Right, and that's the core of today's episode. We're not talking about a rogue AI — we're talking about a feature, not a bug. Large language models are fundamentally next-word predictors. They don't have a database of verified facts they're consulting. They're constructing plausible-sounding sequences. And when the training data has examples of legal citations, the model knows the form — it just doesn't know which citations are real. Luna: It's like a student who memorized a bunch of real footnotes but doesn't understand that the cases need to actually exist. They just know the pattern: case name, volume, reporter, page. Lucas: Exactly. And the problem goes way beyond law. In medicine, there are documented cases where AI models recommend drug dosages that are dangerously wrong. In finance, a model might cite an SEC filing that doesn't exist. In engineering, a model could suggest a building material with false safety specs. The pattern is the same: confident, authoritative, and fabricated. Luna: So is there any recourse? I mean, if I rely on an AI's advice and get sued or hurt, can I go after the company that made the model? Lucas: That's the multi-billion dollar question. And the answer right now is: it depends, but the odds are stacked against you. Take OpenAI's terms of service. Section 3a says, and I'm paraphrasing, the output may not always be accurate, and you shouldn't rely on it as the sole source of truth. They explicitly disclaim liability. And that's standard across the industry. Luna: So the fine print says 'don't trust us,' but the product is designed to sound like a trusted expert. There's a mismatch. Lucas: There's a term for that tension — it's called the 'alignment tax.' The idea is that making an AI model more helpful and fluent often comes at the cost of making it less honest about uncertainty. A model that says 'I'm not sure, you should check with a professional' is less likely to be used than one that says 'Here's the answer.' So companies optimize for engagement, and hallucinations are part of that bargain. Luna: And the user is left holding the bag. I mean, in the Mata case, the lawyer got sanctioned, not OpenAI. The judge said the lawyer had a duty to verify the citations. Lucas: Correct. Judge Castel wrote in his opinion that the lawyer had abandoned his role as gatekeeper. Which is a fair point — professionals do have a duty of competence. But the broader question is whether AI companies have a duty to design models that are less prone to hallucination in high-stakes domains. And some are trying. Luna: Like how? Are there technical fixes? Lucas: Several. One is retrieval-augmented generation — RAG. Instead of the model generating a citation from scratch, it's given access to a database of verified documents and instructed to pull quotes only from that source. That reduces hallucinations dramatically. Another approach is to train the model to say 'I don't know' more often. But that reduces user engagement, so there's a real business incentive not to do it. Luna: So the question is whether regulation will force them to. The EU AI Act, for example, classifies high-risk AI systems and requires human oversight. Could a legal advice chatbot be considered high-risk? Lucas: Potentially, yes. If the AI is used in the practice of law, which involves significant legal effects on individuals, it could trigger the high-risk category. That would require transparency, accuracy benchmarks, and human review. But we're still in early days of enforcement. The EU AI Act only started applying some provisions in phases from 2024 onward, and full enforcement is years away. Luna: So in the meantime, what should a listener do if they use AI for work-related research or advice? I mean, I know a lot of small business owners who use ChatGPT to draft contracts, write marketing copy that references regulations, or even suggest tax strategies. Lucas: The baseline rule is: verify anything that sounds like a fact or a source. Especially if it's a citation, a statute, a regulation, or a numerical claim. Assume the model is trying to be helpful but has no concept of truth. It's like asking a very confident intern who has read everything on the internet but has no judgment about what's real. Luna: And if you're a professional — lawyer, doctor, accountant — you probably have an ethical duty to verify. The Mata case is a good warning. Lucas: Right. And I'd add: be careful with the prompts you use. If you ask 'Give me three cases that support my argument,' you're essentially asking for something that may not exist. Instead, ask 'Are there any real cases that support this argument?' And then ask the AI to cite its source, and then check that source yourself. Luna: It's a shift in mindset. Instead of treating AI as an oracle, treat it as a first draft generator. A research assistant you have to double-check. Lucas: Exactly. And that's actually a healthy way to use these tools. They're incredible for brainstorming, summarizing, and drafting. But the moment the output becomes a decision that affects your job, your money, or your health, you have to bring your own expertise to the table. Luna: Or hire a human expert. Which is still a thing. Lucas: It is. And speaking of things that keep this show going — we should mention that the reason we can spend time digging into cases like Mata v. Avianca without chasing sponsors is that listeners like you support us directly. It's a small thing, but it genuinely makes a difference. Luna: Yeah, a couple of dollars a month helps keep the servers running and the research time covered. If you've gotten something out of these episodes, you can find us at buy me a coffee dot com slash fexingo. Lucas: And it keeps the show ad-free, which means we can talk about cases like this without having to soften anything for a sponsor. So thank you to everyone who contributes. Luna: Alright, back to the topic. One more thing I want to touch on: what about AI systems that are specifically marketed as 'professional grade' — like legal research tools that claim to be trained on case law? Are they safer? Lucas: Generally, yes — but not foolproof. Tools like Casetext's CoCounsel or LexisNexis's AI are built with RAG and trained specifically on legal databases. They hallucinate less because they're constrained to a curated corpus. But even those systems can misread a case or miss a key distinction. No AI is perfect. The best approach is to treat AI output as a starting point, not a final answer. Luna: So the message is: AI can be a powerful tool, but it's not a replacement for professional judgment. And the responsibility for the final output rests with the human. Lucas: Exactly. And I think that's a good note to end on. We'll keep tracking this space as regulation evolves and as new cases emerge. Thanks for listening. Luna: See you next time.