Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Therapist Gives Bad Advice Based on Biased Training Data
Transcript
- Lucas: So there's this moment in a 2025 Stanford AI Ethics Lab study that I keep thinking about. They tested a popular therapy chatbot — one that's been downloaded over two million times — and found that for non-white, non-English-speaking users, the chatbot gave inappropriate or potentially harmful responses thirty-four percent more often than for white, English-speaking users. Luna: Thirty-four percent — that's not a rounding error. That's a systematic failure. Lucas: Right, and the thing that makes it stick for me is the root cause. When the researchers dug into why, they found that the training data was scraped almost entirely from Reddit's therapy forums — r/psychotherapy, r/CBT, r/TalkTherapy — which, if you've spent any time there, skew heavily toward a specific demographic: young, white, American, and predisposed to cognitive-behavioral therapy as the default framework. Luna: So the AI learned that CBT is the answer to everything, and if your worldview doesn't fit neatly into that box, it either misses the mark or, worse, pathologizes you. Lucas: Exactly. And this is where the episode's angle comes in. It's not just another 'AI has bias' story. It's about what happens when the training data itself has a therapeutic orthodoxy, and no human is checking that orthodoxy against the user's actual needs. The study had a specific case: a user who was grieving the death of a parent — not clinically depressed, just grieving — and the chatbot immediately started pushing CBT exercises for depression. The user reported feeling invalidated and stopped using the service. Luna: And that's the kind of harm that doesn't show up in a simple accuracy metric. It's relational harm. The AI broke trust at a moment of vulnerability. Lucas: If today's conversation gave you something usable — a framework for thinking about AI training data, a reason to ask more questions before trusting a mental health chatbot — the way we keep these episodes ad-free and focused is through listener support. It's buy me a coffee dot com slash fexingo. Just a simple way to say this kind of deep-dive matters to you. Luna: Yeah, and it genuinely helps us stay independent — no sponsors to please, no pressure to soften the conclusions. So if that resonates, we appreciate it. Lucas: Back to the study. The researchers categorized the chatbot's failures into three buckets: cultural mismatch, where the AI didn't understand norms around family or community; language mismatch, where idioms or emotional expressions got lost; and framework mismatch, where the therapy model itself was inappropriate. The last one is the most insidious, because it's not about translation errors — it's about the AI deciding what kind of help you need based on what it saw most in its training data. Luna: And that's a design choice, not an accident. Someone decided to scrape Reddit forums without auditing for therapeutic diversity. Lucas: The company behind the chatbot — I won't name them because they've since updated their data practices — originally argued that CBT is evidence-based and therefore a safe default. But that misses the point. The evidence base for CBT is heavily studied in Western, educated populations. Applying it universally without adaptation is like using a drug that was only tested on men and prescribing it to women. Luna: Especially when the stakes are mental health. A wrong move can reinforce a negative self-narrative or delay someone from getting real human help. Lucas: Let's talk about the regulatory piece, because it's fascinating. Mental health chatbots currently occupy this gray zone. The FDA has said it will regulate ai driven medical devices, but most therapy chatbots are classified as 'general wellness' products — not medical devices — because they don't claim to diagnose or treat specific conditions. So they're essentially unregulated. Luna: Which means the burden falls on the user to figure out whether the AI is giving good advice. And most users don't have the expertise to evaluate that. Lucas: The Stanford study recommended three things that I think are worth remembering. First, training data transparency — companies should disclose the demographic and therapeutic composition of their datasets, so users and clinicians can assess fit. Second, real-time safety monitoring — if the AI detects that it's moving outside its training domain, it should flag that and offer a referral to a human. Third, a requirement for human-in-the-loop for high-stakes interactions. Luna: The human-in-the-loop part is interesting, because it's expensive. But the cost of not having it is harm, and that harm is disproportionately borne by already marginalized users. Lucas: There's also a broader lesson here for any AI that deals with subjective human experience — not just therapy, but coaching, education, even financial advice. The training data carries implicit values. If you don't interrogate those values, you bake them into the system. Luna: I think about the financial advice parallel. If you train an investment chatbot on data from a bull market dominated by tech stocks, it's going to give very different advice than one trained on a more diversified historical set. Lucas: Exactly. The same principle applies. So what do we do? As listeners, as potential users, what's a practical takeaway? Luna: I'd say: before you trust an AI with something intimate — your mental health, your finances, your legal questions — ask yourself, 'What was this trained on?' If the answer is vague or opaque, that's a red flag. Lucas: And for the companies building these systems, the message is clear: diversify your training data, test across populations, and build in safety nets. Not as an afterthought, but as a core design requirement. Luna: The Stanford study ends with a line I keep coming back to: 'The most dangerous AI is not the one that fails often, but the one that fails consistently in a way that reinforces existing inequities.' Lucas: That's the episode right there. Thanks for listening. We'll be back next week with another angle on AI and ethics.