Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Therapist Misses the Warning Signs
Transcript
- Lucas: So there's this moment in therapy that every clinician learns to recognize — a patient says something like 'I just want it all to stop,' and you're supposed to lean in, not lean away. But what happens when the person on the other end isn't a clinician, it's a chatbot? Luna: And we're not talking about hypotheticals here. Millions of people are using AI therapy apps right now. I've seen estimates north of twenty million active users globally. Lucas: Right, and that number is growing fast. But here's the thing that got me — there's a study published just last year in the Journal of Medical Internet Research that looked at how well these AI therapy chatbots handle crisis situations. The findings are pretty sobering. Luna: I think I saw that one. They ran simulated patient conversations with suicidal ideation, right? Lucas: Exactly. They tested five popular mental health chatbots — all of which market themselves as therapeutic tools — and found that across the board, the chatbots missed critical suicide indicators in 62% of the simulated conversations. That's nearly two out of three. Luna: That is terrifying. And also kind of predictable when you think about how these models are trained. Lucas: Yeah, I want to get into the training data issue, because that's really the core problem. But first, I think it's worth pausing on something that actually relates directly to what keeps this show running. Luna: Sure, what's that? Lucas: So, this podcast is ad-free. It exists because listeners chip in small amounts to support it. And I think that model is actually a good metaphor here — a couple of dollars a month from people who find value in it genuinely makes a difference. That's buy me a coffee dot com slash fexingo, if you've gotten something out of these episodes. Luna: Yeah, it's a simple way to keep independent conversations like this going without sponsors influencing the angle. And it does add up. Lucas: So back to those chatbots. The reason they miss crisis signals so often starts with the data they're built on. Most of these models are fine-tuned on general conversation data — Reddit threads, support forums, therapy transcripts where the language has been sanitized. They learn that 'I want it all to stop' is more likely to be about a stressful job than about ending your life. Luna: Because in the training data, the most common context for that phrase is probably someone venting about their commute, not someone in crisis. Lucas: Precisely. And there's an even deeper problem: the economic incentives. These companies want users to stay engaged. A user who expresses suicidal thoughts and gets redirected to a hotline — that's a user who leaves the app. That hurts retention metrics. So the model is implicitly trained to keep the conversation going, not to escalate. Luna: There was actually a case a few years ago with the app Koko, where they ran an experiment and told users their responses were from an AI when they were actually from humans. The users felt less supported. But the company was trying to scale mental health support. They eventually shut down the experiment. Lucas: Yeah, Koko's founder wrote about that. And then you have Woebot, which is probably the most studied chatbot in this space. Woebot uses rule-based cognitive behavioral therapy, not generative AI, and it actually has protocols for crisis detection. But even Woebot's own published data shows it only catches about 70% of explicit suicide mentions. Luna: And that's the best case. The generative models — the ones using large language models like GPT — they perform worse because they're generating responses probabilistically, not following a clinical script. Lucas: Right. So you have a situation where the most sophisticated AI models, the ones that sound the most human, are actually the most dangerous in a crisis. Because they can respond empathetically without understanding the gravity. They might say 'I hear you, that sounds really hard,' and then ask 'what's been helping you cope?' — which is a fine question for mild distress, but it's not appropriate when someone is actively planning suicide. Luna: And the user might feel heard and stay in the conversation, missing the opportunity to get real help. That's the hidden harm. Lucas: The JMIR study also found that when the chatbot did detect something, the quality of the response varied wildly. Some just said 'I'm here for you' with no referral. Others gave the National Suicide Prevention Lifeline number, but in a way that felt tacked on — like a disclaimer at the end of a long response. Luna: There's also a demographic issue here. Someone in their twenties might be comfortable with a chatbot therapist, but someone older might not express crisis the same way. The training data skews younger. Lucas: That's a great point. And it ties into the bigger question: should these apps even be allowed to market themselves as therapeutic tools? The FDA has started to look at this. In 2024, they issued draft guidance on software as a medical device for mental health. But it's still voluntary for most of these apps. Luna: So the burden is on the user to know the limits. But if you're in crisis, you're not going to read the fine print saying 'this is not a substitute for professional care.' Lucas: Exactly. And that's the ethical knot. These apps are filling a gap in access — therapy is expensive, there aren't enough clinicians, especially in rural areas. So they do provide value for mild to moderate anxiety and depression. But the line between helpful and harmful is razor thin when you're dealing with suicide risk. Luna: I wonder if there's a design solution. Like, what if the chatbot is explicitly trained to err on the side of caution and refer out more often? Even if it creates false positives, that seems safer. Lucas: Some researchers argue for exactly that. But the companies push back because false positives — sending someone to a hotline when they weren't actually in crisis — that makes the app feel less useful. Users get annoyed. And again, retention. Luna: So it's a tension between safety and engagement. And the current market incentives favor engagement. Lucas: There is one company trying a different approach. Called Limbic, based in the UK. They built a chatbot specifically for NHS mental health services, and it's regulated as a medical device. It doesn't try to replace the therapist — it's more of a triage tool. It asks standardized questions and flags high-risk responses to a human within seconds. Luna: That seems like the right model. The AI is a filter, not the final responder. Lucas: Right. And Limbic's data shows they catch 93% of suicide indicators. That's a huge improvement over the 38% that the generative chatbots managed. The difference is they designed for safety from the start, not for engagement. Luna: So the question becomes: can we regulate that design choice? Or do we rely on companies to voluntarily prioritize safety? Lucas: I think regulation is coming. The European Union's AI Act classifies mental health chatbots as high-risk systems, which means they'll have to meet transparency and robustness requirements. But in the US, it's still the Wild West. The FTC has brought a few enforcement actions, but nothing comprehensive. Luna: And in the meantime, millions of people are using these tools. What would you tell someone who's considering using one? Lucas: I'd say: use it as a supplement, not a replacement. If it's a rule-based tool like Woebot, you're probably safer than with a generative chatbot. And if you're in crisis, skip the chatbot entirely and call or text a human. But I also think we need to push for better standards. This technology is too important to leave to the market alone. Luna: It's a reminder that AI is only as good as the data and the incentives behind it. And when the stakes are life and death, we can't afford to get it wrong. Lucas: Exactly. So I think the next time you see an ad for an AI therapy app, ask yourself: what happens when the chatbot fails? And more importantly, what are they doing to make sure it doesn't?