Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Chatbot Gives Medical Advice Without a License
Transcript
- Lucas: You wake up with a sharp pain in your lower right abdomen. It's two in the morning. You're not sure if it's serious. So you pull out your phone and ask an AI symptom checker what's going on. Luna: I've done that. I think most people have at this point. It's just faster than waiting for a doctor's appointment. Lucas: Right. And the answer you get could be anything from 'it's gas' to 'you need to go to the emergency room now.' But here's the thing: the AI giving you that advice is not a doctor. It's not even a medical device in most cases. And a growing body of research suggests these tools are wrong more often than they're right. Luna: We're talking about the symptom checkers built into apps like WebMD, Babylon Health, Ada, and the newer gpt based chatbots that some companies are marketing directly to consumers, right? Lucas: Exactly. And specifically, I want to focus on a study published in October 2025 by researchers at Stanford Medicine. They tested four of the most popular symptom-checker apps — including Babylon Health and WebMD's Symptom Checker — by feeding them 200 standardized medical cases. The cases were real patient scenarios with confirmed diagnoses. Luna: What did they find? Lucas: Across all four apps, the correct diagnosis was listed in the top three suggestions only 32 percent of the time. That means more than two-thirds of the time, the app's best guesses were wrong. And in 18 percent of cases, the app recommended an action that a human doctor would consider potentially harmful — like advising a patient with early-stage appendicitis to take an antacid and wait it out. Luna: Eighteen percent recommending harmful action — that's not a minor error rate. That's the kind of mistake that could kill someone. Lucas: And yet none of these apps carry anything like the regulatory burden that a blood pressure cuff or a pregnancy test does. The FDA classifies most symptom checkers as 'general wellness' tools or 'health information resources' — not medical devices. That means they can go to market without clinical trials, without proof of accuracy, without any post-market surveillance. Luna: There's a company called Babylon Health that has been particularly aggressive here. They've raised over a billion dollars and they've positioned their AI as a replacement for primary care in some markets. In the UK, they had a contract with the NHS for a while. Lucas: That's right. And the Stanford study found that Babylon's app performed worse than the average of the four apps tested. It had a correct diagnosis rate of just 28 percent. Now, to be fair, Babylon has argued that their AI is not meant to diagnose — it's meant to triage, to help patients decide whether they need to see a doctor. But the line between triage and diagnosis is blurry when the app is telling you 'this is likely a urinary tract infection, here's a prescription.' Luna: So who's liable if the AI gets it wrong? The company? The doctor who signed off on the algorithm? The patient who followed the advice? Lucas: That's the million-dollar question. And right now, the answer is largely nobody. In the United States, there's no clear legal framework for AI malpractice. The companies argue that they're just providing information, not practicing medicine. The terms of service usually say something like 'this tool is for informational purposes only and does not constitute medical advice.' But the way these tools are marketed — with phrases like 'your AI doctor' or 'get a diagnosis in seconds' — contradicts that disclaimer. Luna: It feels like we're in a regulatory gap that's been there for years, but the pace of AI deployment is making it much more dangerous. Lucas: Exactly. And this brings us to something I think is worth mentioning. If today's conversation about the real risks of AI in healthcare gave you something useful, something you might want to share with a friend who uses these tools — that's exactly the kind of value we try to deliver on this show every week. And we keep the show free, without ads, because we believe these conversations should be accessible to everyone. If that matters to you, you can support the show at buy me a coffee dot com slash fexingo. No pressure, but every bit helps us keep doing the research. Luna: Yeah, and it's listeners like you who make it possible for us to dig into stories like this one. So thanks for being part of that. Lucas: Alright. So back to the regulatory picture. The EU AI Act, which came into force in stages starting in 2025, does classify some health AI as 'high-risk' — meaning it would have to meet transparency and accuracy standards before being deployed. But the Act has a loophole: AI systems that are 'intended to provide general health information' rather than 'diagnose or recommend treatment' may fall into a lower-risk category. And that's exactly the classification that companies like Babylon are arguing for. Luna: So the companies are lobbying to keep their tools classified as general information, not medical devices, to avoid regulation. Lucas: Exactly. And the Stanford study suggests that the public health cost of that classification is already measurable. Another finding from the study: when the symptom checkers were wrong, they were overconfident. They presented incorrect diagnoses with the same level of certainty as correct ones. So the user has no way of knowing whether they're getting good advice or bad advice. Luna: I've seen that in my own use. The app will say 'We are 85% confident this is a sinus infection.' It feels very authoritative. You don't second-guess it. Lucas: And that's the psychological risk. The interface is designed to inspire trust. The language is clinical. The design mimics a medical encounter. But the underlying model is just a pattern-matching engine trained on text — it has no understanding of anatomy, no ability to examine a patient, no awareness of its own limitations. Luna: There's another dimension here. These AI tools can also encode biases. If the training data underrepresents certain demographics — say, people of color or women with atypical heart attack symptoms — the AI will be less accurate for those groups. A study from 2024 found that some symptom checkers were significantly less accurate for Black patients than white patients. Lucas: Right. And that's not just a fairness issue — it's a safety issue. A missed heart attack in a woman because the AI was trained on male-pattern symptoms is a life-threatening error. And there's no mechanism for reporting those errors because the apps don't track outcomes. They don't follow up to see if the user got better or worse. Luna: So what would meaningful regulation look like? What's the right standard? Lucas: The Stanford researchers propose something straightforward: at minimum, any AI tool that makes a diagnosis or triage recommendation should be required to publish its accuracy rates — broken down by condition, by demographic group, and by severity. And it should be required to have a human-in-the-loop for any recommendation that involves serious illness. That's not a radical idea. It's similar to what we require for over-the-counter pregnancy tests, which have to state their accuracy rate on the box. Luna: But the tech industry pushes back. They say regulation will stifle innovation, that it will slow down the development of AI that could save lives. How do you respond to that? Lucas: I'd say that an AI that kills people because it was never tested isn't innovation — it's negligence. The real innovation is in building systems that are safe enough to trust. And there are companies doing that well. For instance, the Mayo Clinic has developed an AI tool for detecting atrial fibrillation that was tested in a randomized clinical trial with over 2,000 patients before deployment. That's the standard we should expect. Luna: So the problem isn't AI in healthcare. It's AI in healthcare without accountability. Lucas: That's exactly right. The technology itself has enormous potential. But right now, the incentives are misaligned. Companies are rewarded for speed of deployment and user growth, not for accuracy or safety. And the regulatory framework hasn't caught up to the reality that millions of people are using these tools every day. Luna: It's worth noting that this isn't just a consumer issue. Some health insurance companies in the US have started using AI to pre-authorize treatments. There's a well-documented case from 2024 where an AI system used by Cigna was denying claims at a rate of 90 percent — much higher than human reviewers. And patients had no idea their claim was being decided by an algorithm. Lucas: That's a whole other layer. When the AI is not advising the patient but making decisions behind the scenes — that's arguably even more concerning because the patient has no ability to question it. There's a class-action lawsuit against Cigna over that. But the broader point is that we're entering a world where AI is woven into every part of the healthcare system, and the public has very little visibility into how it works or how well it performs. Luna: If listeners take one thing away from this episode, what should it be? Lucas: I think it's this: if you use a symptom checker or a health chatbot, treat it as a starting point — not a diagnosis. Ask yourself: would I trust this same level of certainty from a random person on the internet? And if the answer is no, then consider whether you should trust it from an AI. The technology is improving, but it's not there yet. And until regulators close the gap, the responsibility falls on us as users to stay skeptical. Luna: Fair point. And maybe the next time you feel that pain in your abdomen at 2 AM, you call a nurse hotline instead. Lucas: Or at least cross-reference with a second AI — and then call a doctor in the morning. Alright, that's a wrap for this episode. Thanks for listening.