Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When AI Recommends Your Medical Treatment Without Clinical Trials
Transcript
- Lucas: So you go to the ER with a fever and confusion. The doctor orders blood work, and within hours, an algorithm in the hospital's electronic health record system spits out a recommendation: start antibiotics now, this patient has a high risk of sepsis. Luna: And that algorithm was never tested in a clinical trial the way a new drug or a medical device would be. Lucas: Exactly. That's the reality for a growing number of AI systems embedded in patient care. And the most famous example is Epic Systems' Sepsis Model, which was deployed in hundreds of hospitals starting around 2016. Luna: Epic is the giant of electronic health records. Their sepsis model was meant to flag patients early, before their condition turned critical. Lucas: Right. And it wasn't until 2021 that a study in JAMA Internal Medicine actually evaluated its real-world performance. The results were pretty damning. The model only identified about 63 percent of sepsis cases — and more troubling, it generated alerts for patients who didn't have sepsis, leading to alert fatigue and unnecessary antibiotic use. Luna: So the algorithm that was supposedly saving lives was, in many cases, just adding noise. And the hospitals had been relying on it for years. Lucas: And that's the central problem. We have this regulatory framework for drugs and devices that demands randomized controlled trials, evidence of efficacy and safety, before they ever reach a patient. But AI systems that influence treatment decisions have largely bypassed that process. Luna: The FDA has been trying to catch up. They first approved an ai based medical device for detecting diabetic retinopathy back in 2018. But those are diagnostic tools. The newer systems are actually recommending treatments. Lucas: Right. And the regulatory category is called Software as a Medical Device, or SaMD. The FDA has approved hundreds of SaMD products, but most are for imaging or diagnostics — reading X-rays, spotting tumors. The treatment recommendation systems operate in a grayer area. Luna: Because they're often marketed as clinical decision support, not as medical devices. And there's a regulatory carve-out for decision support that gives clinicians the final say. Lucas: Exactly. The 21st Century Cures Act, passed in 2016, explicitly excluded certain clinical decision support software from FDA oversight, as long as the clinician can independently review the basis for the recommendation. But in practice, when an algorithm says 'this patient is septic,' and the ER is chaotic, how many doctors are really going to second-guess it? Luna: And the algorithm's reasoning is often a black box. The Epic model, for instance, uses dozens of variables — heart rate, white blood cell count, age, things like that — but the exact weighting isn't public. So the doctor can't really evaluate the logic. Lucas: That's the opacity problem. And it gets even trickier when the AI is trained on data from one population but deployed on another. A model trained on academic medical center data might perform poorly at a community hospital with a different patient mix. Luna: There was a well-known case at the University of Michigan where a sepsis model had to be retrained because it was missing cases in younger patients. The algorithm was essentially calibrated to older adults. Lucas: And that's not just a sepsis problem. We're seeing AI models for cancer treatment recommendations, for predicting which patients will respond to immunotherapy, for determining whether a patient needs a CT scan. These are high-stakes decisions, and the evidence base for these models is often thin. Luna: So what would proper validation look like? Should every AI that influences treatment go through a randomized controlled trial like a new drug? Lucas: Some experts argue yes. But that's impractical — trials are expensive and slow, and the technology evolves faster than the trial cycle. The FDA has proposed a middle ground: something called a Predetermined Change Control Plan, or PCCP. Luna: That's the framework where the manufacturer specifies in advance what kinds of updates the AI can make and how it will be validated, and the FDA pre-approves that plan. So the AI can improve without needing a new approval every time. Lucas: Right. The FDA issued final guidance on PCCPs in 2024. It's a thoughtful approach, but it still relies on manufacturers to do rigorous testing. And the agency has limited resources to audit compliance. Luna: The other challenge is that many of these systems aren't even submitted to the FDA. They're marketed as quality improvement tools or research prototypes, so they fall through the regulatory cracks. Lucas: And that's where we are today. There are thousands of AI models running in hospitals, recommending treatments, and the majority have never been validated in a peer-reviewed study that reflects real-world conditions. Luna: So the patient in the bed has no idea that the algorithm influencing their care might be untested. And the doctor might not either. Lucas: That's the core ethical problem. We need transparency — not just about the model's existence, but about its performance in the specific setting where it's being used. And we need regulatory mechanisms that can keep pace with the technology without stifling innovation. Luna: Speaking of keeping things going — you know, this show exists because of a small group of listeners who chip in monthly. It keeps us ad-free and independent, which is especially important when we're covering stuff like this, where there's no commercial pressure to pull punches. Lucas: It's true. We don't run ads, we don't have sponsors, and that lets us dig into topics like AI in clinical trials without worrying about upsetting a healthcare advertiser. If you've gotten something useful out of these conversations, you can join that group at buymeacoffee.com/fexingo. It's a simple way to keep the podcast sustainable. Luna: And we mean it — no pressure, just a genuine thank you to everyone who already does. It makes a real difference. Lucas: Alright, back to the regulatory picture. One promising development is the growing use of 'silent trials' or 'shadow mode' deployments, where the AI runs alongside standard care but its recommendations aren't acted on. That lets researchers gather data on accuracy without risking patient harm. Luna: The Veterans Health Administration has done this with several models. They run the AI in the background for months, compare its predictions to actual outcomes, and only then decide whether to integrate it into clinical workflow. Lucas: That seems like a sensible middle path. It generates real-world evidence without putting patients at immediate risk. But it takes discipline and resources that not every hospital system has. Luna: And it still doesn't solve the problem of models that degrade over time. A sepsis model trained on 2019 data might not work as well in a post-pandemic environment where patient populations and treatment protocols have shifted. Lucas: Right — that's the 'data drift' problem. The model's performance can silently erode, and if nobody is monitoring it, patient care suffers. Some hospitals have set up AI monitoring committees, but that's far from standard practice. Luna: So what would you say to a listener who just got a treatment recommendation from an AI and wants to know if it's trustworthy? Lucas: I'd say ask your doctor: is this recommendation based on a clinical trial or on a model? And if it's a model, has it been validated on a population like you? It's a fair question, and clinicians should be able to answer it. Luna: But right now, most probably can't. And that's the gap we need to close. Lucas: Exactly. The technology is moving fast, but the safeguards are still playing catch-up. And until we have better transparency and real-world validation, we're essentially running an experiment on patients — just without their consent.