Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When AI Models Are Trained on Your Medical Data
Transcript
- Lucas: Imagine you go to the hospital for a routine checkup, blood work comes back fine, the doctor says you're healthy. But unbeknownst to you, your medical records — anonymized, stripped of your name and date of birth — get fed into an AI model that a tech company is building to diagnose sepsis. Luna: And you'd never know, because technically you gave consent when you signed that generic form at registration, right? Lucas: Exactly. That broad 'treatment, payment, and operations' consent — it's a catch-all. And it's become the legal foundation for one of the most ethically tangled practices in AI today: training commercial models on patient data without specific, informed permission. Lucas: Today I want to look at a concrete case. In 2023, a major academic medical center — let's call it University Health System — struck a deal with a large cloud provider. They gave the company access to over 3 million de-identified patient records to train a diagnostic algorithm for early detection of acute kidney injury. Luna: I remember that. The class action came about a year later, right? Patients argued that 'de-identified' isn't really anonymous anymore — especially when you combine it with other data sources. Lucas: That's the core tension. The hospital argued that HIPAA allows sharing de-identified data without individual consent. But researchers have shown you can re-identify people from supposedly anonymous health data with something as simple as a zip code and a rare diagnosis. Lucas: And this isn't a hypothetical. In 2019, a study from the University of Melbourne successfully re-identified 99.8 percent of patients in a de-identified dataset using just their age, gender, and postcode. Luna: So the legal shield — de-identification — is actually pretty thin. Yet hospitals and tech companies rely on it constantly. Lucas: Right. And the incentive structure is huge. The global market for AI in healthcare is projected to hit 188 billion dollars by 2030. Every major cloud provider — Amazon, Microsoft, Google — has a healthcare AI division. They need training data, and hospitals have the best data. Luna: Which brings us to the crucial question: should patients have the right to opt out of having their data used for AI training, even if it's de-identified? Lucas: That is exactly the debate the Health Data Rights Act, proposed in the U.S. Senate last month, is trying to address. It would require explicit consent for any secondary use of health data — including AI training — and it would mandate that patients be told exactly which models their data is being used to train. Lucas: But here's the pushback from the medical establishment: they argue that if you require opt-in for every secondary use, you cripple research. Many life-saving discoveries have come from analyzing large datasets without individual consent. Luna: And there's a fairness angle too — if only certain people opt in, the model might not generalize across populations. Lucas: Exactly. That's the bias risk. If the training data skews toward patients who are more engaged, more educated, more trusting of the system — then the AI might perform worse for the people who need it most. Lucas: So on one hand, you want broad, representative data to train accurate models. On the other, you want individual autonomy over something as intimate as your medical history. There's no clean answer. Luna: Makes me think about this in terms of the donation model for the show. If today's tech conversation gave you something usable — a framework, a concrete question to ask — honestly, if it was worth a coffee to you, that's the link: buy me a coffee dot com slash fexingo. Lucas: Yeah, it's a small way to keep the show ad-free and independent. And it's exactly the kind of voluntary, informed choice we're talking about — you decide if the value is there. Luna: Exactly. And we appreciate every single one. Now, back to the legal landscape — what's actually happening in courts right now? Lucas: So the University Health System case settled in early 2025 for an undisclosed sum, but the settlement included a transparency clause: they have to publish a public register of all AI training partnerships involving patient data. That's actually a meaningful precedent. Lucas: Meanwhile, a separate case in Illinois is testing the state's Biometric Information Privacy Act against health data. The plaintiff argues that de-identified data is still a biometric identifier if it's used to train a facial recognition model for patient identification. Luna: That's a creative legal argument — treating a health record as a biometric marker because it's unique to you. Lucas: Right. And it could have huge implications. If courts start treating de-identified health data as a protected biometric, it would effectively require opt-in consent for any AI training use. That would slow down a lot of current projects. Luna: Is there any middle ground? Some kind of dynamic consent model? Lucas: There are experiments with that. A few hospitals are piloting a system where you get a notification on your patient portal — 'Your data could help train an AI to detect kidney failure. Click here to consent, or click here to opt out with zero impact on your care.' Lucas: Early data from one pilot in California shows that when you explain exactly what the AI will be used for — and guarantee that opting out won't affect treatment — about 68 percent of patients consent. That's actually higher than the typical blanket consent rate. Luna: So transparency builds trust, and trust drives participation. That's a hopeful data point. Lucas: It is. But the implementation challenge is real. Most hospitals don't have the IT infrastructure to offer granular consent choices. Their electronic health record systems are ancient — some still run on COBOL. Luna: So we're talking about a systemic upgrade, not just a policy change. Lucas: Exactly. The Health Data Rights Act includes a 2 billion dollar fund to modernize hospital IT systems for exactly this purpose. But even if it passes, it'll take years to roll out. Lucas: In the meantime, what can a patient actually do? I'd say: when you go for your next appointment, ask the registration desk — 'Will my data be used to train any AI models? Can I opt out?' Most front desk staff won't know the answer, but the question itself creates pressure for hospitals to develop answers. Luna: That's a concrete action. And it's surprisingly similar to asking your bank if they sell your transaction data — the answer is often 'we don't know' because the policy is buried. Lucas: Right. The ethical burden shouldn't fall entirely on the patient. But until regulation catches up, informed questions are one of the few levers we have. Luna: So what's the timeline on the Health Data Rights Act? Any chance it becomes law this session? Lucas: It's bipartisan in committee, which is rare these days. But the healthcare lobby is pushing hard against the opt-in requirement — they argue it would add billions in administrative costs. I think the most likely outcome is a compromise: opt-out instead of opt-in, with a clear disclosure requirement. Luna: Opt-out is better than nothing, but it still puts the burden on the patient to act. Lucas: That's the tension at the heart of this whole episode. The technology is developing faster than our ethical frameworks and legal systems can adapt. And the data at stake is the most personal data there is. Lucas: I'll be watching the Senate markup next month. If the bill passes, it could set a global standard — similar to how GDPR reshaped data privacy worldwide. Luna: And if it doesn't, we'll probably see more class actions and more state-level laws, creating a patchwork that's hard for anyone to navigate. Lucas: Exactly. Either way, the genie is out of the bottle. AI is being trained on our health data. The only question is whether we'll have a say in how.