Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / The Hospital Algorithm That Fails Black Patients
Transcript
- Lucas: There's a hospital algorithm that's been making decisions about patient care for years — and it turns out it's been assigning healthier risk scores to Black patients than to white patients with the same chronic conditions. Luna: Wait — so it's systematically rating Black patients as less sick than they actually are? Lucas: Exactly. And we're not talking about some experimental model in a lab. This algorithm was used by millions of patients across major health systems to manage care for chronic diseases like diabetes and hypertension. It was built by a company called Optum, which is part of UnitedHealth Group. Luna: Okay, so how did this bias creep in? Was it the training data? Lucas: Yes — but not in the obvious way where the data just doesn't have enough Black patients. The data had plenty of Black patients. The problem was the proxy the algorithm was trained on. Luna: The proxy? You mean the thing it uses to predict risk? Lucas: Right. The algorithm was designed to predict which patients would need extra care coordination — more doctor visits, more nurse check-ins, more medication management. The intended outcome was 'who will have high healthcare needs.' But the label it was trained on was 'who has high healthcare costs.' Because costs are easy to measure and they usually correlate with needs. Usually. Luna: But not for Black patients. Lucas: Not equally, no. The researchers — a team from Berkeley, Stanford, and other institutions — found that at any given level of chronic illness, Black patients incurred about 20 percent lower healthcare costs than white patients. So the algorithm learned that lower costs meant lower needs. For Black patients, that was wrong. Luna: Why were costs lower? Less access? Less trust? Different treatment patterns? Lucas: All of the above. Black patients are more likely to be uninsured or underinsured, so they get fewer billable services. They may avoid care because of historical mistrust. And there's evidence that doctors recommend fewer procedures for Black patients even when clinically appropriate. So the spending gap isn't about health — it's about systemic disparities. But the algorithm treated spending as a direct measure of health. Luna: So the algorithm was essentially baking in existing disparities and calling them predictions. Lucas: That's exactly what happened. The study, published in Science in 2019, estimated that this bias reduced the number of Black patients identified for extra care by more than half. At the specific threshold the algorithm used, 17.7 percent of white patients were flagged for extra care versus only 13.5 percent of Black patients — even though Black patients had a higher burden of chronic disease. Luna: That's a staggering gap. Was the hospital aware of this? Lucas: Not initially. The researchers actually partnered with the hospital system that was using this algorithm — and when they showed them the results, the clinical team was shocked. They had assumed the algorithm was race-neutral because they didn't explicitly include race as a variable. Luna: Which is a common misconception — that if you don't feed race into the model, it can't be biased. Lucas: Exactly. And that's one of the most important lessons from this case. Bias doesn't require a protected category as an input. It can come through any correlated proxy. In this case, it was healthcare spending. In other cases, it could be zip code, or lab test frequency, or even the number of missed appointments. Luna: So what did the hospital do when they found out? Lucas: They worked with the researchers to retrain the algorithm. Instead of predicting costs, they predicted 'number of chronic conditions' — a more direct measure of health need. When they did that, the bias largely disappeared. The percentage of Black patients flagged for extra care jumped to 18.2 percent, compared to 18.4 percent for white patients. Almost equal. Luna: So the fix was relatively straightforward. Why didn't they do it in the first place? Lucas: Because costs were easier. They had clean, structured claims data. Chronic condition counts require more data cleaning and integration from electronic health records. And there was no incentive to question the proxy — the algorithm seemed to work well on aggregate metrics. Luna: But aggregate metrics can mask group-level disparities. This is the classic 'fairness through unawareness' trap. Lucas: Right. And that's why the push for algorithmic audits is so important. The researchers proposed a simple test: check whether your model's predictions are calibrated equally across racial groups for the same level of true need. If not, you likely have a proxy bias problem. Luna: And this isn't just about Optum's algorithm. How many other healthcare AI systems have the same flaw? Lucas: That's the scary part. A follow-up survey by the same researchers found that well over half of healthcare algorithms in use had never been audited for racial bias. Many use cost-based proxies because that's what's available. And since the Affordable Care Act expanded insurance, cost data is even more ubiquitous. Luna: So you could argue the algorithm was 'working' for the health system's financial goals — reducing costs by targeting the highest spenders — but failing the clinical goal of improving health equity. Lucas: That's a perfect way to put it. The algorithm optimized for the wrong objective. And the health system didn't realize the mismatch because they assumed cost and need were interchangeable. Luna: What about the company that built the algorithm? Did Optum take any responsibility? Lucas: Optum said they were committed to addressing bias and that they would update their models. But the study was published over six years ago — May 2026 now — and we still don't have widespread transparency requirements for healthcare algorithms. The FDA regulates medical devices, but software that recommends care decisions is often classified as 'clinical decision support' and not subject to the same review. Luna: So it's largely self-policing. Lucas: Yes. And self-policing works only if there's an active culture of auditing. The hospital in the study did the right thing by partnering with researchers, but without mandates, many institutions don't allocate resources for this kind of work. Luna: I want to dig into one more thing. The researchers found that if you simply removed race from the model, the bias didn't go away. But if you also changed the proxy, it did. So the problem wasn't race as a variable — it was the proxy. Lucas: Exactly. That's a crucial nuance. Some people argue for 'race-blind' algorithms, but this case shows that blindness to race can actually perpetuate bias if the proxy is correlated with race. The fix wasn't to ignore race — it was to measure the right thing. Luna: Which brings up a bigger question: how do we define 'the right thing' in healthcare? Is it cost? Is it morbidity? Is it patient-reported outcomes? Lucas: And that's where domain expertise becomes indispensable. The algorithm developers probably weren't clinicians who understood the spending disparities. They were data scientists optimizing a cost-prediction model. The collaboration between researchers and clinicians is what uncovered the bias. Luna: So the lesson for any company building AI in a high-stakes domain is: don't just hire data scientists. Hire people who understand the institutional context where the model will be deployed. Lucas: Absolutely. And test your model on subgroups even if you think it's race-neutral. Because the proxy might be hiding something. Luna: I also wonder about the patients affected. Did any of them know they were being scored by an algorithm that undervalued their health? Lucas: Almost certainly not. There's no requirement to inform patients that an algorithm is influencing their care tier. And that's a separate ethical issue — transparency. If your doctor is using a risk score to decide how often to see you, shouldn't you know what's in that score? Luna: It feels like the same conversation we had about credit scores. But in credit, you can at least check your report. In healthcare, there's no equivalent. Lucas: Right. And the consequences of a bad health score are much more severe than a bad credit score. You could be denied access to care management programs that prevent hospitalizations or complications. Luna: So where does this leave us? Are we seeing progress since 2019? Lucas: There's growing awareness. The Biden administration's executive order on AI in 2023 included specific guidance on algorithmic fairness in healthcare. Some states have started introducing bills that require bias audits for clinical algorithms. But implementation is slow. A 2025 survey of hospital systems found that only about 30 percent had conducted any kind of bias audit on their algorithms. Luna: And the other 70 percent are essentially flying blind. Lucas: Flying blind with algorithms that affect millions of lives. And the irony is that many of these algorithms are designed to reduce costs — but by missing high-need Black patients, they may actually increase long-term costs because untreated chronic conditions lead to expensive emergency care. Luna: So biased algorithms can be bad for both equity and the bottom line. Lucas: Exactly. And that's the argument that often gets hospitals to pay attention. Not just 'this is unfair' but 'this is costing you money and harming patients.' When you put it that way, the case for audits becomes much stronger. Luna: It seems like a no-brainer. But I guess the inertia of existing systems is powerful. Lucas: It is. And that's why episodes like this matter. The more people understand that bias isn't just about malicious code — it's about well-intentioned proxies that accidentally discriminate — the harder it is for organizations to claim ignorance. Luna: So what's the one thing our listeners should take away from this episode? Lucas: When you hear about an AI system making decisions in healthcare, ask: what is it actually predicting? If it's predicting cost instead of need, you should be skeptical. And if the developers can't tell you how the model performs across different racial groups, that's a red flag. Luna: That's a solid rule of thumb. And maybe the first step toward a more accountable AI ecosystem. Lucas: One algorithm at a time.