Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Doctor Misreads Your Genetic Ancestry
Transcript
- Lucas: So you get a DNA test kit for your birthday, spit in a tube, mail it off, and a few weeks later you get a report that says your risk of developing heart disease by age sixty is 'average.' But what if that 'average' is based on a thousand genomes from Iceland and almost none from West Africa? Luna: That's the uncomfortable question behind polygenic risk scores — PRS for short. They're supposed to predict your odds of getting diseases like diabetes or heart failure based on tiny genetic variations. Lucas: Right. And here's the concrete number I want to anchor this on: a 2025 study out of the Broad Institute analyzed the performance of twenty-two common polygenic risk scores across different ancestries. They found that for Black patients, the accuracy of heart disease prediction dropped by forty-one percent compared to white patients. Forty-one percent. Luna: That's not a small gap. That's basically useless for a huge chunk of the population. Lucas: Exactly. The scores are built on genome-wide association studies — GWAS — and historically about eighty percent of GWAS participants have been of European descent. So the algorithm learns correlations that hold well in people with similar ancestry but break down the further you get from that reference population. Luna: And the companies selling these tests — 23andMe, Ancestry.com, even some academic labs — they're not exactly hiding this. If you dig into the fine print, they'll say something like 'this test was developed using data from populations of European origin.' But how many consumers see that? Lucas: Almost none. And the downstream effect is real. Imagine you're a Black woman in her forties and your PRS says your diabetes risk is low. Based on that, your doctor might not push for early screening. Meanwhile, the actual risk could be elevated — the algorithm just couldn't see it because the training data lacked the relevant variants. Luna: So the technology that's supposed to make medicine more precise is actually making it less precise for the people who already face the worst outcomes. That's a brutal irony. Lucas: It's the AI bias problem in a new wrapper. We've talked about bias in hiring, in lending, in criminal justice. Here it's in genomics, which sounds objective and data-driven but is still only as good as the data you feed it. Luna: And what's being done about it? I know there's the All of Us research program from the NIH, which is explicitly trying to recruit a diverse cohort. But that's still years away from producing usable risk scores. Lucas: Yeah, All of Us has enrolled over a million participants, but the data release for polygenic work is still rolling out. Some labs are trying to build 'trans-ancestry' PRS that combine multiple populations, but the math is hard. You need large enough sample sizes from each group, and the genetic architecture of diseases varies — a variant that matters in Europeans may not even exist in Africans. Luna: And there's a funding problem too. Most GWAS funding goes to large European biobanks like UK Biobank. There's less incentive to invest in African or Latin American cohorts because the infrastructure isn't there and the samples are more expensive to collect. Lucas: Which brings us to the economics. The companies that sell these risk scores have a strong incentive to market them as broadly valid. If they put a big disclaimer reading 'this may not work for you if you're not white,' that hurts sales. So they put it in the fine print and hope nobody looks. Luna: And if the podcast gives you something concrete to think about — a number, a story, a way to spot the trap — that's the whole point. We keep this show going because a small group of listeners chips in monthly at buy me a coffee dot com slash fexingo. No ads, no sponsors, just people who find this useful. And that's what lets us do episodes like this one. Lucas: Yeah, and I'll say — the tech angle today is a reminder that even 'hard science' like DNA can have a data quality problem. And the companies that are transparent about their limitations are the ones worth trusting. Luna: So if someone gets a PRS report, what should they look for? Lucas: Check the population reference. If the report doesn't tell you what ancestry the score was built on, that's a red flag. Some companies now provide a 'diversity score' that shows how well the PRS performs across groups. If that metric is missing, assume it's European-only. Luna: And be skeptical of any claim that a single number can predict your health future. The PRS is just one piece. Lucas: Right. The broader point is that AI in medicine — whether it's reading images, predicting risk, or recommending treatments — inherits the biases of its training data. And when that data is mostly one demographic, the tool can widen the gap it was meant to close. Luna: There is some good news. The Broad Institute study I mentioned earlier — they also showed that when you include diverse data, the accuracy for non-European groups improves significantly. It's not a permanent limitation, just a fixable one. Lucas: But it takes time and money. And while we wait, millions of people are getting risk scores that may be misleading. That's the ethical call to action: if you're building or buying a PRS, demand diverse validation data. Luna: And if you're a consumer, ask your doctor whether the test they're ordering has been validated for your ancestry. It's a simple question that most people don't know to ask. Lucas: Okay, let's zoom out for a second. This whole problem is a case study in what happens when AI development is driven by convenience rather than equity. The easiest data to get is from people with access to healthcare and the time to participate in studies. That skews everything. Luna: And the fix isn't just technical. It's institutional. Funding agencies need to require diversity targets. Journals need to reject papers that only test on European populations. Regulators need to set standards. Lucas: The FDA actually released draft guidance in late 2025 for ai based medical devices, and it includes a recommendation that performance be reported across subgroups. But it's not a mandate yet. Luna: So voluntary compliance. Which means the companies with the best PRS for white people have little incentive to make them work for everyone else. Lucas: Exactly. And that's where the market failure lives. The people who are most harmed by inaccurate scores are often the ones least able to demand better products. Luna: One thing I want to flag: there's a difference between a PRS that's less accurate and one that's actively harmful. In the heart disease case, if it underestimates risk, someone might skip a statin. That's direct harm. Lucas: And in some cases, it can overestimate risk, leading to unnecessary worry and invasive testing. Either way, it's a failure of the tool. Luna: So what's the path forward? More data, obviously, but also better algorithms that can handle small sample sizes. There's work on 'polygenic risk score transferability' that uses Bayesian methods to adjust for ancestry. It's promising but early. Lucas: And there's a role for consumer education. The more people understand that these scores are not destiny, the less likely they are to make life-altering decisions based on a flawed number. Luna: I think the bottom line is this: polygenic risk scores are a powerful tool, but only if they work for everyone. As of June 2026, they mostly work for people of European descent. That's a solvable problem, but it requires intentional action. Lucas: And that's the kind of action we should expect from companies, regulators, and researchers. In the meantime, if you get a DNA test, read the fine print. And if the company can't tell you how well the score performs for your ancestry, that's an answer in itself.