Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / AI Systems That Flag Your Child to Child Protective Services
Transcript
- Lucas: When a call comes into a child abuse hotline, someone has to decide whether it's serious enough to send an investigator. In Allegheny County, Pennsylvania — that's Pittsburgh and its surrounding area — that decision has been partly automated since 2016. Luna: Partly automated meaning an algorithm scores each call and tells the hotline worker how risky the situation is. Lucas: Exactly. The Allegheny Family Screening Tool, or AFST, was one of the first predictive models deployed in child welfare in the United States. It takes data from the call itself — what's reported — and merges it with county records: past child welfare cases, incarceration data, mental health services, even birth records. It outputs a score from one to twenty. One means low risk. Twenty means high risk. Luna: And what does the hotline worker do with that score? Do they have to follow it? Lucas: No. The worker is not required to accept the recommendation. They're supposed to use it as one piece of information alongside their own judgment. But studies later showed a strong correlation between higher scores and the decision to screen in a call for investigation. So the model's output clearly influences the outcome, even if it's not binding. Luna: That raises an obvious question: does the model actually identify dangerous situations accurately, or does it just amplify existing biases in the system? Lucas: That's the central debate. The county published a series of evaluations. The model's area under the curve — a common accuracy metric — was around 0.82 to 0.86 during the pilot, which is decent but not perfect. But accuracy alone doesn't address fairness. A 2019 study by researchers at the University of Pittsburgh found that Black families were more likely to be flagged as high risk even when controlling for prior system contact. Luna: So the model was essentially replicating the disparities from the historical data it was trained on. Lucas: Right. The training data came from years of Allegheny County child welfare records. If those records reflect biased decision-making — which they almost certainly do, given well-documented racial disparities in the child welfare system — then the model learns those patterns as if they're objective truth. A family that had a prior CPS call, even if unfounded, becomes riskier in the model's eyes. And since Black families in Allegheny County have historically higher rates of contact with the system, the model disproportionately flags them. Luna: It's a feedback loop. The system investigates more Black families, generates more records of investigations, and then the model uses those records to predict more risk for Black families. Lucas: That's exactly the critique. But the county officials have pushed back, saying the tool actually helps standardize decisions and reduce human bias. They argue that before the AFST, screening decisions were even more inconsistent — one worker might screen in a call that another would dismiss. The model at least provides a consistent baseline. Luna: That's a fair point. Inconsistency is also a form of unfairness. If two identical calls get different responses depending on who answers the phone, that's a problem too. Lucas: Yeah. And the AFST does have some transparency that many other government algorithms lack. The county released a detailed technical report, held public meetings, and even launched a community oversight board. But critics say transparency isn't enough when the stakes are so high. A false positive — a family that's investigated unnecessarily — can be traumatic. A false negative — a family that's not investigated when they should be — can be catastrophic. Luna: Let me play the skeptic here for a moment. If the model is decently accurate and reduces inconsistency, and workers can override it, isn't that a net improvement over pure human judgment? Or are we holding algorithms to an impossible standard? Lucas: That's a really important question. I think the answer depends on how you weigh the costs. The county's own data shows that of the calls screened in for investigation, only about 1.9 percent actually led to removal of a child from the home. So the vast majority of investigations don't result in removal. That suggests the system — model plus human — is already conservative. But that also means a lot of families are being investigated unnecessarily. Luna: What about the families who are investigated but never flagged by the algorithm? How many of those are false positives from the model? Lucas: That's harder to pin down because the evaluation metrics focus on whether a call was screened in, not whether the abuse was confirmed. But one study did look at that: among high-risk scored calls that were investigated, only about 12 percent resulted in a substantiated finding of abuse or neglect. So 88 percent of high-risk investigations turned up nothing. Luna: Eighty-eight percent false positive rate for the highest risk category. That's not great. Lucas: No, it's not. And those numbers aren't unique to Allegheny. Similar tools in other states — like the Rapid Safety Feedback system in Los Angeles — show comparable patterns. The fundamental challenge is that child maltreatment is rare and hard to predict. Machine learning models struggle when the base rate of the event you're trying to predict is very low. You end up with a lot of false positives just because the event is rare. Luna: Right. Even a model with high specificity will still flag many more false positives than true positives when the event happens in, say, 2 percent of cases. Lucas: Exactly. And that's before we even get to the question of what features the model is using. Some of the inputs are arguably proxies for poverty: having received public assistance, having a prior case history, living in a high-poverty zip code. The model may be flagging poverty as a risk factor, which then punishes families for being poor. Luna: But that's not the algorithm's fault alone. Social workers also use those same proxies in their own judgment. Lucas: True. The difference is that when a human uses those proxies, it's often unconscious. The algorithm makes them explicit. And that creates an opportunity for scrutiny. You can audit an algorithm. You can't easily audit the thousands of human decisions made across a county every year. Luna: There's also the question of whether families even know they're being scored. Does Allegheny County inform families that an algorithm has assigned them a risk score? Lucas: Not directly. The score is used internally by hotline workers. Families don't receive a notification saying 'Your case was scored a 14 out of 20.' And there's no mechanism for a family to challenge the score or request a review of the algorithm's inputs. Luna: That seems like a transparency gap. If a score influences whether a worker investigates you, you should at least know that score exists and what data it used. Lucas: I agree. And this is where we start to see a real tension. On one hand, child welfare agencies are trying to use data to make better decisions in a system that's chronically underfunded and overburdened. On the other hand, the tools they're deploying have real limitations and potential harms that aren't always communicated to the people affected. Luna: Before we go further, I want to mention something that's been on my mind while we've been talking about this. And it actually ties into why we can explore these questions with the depth we do. Lucas: Go ahead. Luna: This show is listener-supported. There are no ads, no sponsors — just a small group of people who chip in monthly via buy me a coffee dot com slash fexingo. That support is what lets us spend time on episodes like this, where the topic is important but not necessarily viral. So if you find these conversations useful, that's the only way it keeps going. Lucas: Yeah. It's a quiet thing, but it makes a real difference. And it's why we stay independent. Now, back to Allegheny County — one thing I find really interesting is how the conversation around AFST has evolved over the past decade. Luna: How so? Lucas: In the early years, the debate was mostly about accuracy and bias. More recently, the discussion has shifted to accountability and redress. Who is responsible when an algorithm contributes to a harmful outcome? The software vendor? The county? The hotline worker? And what recourse does a family have if they believe the algorithm led to an unjustified investigation? Luna: Is there a legal precedent for that kind of challenge? Lucas: Not yet at the federal level. But there have been state-level bills proposed. In 2021, a bill in California would have required child welfare agencies to disclose when an algorithm is used and to provide families with a way to contest the score. It didn't pass, but it signals where the policy conversation is heading. Luna: It seems like the core issue is that these tools are deployed in a domain where the risks are asymmetric. The cost of a false negative is huge — a child could die. The cost of a false positive is also huge — a family could be traumatized. And the algorithm can't resolve that trade-off; it just inherits whatever threshold the designers choose. Lucas: Exactly. And in practice, agencies often set the threshold low to avoid missing real abuse. That means more false positives. But the people who suffer those false positives are disproportionately low-income families of color. So the algorithm becomes another mechanism through which structural inequality is reproduced. Luna: Is there any version of this that works better? Could we design a child welfare algorithm that actually reduces bias? Lucas: Some researchers have proposed using fairness constraints during training — forcing the model to have equal false positive rates across racial groups, for example. But there's a catch: if you equalize false positive rates, you might increase false negative rates for one group. And in child welfare, a higher false negative rate could mean more children left in dangerous homes. So you're trading one harm for another. Luna: That's a genuinely hard ethical dilemma. You can't avoid the trade-off; you can only decide which side you're willing to err on. Lucas: Right. And that decision shouldn't be made solely by data scientists or county administrators. It should involve the communities that are most affected. That's why some advocates call for participatory design processes — bringing in parents, social workers, and child welfare experts to set the threshold together. Luna: Has Allegheny County done that? Lucas: To some extent. They had a community advisory board, but critics say it lacked real power. The board could review and discuss, but the county made the final decisions. And the board's membership didn't include many people who had actually been through the child welfare system. Luna: That's the missing voice. If you're building a tool that decides whether someone investigates your parenting, you should have a seat at the table. Lucas: Exactly. And that's a broader lesson for AI ethics in government: the people most impacted by a system should have meaningful influence over its design and oversight. Otherwise, it's just technocracy dressed up as progress. Luna: So what's the current status of the AFST? Is it still in use? Lucas: Yes. As of this year, 2026, it's still operational in Allegheny County. The county has made updates over time — they've changed some features, recalibrated the model, and published annual reports. But the core logic remains the same: a call comes in, the algorithm scores it, and the worker decides. Luna: And other counties and states have followed suit. There are similar tools in Los Angeles, in Colorado, in Washington state. So this is not just a Pennsylvania story. Lucas: No. It's a national trend. And as more jurisdictions adopt these systems, the questions we've been asking become more urgent. How do we ensure transparency? How do we build accountability? And how do we keep the human element central in decisions that affect families' lives? Luna: It feels like the answer isn't to abandon algorithms entirely, but to embed them in a framework that respects the complexity of child welfare. Lucas: Yeah. The technology itself isn't the villain. The danger is deploying it without the necessary safeguards and community input. Done right, a predictive model could help triage limited resources to the families that need them most. Done wrong, it becomes another tool for surveillance and punishment. Luna: And that's the choice we face every time a new government AI system goes live. Lucas: It is. And the clock is ticking. More are coming online every year. The real test isn't whether the model is accurate — it's whether the system that contains it is just.