Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When AI Decides on Your Bail
Transcript
- Lucas: Luna, did you know that in New Jersey, if you get arrested tonight, there's a good chance an algorithm will help decide whether you sleep in a jail cell or your own bed tomorrow morning? Luna: I know bail algorithms have been around for a while, but I thought New Jersey moved away from cash bail entirely after 2017. Isn't that the state that essentially abolished money bail? Lucas: Yes, and that's actually the fascinating part. New Jersey did replace cash bail with a risk-based system. But the risk assessment itself relies on a tool called the Public Safety Assessment — the PSA — which was developed by the Arnold Foundation. It scores defendants on two dimensions: flight risk and new criminal activity risk. Luna: And the scores are generated by an algorithm. So what are the inputs? That's where the ethical question starts. Lucas: Right. The PSA uses nine factors: age at current arrest, whether the current offense is violent, pending charges, prior convictions, prior failures to appear in court, prior sentences to incarceration, and a few others. Notably, it does not include race or income explicitly. The idea was to create a colorblind, class-blind instrument. Luna: But we've seen this story before. COMPAS in Florida looked race-neutral on paper, but the ProPublica investigation in 2016 showed it was labeling Black defendants as higher risk for re-offense at nearly twice the rate of white defendants. So what does the data actually say about the PSA in New Jersey? Lucas: A study published last year in the journal Science Advances analyzed over a hundred and fifty thousand cases from New Jersey between 2017 and 2022. They found that the PSA's predicted risk scores for Black defendants were systematically less accurate than for white defendants. Specifically, the tool underestimated the flight risk for Black defendants by about twelve percent relative to white defendants. Luna: So if you're Black, the algorithm says you're less likely to skip court than you actually are. But that could lead to more people being released who then fail to appear, which might make judges lose trust in the tool. Lucas: Exactly. And here's the twist: the study also found that the algorithm overestimates risk for Black defendants on the new criminal activity scale by a smaller margin. So it's not a simple case of 'biased against' or 'biased for' — it's pattern of error that varies by outcome. But the important point is that the errors are not random. They correlate with race. Luna: And those errors have real consequences. If a judge sees a risk score that's off, they might override it — but they might also just follow the recommendation. Do we know how often judges actually follow the PSA? Lucas: The study tracked that too. Judges follow the release recommendation about seventy percent of the time for both Black and white defendants. But when they override, it tends to be in the direction of more detention for Black defendants and more release for white defendants. So the algorithm's errors are compounded by human discretion. Luna: So the system is supposed to be fairer than cash bail, where poor people just sit in jail because they can't make bond. But the algorithm introduces its own form of bias. And unlike a human judge, you can't cross-examine the algorithm about its reasoning. Lucas: That's the core issue: transparency. The PSA is proprietary. The Arnold Foundation has not released the full algorithm to the public. Researchers have to reverse-engineer it from the scores and the inputs. So there's no way for a defendant or their lawyer to actually challenge the score in court. You can't say 'the formula double-counted my prior conviction' because you don't know the formula. Luna: That feels like a due process violation. If the state is using a tool to deprive you of liberty, you should be able to look under the hood. But the company says it's a trade secret. So we have this tension between commercial confidentiality and civil rights. Lucas: And it's not just the PSA. Similar tools are used in dozens of states — the Laura and John Arnold Foundation's tool, or variations of it. A 2019 report from the ACLU found that at least twenty states use risk assessment tools in pretrial decisions. And most of those tools are not publicly audited. Luna: So what would a fair algorithm look like? Could you even build one that's race-blind and accurate? Lucas: That's the million-dollar question. Some researchers argue that you should include race as a variable to correct for systemic bias, because the data itself is tainted by centuries of unequal policing. Others say that's ethically unacceptable because it codifies race. There's no consensus. But what almost everyone agrees on is that any tool used in criminal justice should be subject to independent auditing, and the results should be public. Luna: And that audit should include not just overall accuracy, but accuracy by demographic group, and also look at whether the tool reduces or exacerbates disparities in the system. Lucas: Precisely. And here's where it gets even more uncomfortable: the PSA has been in use for nearly a decade now. New Jersey's pretrial jail population dropped by about forty percent after the 2017 reforms. That's a huge win. But if the tool is quietly producing biased outcomes, we might be celebrating a reform that's actually creating a new, less visible form of injustice. Luna: It reminds me of the saying that the opposite of a good system is a slightly broken system that people assume is working. The PSA is better than cash bail for a lot of people, but it's not nearly good enough. Lucas: And this brings up a broader point about AI in public policy. We're seeing algorithms creep into everything from child welfare to eviction proceedings. In each case, the pattern is similar: a well-intentioned tool is deployed, early studies show promise, then later research reveals bias. But by then, the tool is embedded in courtrooms or agency workflows, and it's very hard to unwind. Luna: So what's the path forward? Should we pause the use of these tools until we have better standards? Lucas: Some advocates say yes — a moratorium on risk assessment tools in criminal justice until there's a federal audit framework. Others say that would leave us with the older, arguably worse system of cash bail. For me, the most realistic near-term step is mandatory algorithmic impact assessments, like what the EU AI Act is beginning to require for high-risk systems. Those would at least force developers to disclose performance metrics across demographic groups before deployment. Luna: And that's something our listeners can actually look up — if their state uses a pretrial risk tool, they can check whether it's been audited. In many cases, the answer will be no. Lucas: And if it hasn't, that's a story worth reporting. Which is exactly the kind of thing we try to do here. And speaking of that, we want to mention something quickly. You know, we deliberately keep these episodes free of ads. No sponsors, no commercial breaks. That's a conscious choice — we think the topic deserves a clean conversation without someone trying to sell you something. Luna: Yeah, it's a principle we stick to. And the way that works is through listeners who choose to support the show. If you find value in what we do, you can find us at buy me a coffee dot com slash fexingo. It's a simple way to help us keep going ad-free. Lucas: And every contribution genuinely helps — it lets us spend more time on research and less on chasing sponsors. So thank you to anyone who's already supported us. Now, back to the PSA: one thing I didn't mention is that the Arnold Foundation actually did release a limited audit in 2021. They found that the PSA's predictive accuracy was 'consistent across race' in terms of area under the curve — that's a statistical measure. Luna: But that's not the same as saying it's fair. A model can have the same overall accuracy for two groups and still have different error patterns — false positives versus false negatives — that impact groups differently. Lucas: Exactly right. The Arnold Foundation's own report acknowledged that the PSA might have differential validity by race, but they argued that the tool was still an improvement over human judgment. The problem is, the baseline for comparison — human judges — is also biased. So you're comparing a biased algorithm to biased humans, and the algorithm might be slightly less biased. But that's a very low bar. Luna: It's like saying your car is safer because it only crashes ten percent of the time instead of fifteen. You still don't want to drive it. Lucas: Right. And there's another layer: these tools don't just predict the future — they shape it. If the PSA says a person is low risk, they get released and are less likely to be rearrested simply because they're not in jail. Conversely, if someone is detained based on a high risk score, they might plead guilty just to get out, which then becomes a prior conviction that raises their risk score next time. It's a feedback loop. Luna: That's a pernicious cycle. The algorithm creates the reality it claims to predict. So we're not just measuring risk, we're manufacturing it. Lucas: And that's why I think the ethical conversation around AI in criminal justice needs to go beyond bias in the narrow sense — we need to talk about the political economy of risk assessment. Who profits from these tools? Who has access to the data? What happens when the company that owns the algorithm gets acquired and the terms of use change? Luna: Those are questions we probably don't have good answers to yet. But they're the right questions to ask. And for listeners who want to dive deeper, the study we referenced is by Megan Stevenson at George Mason University — she's done some of the best work on pretrial risk assessment. It's worth reading. Lucas: Absolutely. So to wrap up: the PSA in New Jersey is a powerful example of how a well-intentioned reform can create new ethical problems. The tool reduces the jail population, but it does so in a way that's racially uneven and opaque. The lesson is not that we should never use algorithms in criminal justice, but that we need a radically higher standard of transparency and auditability before we deploy them. Luna: And maybe we need to ask if some decisions are too important to be decided by a black box, even if the black box is better than the alternative. Lucas: That's the fundamental question. And it's one that will only become more urgent as AI systems become more capable. Thanks for listening, everyone.