Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When AI Models Police Your Social Media Speech
Transcript
- Lucas: So there's this number that's been stuck in my head since I read the Algorithmic Justice League's latest report: a content moderation model trained on standard American English has a forty percent higher false-positive rate when it encounters African American Vernacular English. Luna: Forty percent higher — that's not a rounding error. That's a systemic failure. Lucas: Exactly. And that's the episode. Because we talk a lot about AI deciding bail or credit scores, but the most frequent way most people interact with algorithmic judgment day-to-day is when they post something and it gets taken down or shadow-banned. Luna: Right. Facebook, TikTok, YouTube — they all use AI to flag hate speech, harassment, misinformation. And the scale is staggering. Meta alone reported taking action on something like forty million pieces of content in a single quarter last year. Lucas: And almost all of that triage happens before a human ever sees it. So the stakes are: what counts as hate speech? What counts as misinformation? And who gets to decide? Luna: Let's talk about that forty percent number. How did the Algorithmic Justice League measure it? Lucas: They ran a controlled test. They took a standard commercial moderation model — the kind a platform like Twitter might license — and fed it thousands of sentences. Half were written in what linguists call mainstream American English. The other half used grammatical structures and vocabulary from African American Vernacular English, which has its own consistent rules. Luna: And the model flagged the AAVE sentences as hate speech or misinformation at a much higher rate, even when the content was neutral or positive. Lucas: Yeah, a sentence like 'She been had that job' — perfectly grammatical in AAVE — was flagged as potentially hateful because the model misread 'been' as an intensifier with negative intent. The model has no cultural context. Luna: It's not just about dialect. It's about topics too. Content about police brutality, racial justice, even discussions of systemic racism — those get flagged at disproportionate rates because the training data associates certain keywords with hate speech. Lucas: There was a well-documented case on TikTok in 2024 where LGBTQ+ content in the Middle East was being automatically removed. The AI had learned to associate words like 'gay' or 'transgender' with content that violated local hate-speech laws, but it couldn't distinguish between a slur and a term of identity. Luna: So you end up with platforms that are over-moderating marginalized voices while letting actual hate speech through — because hate speech from dominant groups often uses coded language the model hasn't been trained to catch. Lucas: There's research showing that white supremacist content often flies under the radar because it uses euphemisms and historical references. Meanwhile, a Black teenager saying 'this is fire' gets flagged as promoting violence. Luna: So what's the solution? Better training data? More transparency? Lucas: Both. But let's talk about transparency first. Every major platform publishes a transparency report now. They'll tell you how much content they removed and why. But they won't tell you the false-positive rate broken down by dialect or demographic group. Luna: Because that would show the bias. And if they showed it, they'd have to fix it. Lucas: Exactly. There's a proposal from the Algorithmic Justice League that any platform over a certain size should be required to publish what they call an 'adverse impact analysis' — modeled on employment discrimination law. Show us the disparate impact your moderation AI has on protected groups. Luna: That would be a game-changer. But is there any movement on the regulatory side? Lucas: The EU's Digital Services Act already requires some transparency, but it's more about process than outcomes. A few US states have introduced bills — California's AB 2761 would mandate impact assessments for automated decision systems. But it's early. Luna: And in the meantime, people are getting silenced. I think about the Arab American writer who got banned from Instagram for using the word 'shaheed' — which can mean martyr but is also a common term for a fallen loved one. Lucas: The model saw 'shaheed' in a database of terrorism-related terms and pulled the trigger. No context. No nuance. Luna: And that's the core problem. These models are trained on massive datasets scraped from the internet, and the internet is not a neutral representation of human language. It's over-indexed on certain voices and under-indexed on others. Lucas: So what would it take to build a moderation AI that's actually fair? Let's look at one promising approach: participatory auditing. Luna: Tell me about that. Lucas: Instead of the platform hiring a few outside experts to poke at the model, you let affected communities design test cases. So a group of AAVE speakers creates a set of sentences they know are benign but the model is likely to flag. Then the platform has to either fix the model or justify the false positives. Luna: That puts the power back in the hands of the people who are most harmed. Lucas: Right. And there's a pilot program doing this right now — the Civil Rights Corps is working with a coalition of community groups to audit a major platform's hate-speech model. We might see the results later this year. Luna: If today's conversation gave you something useful — maybe a new way to think about what happens when you hit 'post' — the way these episodes stay ad-free is listener support. It's a small thing, but if you go to buy me a coffee dot com slash fexingo, that helps us keep digging into these stories without sponsors. Lucas: Yeah, and we really do read every message that comes with it. It's a direct line to what matters to you. So if this episode clicked, that link is there. Luna: Now, back to auditing. One thing that gives me hope is that some of the model providers are starting to release what they call 'model cards' — a standardized fact sheet about how a model performs across different groups. Lucas: Google's Model Cards toolkit is one example. But it's voluntary. And a model card doesn't mean much if it's not updated regularly with real-world performance data. Luna: So the gap is between intention and enforcement. Lucas: Exactly. And that's where regulation could step in. If you had to certify your moderation model annually — like a car inspection — you'd see a lot more investment in fairness. Luna: Let's talk about one more specific case: the 2025 controversy around Reddit's automated anti-harassment bot. Lucas: Oh, the one that was banning accounts for using the word 'stupid'? Luna: Yeah. The bot was trained on a dataset of toxic comments, but it over-learned. Anything that resembled a personal attack got flagged. People got banned for saying 'That's a stupid idea' in a policy debate. Lucas: And Reddit's response was to tweak the threshold — which is the classic move. But thresholds don't fix the underlying distributional problem. If your model is bad at distinguishing criticism from harassment, raising the threshold just means more harassment gets through. Luna: So it's a trade-off. And the trade-off is always borne by the people whose language is furthest from the training data. Lucas: Look, I don't want to sound like I'm saying moderation is bad. Platforms need to deal with harassment, misinformation, child exploitation. The question is whether we can build systems that are both effective and equitable. Luna: And the answer so far seems to be: not without a lot more work. Lucas: Yeah. But there are concrete steps. First, diversify training data — actively collect examples from underrepresented dialects and communities. Second, implement participatory auditing. Third, mandate transparency that includes disaggregated false-positive rates. Luna: And fourth, give users a meaningful appeals process. Right now, if you get banned, you often have no idea why, and no recourse. Lucas: Some platforms are experimenting with showing users the specific phrase that triggered the flag. But even then, it's a black box explanation. The model doesn't know why it thinks the phrase is hateful. Luna: So we're left with a system that's fast, scalable, and deeply flawed. And the people who suffer most are the ones who already face the most surveillance. Lucas: If you take one thing from this episode, I hope it's that the next time you see a post removed and think 'good, they caught that' — ask yourself whether the same system would catch a dog whistle from a powerful group, and whether it might be silencing someone just for speaking their dialect. Luna: That's the question. And it's not just an academic one. It's happening millions of times a day. Lucas: All right. That's the episode. Thanks for sticking with us.