skinny.

Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence

How AI Content Moderation Models Learn Bias

In this episode, we explore how content moderation AI systems inherit bias from the human decisions used to train them. We start with a concrete case: a 2021 audit of Facebook's moderation algorithm found that posts written in African American Vernacular English were flagged as hate speech at 1.5 times the rate of comparable posts in standard English. Lucas explains how this happens—through skewed training data, ambiguous labeling guidelines, and the pressure on human moderators to err on the side of removal. Luna adds data from Twitter's moderation struggles and asks whether neutral language…

The skinny

The skinny isn't ready yet — notes appear once the transcript is processed.