Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence

How AI Chatbots Are Learning Bias from User Feedback
In this episode of AI Ethics with Fexingo, Lucas and Luna explore how reinforcement learning from human feedback (RLHF) can introduce subtle biases into language models. They dive into a real-world example from early 2026, when a major chatbot began refusing to answer questions about certain historical figures after receiving thousands of user flaggings — even when those flaggings were incorrect or politically motivated. The hosts explain how RLHF works, why crowd-sourced feedback amplifies majority viewpoints, and what companies like Anthropic and Google are doing to fix it. They also…
The skinny
The skinny isn't ready yet — notes appear once the transcript is processed.