Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Assistant Misidentifies Your Accent
Transcript
- Lucas: So a couple of weeks ago I'm trying to get my smart speaker to add 'pecan' to the shopping list. I say it the way I've always said it — 'pee can' — and it adds 'peacock.' I tried three times. On the fourth try, I said 'pee kahn' and it worked instantly. Luna: Oh, the pecan wars. My family is from Georgia, so for us it's 'puh kahn.' My phone autocorrects it every time. Lucas: Right, and it's funny until it's not. Because there's a real body of research now showing that voice AI systematically misrecognizes speech from people with non-standard accents — and I'm not talking about foreign accents, I'm talking about perfectly fluent native varieties of English. Luna: You mean like African American Vernacular English, or certain Southern dialects? Lucas: Exactly. Stanford released a study last year that tested speech to text systems from Amazon, Apple, Google, IBM, and Microsoft. They found that for speakers of African American Vernacular English — what linguists call AAVE — the error rate was nearly double what it was for speakers of standard American English. For some systems, it was triple. Luna: And these are the same systems powering everything from your smart speaker to automated captions on video calls to transcription in courtrooms. Lucas: That's what makes it serious. The researchers used a data set called the Corpus of Regional African American Language, which has hours of natural speech from communities across the U.S. And they found that the gap persisted even when they controlled for background noise, speech rate, and clarity of enunciation. Luna: So it's not about mumbling or noise. It's purely a bias in the training data. Lucas: That's the conclusion. The models are overwhelmingly trained on audio from speakers of standard white American English — think NPR announcers, YouTube tutorials from Silicon Valley, things like that. AAVE and Southern dialects are dramatically underrepresented in the training corpus, and the models simply don't learn those phonetic patterns. Luna: I want to get concrete here. What does a two-times or three-times error rate actually look like in practice? Lucas: One example from the study: the phrase 'I ain't got none' was transcribed as 'I am not going to any' by the leading commercial model. That's a complete semantic shift. Now imagine that happening in a deposition or a police interview transcript. Luna: It could change the entire meaning of a testimony. That's a due process issue. Lucas: Exactly. There was a case in 2023 — a man in Detroit was arrested after a facial recognition match, but the arrest warrant affidavit was based in part on an automated transcription of a phone call. The transcription mangled his speech in a way that made him sound more incriminating. He was held for 30 hours before the charges were dropped. Luna: And facial recognition bias gets a lot of attention, but this feels like a quieter, less visible version of the same problem. Lucas: It is. And it's much more widespread. Every time you dictate a text message, every time you use voice search, every time you talk to a customer service bot — there's an accent bias baked in. And most users don't know it's happening because they assume the problem is with their own speech. Luna: That's the insidious part. If your smart speaker constantly mishears you, you start to feel like you're the one who's wrong. Lucas: The Stanford study actually included a survey of AAVE speakers, and a majority said they had consciously altered their speech when interacting with voice assistants — adopting a more standard accent to be understood. That's a form of cognitive load that standard speakers never have to deal with. Luna: So the solution seems straightforward: train on more diverse data. But I imagine it's not that simple. Lucas: It's not. First, collecting high-quality audio data from underrepresented dialects at scale is expensive. Second, there's a privacy concern. Communities that have been historically surveilled are understandably wary of donating their speech data to tech companies. Luna: And even if you get the data, there's the question of how you label it. Dialect features are complex — they're not just about pronunciation, they involve grammar and word choice. Lucas: Right. A model that's been trained to map every input to standard English is going to struggle with constructions that are perfectly grammatical in AAVE but don't exist in standard English. The test of a truly inclusive system would be that it doesn't just transcribe accurately, it's able to transcribe in the speaker's own dialect if that's what they prefer. Luna: That raises an interesting ethical question: should users have the right to know what accent data was used to train the voice AI they're using? Lucas: I think they should. If a company is deploying speech to text in a high-stakes setting like healthcare or law enforcement, they should be required to disclose the demographic makeup of their training data and the error rates across different dialects. Right now, none of them do. Luna: There's actually a bill being drafted in California — the Voice AI Transparency Act — that would mandate exactly that kind of disclosure for any speech system used in government or public services. Lucas: I didn't know about that. That's encouraging. But it only covers government use. Most of the harm happens in consumer products where there's no regulation at all. Luna: So what can a consumer do in the meantime? Is there any way to check if your accent is being fairly handled? Lucas: You can run a simple test yourself. Take a short paragraph that includes some dialect-specific features — like double negatives or certain vowel sounds — and dictate it into three different voice assistants. Compare the transcriptions. If one system consistently performs worse, that's a data point. And then you can decide whether to use that product. Luna: That's actually a really practical takeaway. I'm going to try that with my own system tonight. Lucas: And if you found that test useful — or this whole conversation — it's exactly the kind of thing that listener support makes possible. We keep the show ad-free because of people who chip in at buy me a coffee dot com slash fexingo. It's just a small way to say this kind of deep-dive journalism matters. Luna: Yeah, and it genuinely helps us keep going deeper on these stories instead of chasing clicks. So if today was valuable, that's the place. Lucas: Alright, back to the tech. One development I've been watching is a project called Mozilla Common Voice, which is building an open-source, multi-accent voice data set. They've got contributions from over 90 languages and they're specifically trying to recruit speakers of underrepresented dialects. Luna: So it's a community-driven alternative to the big proprietary data sets. Lucas: Exactly. And early results show that models trained on Common Voice data have more balanced error rates across dialects. It's not perfect, but it's proof that the problem is solvable if you prioritize inclusion in the data collection phase. Luna: But the big players have a massive incentive to stick with their existing data pipelines. Diversity is expensive. Lucas: It is. But the cost of not doing it is also real — in terms of customer frustration, lost trust, and potential litigation. I think we'll see more companies quietly expanding their training data over the next couple of years, partly because they're starting to feel pressure from consumer advocacy groups and partly because the tech itself is getting cheaper. Luna: So where does that leave us? A year from now, will my Georgia relatives be able to ask their smart speaker for a recipe without it turning into a game of charades? Lucas: I think slowly, yes. But the real change won't come from the companies voluntarily — it'll come from regulation and from user demand. So run that test, and if you get bad results, write to the company. That kind of feedback actually lands. Luna: And for the folks building these systems: get out of the lab and listen to how people actually talk. It's not hard, it's just intentional. Lucas: That's the headline. Intentional inclusion. Thanks, Luna.