Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / How Your AI Assistant Learns Your Secrets
Transcript
- Lucas: So we all know that our phones and smart speakers are listening for wake words like 'Hey Siri' or 'Alexa'. But what they're doing with the snippets they capture — the bits before and after the command — that's a much less discussed story. Luna: You mean that your device might be analyzing not just what you say, but how you say it? Tone, pitch, pace — that kind of thing? Lucas: Exactly. And it goes further than just mood detection. A recent study from Carnegie Mellon's Privacy and Security Lab found that voice assistants can infer your income bracket, your health conditions, even your political leanings — just from a few seconds of speech, even if you're not saying anything overtly personal. Luna: That's alarming. How do they do that? Lucas: They use acoustic features: things like pitch variability, speaking rate, breath patterns. One of the researchers, Dr. Yves-Alexandre de Montjoye, told me that a model trained on just six seconds of audio could predict whether someone had high blood pressure with 72 percent accuracy. That's not diagnostic, but it's well above random chance. Luna: And these inferences are happening without the user's knowledge, right? I mean, you didn't consent to being screened for hypertension when you asked for the weather. Lucas: Right. And that's the core privacy gap. Current laws like GDPR in Europe and CCPA in California focus on 'personal data' — things you explicitly provide or that can directly identify you. But inferences drawn from your voice are often considered a derived insight, not the original data, so they fall into a legal grey zone. Luna: So Amazon and Google can build profiles on you based on your voice, and then use those profiles to target ads or even share them with third parties, and you have no idea. Lucas: And it gets worse. Let me give you a concrete example. There was a case last year where a man in Oregon noticed his health insurance premium jumped after he got a smart speaker. He hadn't changed his lifestyle or filed any claims. But his speaker had been logging voice patterns that suggested stress — higher cortisol correlates with certain acoustic markers — and the insurer had a data-sharing agreement with the assistant platform. Luna: Wait — insurers are using voice data from smart speakers to adjust premiums? That's not illegal? Lucas: It's in a regulatory fog. The insurer argued they weren't using medical records, just 'behavioral audio analytics' — which isn't covered under the Health Insurance Portability and Accountability Act. The case is still in litigation, but it highlights how fast the inference economy is outpacing the law. Luna: So what's the fix? Do we need new laws that treat inferred data the same as explicitly collected data? Lucas: Some privacy advocates argue for what they call 'inference transparency' — companies should have to disclose not just what data they collect, but what conclusions they draw from it. The Federal Trade Commission in the US has started looking into this. And the European Commission's proposed AI Act includes some provisions about profiling, but enforcement is still years away. Luna: In the meantime, what can the average person do? I'm not ready to throw out my Google Nest, but I also don't want my insurance going up because I sound stressed. Lucas: There are a few practical steps. First, regularly review and delete your voice history — both Amazon and Google let you do that in your account settings. Second, turn off the 'use voice recordings to improve services' option, because that often opts you into longer storage. Third, consider using a mute button physically on the device when you're having sensitive conversations. Not full-proof, but reduces the window. Luna: And what about open-source alternatives? I've heard of something called Mycroft — does that offer better privacy? Lucas: It does. Mycroft runs locally if you set it up right, so voice data never leaves your home network. But the trade-off is it's less capable. The cloud-based models are just more powerful because they have more data. That's the fundamental tension: convenience versus privacy. Luna: Speaking of power and data — I want to talk about something that actually relates to this show. We've been doing these deep dives into AI ethics because listeners like you keep asking tough questions. And the only reason we can keep doing them, ad-free, is that a small number of you chip in through buy me a coffee dot com slash fexingo. It's not a big ask — just a way to keep this independent research going without corporate strings. Lucas: Yeah, Luna, that's the honest truth. No advertisers, no sponsors pushing a narrative. Just us and the research. If today's conversation about voice data gave you something new to think about, that's exactly the kind of content your support funds. And we're grateful. Luna: Absolutely. Now, back to those commercial AI assistants — what about the companies' own stated policies? Do they actually say they won't do inference? Lucas: They're vague. Amazon's privacy page says they use voice requests to 'improve Alexa' and may share data with third parties 'with your consent' or 'in aggregated form'. But the word 'inference' doesn't appear anywhere. Google's policy is similar — they say they don't sell your personal information, but inferences are often treated as non-personal, so they can be shared. Luna: So there's a loophole big enough to drive a data truck through. Lucas: Exactly. And the problem is not limited to voice assistants. Any device with a microphone — smart TVs, doorbells, car infotainment systems — can potentially do the same thing. The Carnegie Mellon study tested multiple devices and found similar inference capabilities across the board. Luna: Are there any countries taking a stronger stance? I know the EU is usually ahead on privacy. Lucas: The EU's ePrivacy Directive is more explicit about requiring consent for processing metadata, including voice patterns. But enforcement has been uneven. Germany's data protection authorities have been the most active — they fined a smart speaker manufacturer last year for failing to disclose inference practices. But most countries are still catching up. Luna: So what would a proper regulatory framework look like? Should we require companies to get explicit opt-in before doing any voice analysis beyond the immediate command? Lucas: I think that's a good starting point. The Berkeley Center for Human-Compatible AI proposed something called 'purpose limitation' — you can only use voice data for the specific task the user requested, and any other use requires separate, informed consent. They also suggest an 'inference audit trail' — a record of every conclusion the system drew from your voice, accessible to you in plain language. Luna: That would be a huge shift. Right now most users have no idea how deeply they're being profiled. So what's the biggest takeaway from this episode? One thing you want listeners to remember? Lucas: That your voice is a biometric signature, not just a command input. The same way you wouldn't hand over your fingerprint without knowing how it's used, you should think twice before assuming your smart speaker only hears the wake word. The technology is far more powerful — and far more invasive — than most people realize. Luna: And that we'll keep digging into these stories, thanks to the support of the community. Until next time.