Go back

Episode 1: How do voice AI Assistants work?

17m 26s

Episode 1: How do voice AI Assistants work?

The podcast series "Always Listening, Can I Trust My AI Assistant" by the SAZE Research Project delves into the security and privacy aspects of voice AI assistants like Amazon Alexa. These assistants function by recording commands post the wake-up word, sending them to the cloud for processing, and executing tasks using selected skills. Users often perceive these assistants as standalone intelligent devices, unaware of the extensive cloud processing and potential data sharing with third parties. While voice AI assistants listen for wake-up words, recordings are stored and accessible to users, with accidental recordings possible. Privacy in AI assistance emphasizes transparency, user data control, and respect for customer rights, as highlighted in the interdisciplinary SAZE Research Project involving academia and industry partners.

Transcription

2515 Words, 14012 Characters

Welcome to always listening, can I trust my AI Assistant, a podcast series from the Secure AI Assistance Research Project? The SAZE Research Project focuses on investigating the security of AI assistance and privacy of its users. In this podcast series, we will be discussing in particular voice AI assistance with the researchers at SAZE and some of their partners to answer the questions, how do AI assistants really work, how do they use or possibly misuse data, and we will start to unravel the question of can I really trust my AI Assistant? Voice AI assistants are becoming more and more common, they are available on every device probably even on the one you're listening to me on today, and they are increasingly appearing in homes across the world in the form of smart speakers. In fact, 4.2 billion voice AI assistants are used around the world. Researching these statistics, I couldn't help but wonder if I should get one. I spoke to Professor Jose Suk, the lead researcher on the SAZE project about why AI assistants have become so popular. I think they're very convenient, they use a way of interacting with us that is quite natural which is just conversation. So that's quite human and I think that's why it's so convenient. They have advanced quite considerably in the last few years, leveraging the advances on natural language processing, and they can understand now lots of things that we tell them. They offer lots of capabilities from checking the weather, to play music, to actually doing things a bit more complicated. They're so called skills that they have that allow you to even manage your bank account, just talking to your assistant. Ok, so that does sound really useful, but to be honest, I don't really understand how they work, so I give a command like "Hey Alexa, what's the weather today?" In London, England, it's 6 degrees Celsius with partly cloudy skies. What's actually going on behind the scenes for her to tell me that? So the way they work is that, first of all, they need to understand what you are saying and then they need to serve that. This instance is in different ways depending on the architecture, but normally is something like, there is a smart speaker, for instance, the Amazon Echo, basically what it does is wait for the wake up work, in this case, Alexa. And when it hears that wake up work from that moment on, is to record everything you say, and then send it to the cloud. So up to that moment where you have not uttered the wake up word, nothing in theory is getting out of your house, at that point where the smart speaker hears the wake up word then starts recording what you are saying and then sends that to the cloud. So it goes to the cloud and then that audio, first of, is turned into text. So it goes through a pipeline of a natural language processing methods powered by machine learning where the audio is first turned into text and then from the text, it does what is called natural language understanding, which is basically from the text, understand what the user's intent is, what the user's want. Once it knows what the user wants, then it is called to serve what the user wants. In that case, then what it does is to decide what's the best skill to serve that, so assistants have skills to do stuff. So a skill is the name given to the applications used in AI assistance, just like mobile phones have apps, voice assistants have skills. Does it explain a bit more about the two types of skills? There are some of them that are native, so Amazon will have put into Alexa some native skills that check in the weather or answering knowledge questions or stuff like that, that are on Amazon Alexa and then there are also plenty, thousands of other skills that third party developers have created as an add-on to Alexa to give Alexa more capabilities so it can do more stuff for users. So then there is also a process, which is also facilitated by some machine learning models where the best skill for what the user wants is selected, and the moment on the actual command what the user wants is passed on to that skill. If it is an native skill, it keeps within the Amazon architecture and systems, if it is a third party skill in mego anywhere in the internet. No, no, no way. People do not usually understand if they are talking to a native or to a third party skill because they don't see that when they are talking to the assistant. They may see when they are looking for skills, they go to the market, but normally that's not even needed for you to interact and ask for things from Alexa. So there's been lots of research including hours that shows that people just don't know that they are talking to an native skill or to a third party skill. So once the information reaches the skill, then what happens? It will then process the command and do what you want to do. The check the weather will do the can then bring back what the weather will be for today. And that's brought back to the Amazon Alexa infrastructure, convert it again to voice. And then that's what you hear back. In addition to that, there may be other things that happen, for instance, when you have your Alexa integrated with your smart home devices, you could say things like Alexa turn up or turn down the temperature, right? So at that point, the third party skill to process your command, what is doing is to contact your actual cloud account of that smart home device, and then from the call, it will be sent to your smart home device that the temperature needs to be brought down. What you see is that normally there is not a direct also interaction that is something users usually don't understand as well. There is not a direct interaction between the Alexa and your smart home device. It depends, but usually what happens is that everything goes to the cloud and then we'll go back to your home to actually having some action in your smart device. So you are to the wake up word along with your command, and that goes to the cloud to be translated by natural language processing. Once it's been understood, a skill is chosen by the native skill that is owned by the service provider like Alexa or Siri, or a third party skill which could have been made by anyone. Then the action is taken, but actually most people don't understand that this is the mechanism behind a voice AI assistant. So as a user, what we found in the studies we've found, we found that across platforms, I mean, didn't matter what that is and you were talking about, there was some consistent mental model from users, a very prominent one that was the picture that they considered the assistant as this all in one thing device, like small brain. So for instance, if you had your Amazon Echo, everything intelligence, everything would be in there, all your data that you tell Alexa would be inside there, this device would be very clever and would respond to everything you say. That was the most prominent mental model that we saw across users of different types, people who've been using as instance for in some cases years. There were only few people that got to the point where they understood the assistant and really it connects to the internet. It uses the internet, it uses the cloud. For many of the things I asked the assistant to do, they've got to connect to the cloud and do stuff there. And I'm not sure they were really understanding that actually this speaker itself does very little, basically just records and sense to the cloud. Nothing else. All the processing of what the user wants happens there as well. I'm not sure those people were understanding that, but at least they had this intuition that the smart speaker in itself was not enough to do everything. It needed to connect to the internet and it needed to connect to the cloud to do stuff. And what was really surprising is that nobody was talking about third parties getting involved here. What was the surprising bit in the sense that, and most they were talking about some queries, some data going to the cloud, some processing happening in the cloud. But very little people actually, nobody in our studies that we did, they were realizing that what was happening is that data needed to go not just to the Amazon cloud infrastructure, but also to other infrastructures in order for their command to be properly served. That was surprising because it really shows that people were unaware at the end of the day that their data does not only go to the cloud, but also to all sorts of other entities that have nothing to do with Amazon in that sense. So your voice assistant really could be sending data anywhere, depending on the skill that you're using and the information it has access to. But isn't this like any app on your phone or program on your computer? They all ask you for access to data about you, whether that's your name or your location or something else. And sometimes it sends that data on to somewhere else, isn't this just the same? It's very interesting that there are similarities between assistants and this idea of the skills. And it is true that there are some similarities with other concepts in other platforms, like your apps in your mobile phone. It's kind of like if you get your mobile phone alone, it does stuff, but it doesn't do much. You usually put apps into it so that you can do more stuff with your phone, right? So that's quite similar to the concept of skills. But there's a difference because apps usually run in your device, in your mobile phone. That doesn't happen with skills. They don't run in your smart speaker. They run somewhere in the internet and can even be other platforms that have nothing to do with Amazon at all. So the only way you really interact with those skills or that you have some control of it is basically what you just say to the device because when it's gone, the data, it's no longer anywhere in a place where you can control it that way. It is true that in our laptops, in our computers and mobile phones, many of the apps we have nowadays, if you take them off line, you won't be able to do much because most of them also have cloud backend that's doing some of the processing or stuff like that. But still, you have them running in your device and you can basically see where data is going from there to some extent. And that's very thoroughly improved with tools like Apple's privacy report and stuff like that. It's actually even better to show you where data may be going, your data may be going. That doesn't happen with skills because they basically run somewhere else. You have no physical access to where skills may be running at all. So data really can go to any place on the internet, depending on the skill that we're using. The design of a voice AI assistant means that we have less agency in finding out where that data is going and what data is being collected. But what about before we even choose a skill? Are the assistants always listening to you? Are they recording your conversations and potentially sending them somewhere else? Well, okay. So when it comes to what gets recorded and the idea and that these systems and what more systems do is that nothing in principle is recorded. The assistant in principle, it's not continuously listening to what you are saying. In some cases, there are things that you do for that to happen, like pressing a button or basically the wake-up word. For instance, some speakers what they have is a very basic AI model that just dedicates itself to pick up the wake-up word, doesn't do anything else. It's not able to do anything else. So you're at home and if you don't say, for instance, Alexa in principle, nothing should be recorded. In practice, there's been some studies of what's called misactivations where the model that is in those speakers made a mistake. It thought somebody had uttered the wake-up word and they didn't. So it's happened that there's been some accidental recording of data that will, of course, go to the club and be processed and then some people sometimes get things like Alexa suddenly saying something. But that can happen because these systems aren't perfect. So there can be what's called misactivation. However, everything that you say after the wake-up word, everything goes to the club, everything gets stored, everything gets processed and you can actually go, we are talking about Amazon Alexa, go to your Amazon account online and you can see the history of everything if said. It's there for you. You can also go there and remove anything that you don't want to. So there's a simple AI model in a voice assistant whose only job is to pick up the activation word, whether that's Alexa, Siri or OK Google. And only when that AI has been activated will anything else happen in the assistant. So is my assistant always listening? Technically yes, but nothing happens until I say the activation word. Next time, we will explore more about the privacy of AI assistants and what happens to user data. Can users control how their data is being used? And we will hear from a data privacy expert about how these issues are affecting businesses. privacy is not actually really that much about privacy, it's about respecting the rights of your customers and acting with transparency and acting with knowledge about what you're doing with, effectively, their property because the data is theirs, it's not yours. That's in episode two of Always Listening, can I trust my AI assistant? The SAIS project is across disciplinary research project between the Department of Informatics, Digital Humanities and the Policy Institute at King's College London and the Department of Computing at Imperial College London, working with non academic industry partners and policy experts, including Microsoft and Securities, who you will hear from in this podcast. If you would like to find out more about SAIS, you can visit us on our website or contact us on Twitter or LinkedIn. All the links are in the show notes. The music in this podcast is by Sage Quadrero.

Podcast Summary

Key Points:

  1. The podcast series "Always Listening, Can I Trust My AI Assistant" is part of the SAZE Research Project focusing on AI assistance security and user privacy.
  2. Voice AI assistants like Amazon Alexa work by recording commands after hearing the wake-up word, sending them to the cloud for processing, and executing tasks based on selected skills.
  3. Users often perceive AI assistants as standalone intelligent devices, unaware of the extensive cloud processing involved and the potential sharing of data with third parties.
  4. Voice AI assistants are designed to listen for wake-up words, and recordings are stored and accessible to users, with accidental recordings possible due to misactivations.
  5. Privacy in AI assistance requires transparency, user data control, and respect for customer rights.

Summary:

The podcast series "Always Listening, Can I Trust My AI Assistant" by the SAZE Research Project delves into the security and privacy aspects of voice AI assistants like Amazon Alexa. These assistants function by recording commands post the wake-up word, sending them to the cloud for processing, and executing tasks using selected skills. Users often perceive these assistants as standalone intelligent devices, unaware of the extensive cloud processing and potential data sharing with third parties.

While voice AI assistants listen for wake-up words, recordings are stored and accessible to users, with accidental recordings possible. Privacy in AI assistance emphasizes transparency, user data control, and respect for customer rights, as highlighted in the interdisciplinary SAZE Research Project involving academia and industry partners.

FAQs

Voice AI assistants work by first understanding the user's command, then selecting the appropriate skill to fulfill that command, and finally processing the command to execute the desired action.

Skills in voice AI assistants are applications that enable the assistant to perform various tasks, such as checking the weather, playing music, or managing a bank account.

Most users do not realize whether they are interacting with native or third-party skills when using voice AI assistants, as this distinction is not visible during interactions.

Voice AI assistants process user commands in the cloud, where natural language processing methods are used to understand the user's intent and select the appropriate skill.

Voice assistants are technically always listening for the wake-up word, but they do not record or process any data until the activation word is spoken by the user.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.