goJumboGPT

Devices Smart home and connected devices

Smart speakers: what they hear and what gets stored

How wake word detection works, what actually leaves the device, which recordings are kept, and the settings to change on day one.

8 min read How we write

The short answer

  • A smart speaker listens continuously but only transmits once its wake word matches, so the real privacy issue is false wakes rather than constant recording.
  • Each interaction is usually stored as both an audio clip and a text transcript attached to your account until you change the setting.
  • Turning off audio storage does not stop audio being sent, it stops it being kept, and transcripts are often a separate control.
  • A small sample of clips is reviewed by people to improve recognition, and every major assistant now lets you opt out of that.
  • Voice matching is a convenience feature rather than a security control, so disable voice purchasing or protect it with a code.
  • The mute button physically cuts the microphone circuit, which makes it the most reliable control on the device.

A smart speaker listens all the time but records almost none of it. The microphone is live, and a small piece of software on the device itself checks a couple of seconds of sound over and over for one specific pattern: the wake word. Nothing leaves the speaker until that pattern matches. When it does, the device streams to the company's servers, and that stream is what gets transcribed, answered and often kept. So the privacy question is not whether it hears you, because it does. It is how often it wrongly decides you said its name, and what happens to that recording afterward.

What always listening actually means

Inside the speaker is a low power chip running a tiny detection model. It holds a rolling buffer of a second or two of audio in memory, compares it against the sound pattern of the wake word, then overwrites it. That buffer is never stored or transmitted. It works this way for practical reasons as much as principled ones: streaming a household's continuous audio would cost a fortune in bandwidth and would be obvious on any network monitor.

When the pattern matches, three things happen at once. The indicator light turns on, the device opens a connection, and it sends a short piece of audio from just before the match so the start of your sentence is not clipped. From then until it decides you have finished, the audio goes to the servers, where speech recognition turns it into text and a system works out what you wanted.

StageWhere it happensWhat is kept
Listening for the wake wordOn the device, in memoryNothing, the buffer is continuously overwritten
Wake word matchedOn the deviceA short pre-roll of audio, sent with the request
Your commandStreamed to the company serversThe audio clip, usually attached to your account
Transcription and answerCompany serversA text transcript and the action taken
Activity historyYour accountClips and transcripts until you delete them or auto delete removes them
Quality review samplingCompany staff or contractorsA small proportion of clips, if you have not opted out

The indicator light is the detail worth learning. On every mainstream speaker it is wired to the streaming state, so a lit ring means audio is going out. If it lights up when nobody spoke to it, that is a false wake, and you can go and look at exactly what it heard.

False wakes and what they capture

Wake word detection trades off missing you when you call against triggering when you did not. Tuned strictly it feels broken, so makers tune it loose, and the cost is false activations. Words that rhyme or share a rhythm set it off, as do television dialogue, a similar sounding name and background conversation at the wrong volume.

What gets captured is not a whole conversation. It is a clip that runs from a moment before the false trigger until the device gives up, typically a few seconds. But a few seconds of an unguarded conversation can contain a name, an address, a health detail or a password read out loud, and that clip is stored against your account exactly like a deliberate request.

The useful response is to check rather than to worry. Every major assistant has an activity view listing each interaction with the audio attached. Open it once and play the entries you do not recognize. You will learn how often false wakes happen in your home, usually less than the headlines suggest and more than you assumed, and whether the room you chose was a good idea. A speaker facing a television is a false wake machine.

What gets stored and for how long

By default, most assistants keep both the audio clip and its transcript against your account, so recognition can be improved and so you can review what was asked. Retention is now a setting rather than a fixed policy, and the options are broadly the same everywhere: keep recordings until you delete them, delete automatically after a set period of a few months or longer, or do not save audio at all.

The last option deserves explanation, because it is not a magic switch. Choosing not to save audio does not mean the audio is never sent. It is still streamed to be understood, it is simply not retained afterward. Text transcripts and the record of what you asked for often persist separately, because they are what powers your reminder list, your shopping list and the history a household member can see. If you want both gone, you usually need to turn off audio storage and clear the activity history, which are two different controls in two different places.

Assistants are also increasingly built on large language models, which changes the shape of the question. Your spoken request becomes a prompt sent to a model, and the rules that apply to a typed chatbot conversation now apply to your kitchen. It is worth knowing what an AI assistant does with personal data you give it, and separately where prompts are processed once you press send, because the answer decides which controls exist at all.

Deletion also means different things in different places. A clip removed from your history may persist in backups for a while, and aggregated statistics derived from it usually do not disappear. That is normal across the industry rather than a dark pattern, but it is worth understanding what deleting AI data actually removes before assuming a cleared history is the end of it.

Why people listen to some clips

Speech recognition improves by being fed examples of where it went wrong, and the most effective way to find those examples has been to have people listen. Every large assistant has run a program in which a small sample of clips is transcribed or graded by employees or contractors, stripped of the account name but not of the content.

This became controversial because people learned about it from news reports rather than the setup screen, and the outcome was worth having: review programs are now disclosed and every major assistant offers an opt out, often phrased as helping improve the service. Turning it off costs nothing you will notice and removes the chance that a stranger hears a clip captured by mistake. It is commonly on by default, so look for it.

Voice profiles, guests and children

Voice matching links a particular voice to a particular account, so that asking about your calendar returns yours and not your partner's. It is genuinely useful and it is also a second thing the device now knows about you: a voice model, stored with your account.

The failure mode surprises people. Voice matching is a convenience feature, not a security control. It can be fooled by a similar voice and sometimes by a recording, so it should never be the only thing between a visitor and your messages or your card. Set a spoken code or disable voice purchasing, and check what an unrecognized voice is allowed to do, because the device answers whoever is in the room.

Children are the clearest case. A child can ask the speaker anything, order things, call contacts and hear results that were never meant for them, so the family controls are worth setting deliberately rather than discovering later. The broader arrangement of shared devices, accounts and limits is easier to do once, and a family technology setup that prevents most problems covers how to make those choices stick. If the assistant is a conversational model rather than a command system, the extra considerations in what parents should set up around AI apply to it too.

Guests raise a smaller but real question. Most people do not expect a microphone in a spare bedroom, and the polite answer is to say so or not to put one there. In some places it is also a legal question: recording conversations without consent is restricted in several US states and across much of Europe, and although a wake word device is not a recorder in the ordinary sense, the courtesy is the same.

The day one settings pass

Twenty minutes, once, for each assistant account.

  1. Open the activity or voice history and read it. This tells you more about your own device than any setting description.
  2. Turn off audio storage, or set automatic deletion to the shortest period offered.
  3. Clear the existing history, which is separate from changing the retention setting.
  4. Opt out of human review and quality improvement sampling.
  5. Disable voice purchasing, or require a spoken code for it.
  6. Review the connected apps and services with access to the assistant and remove anything you no longer use.
  7. Check who else has access to the household account, including anyone who moved out.
  8. Turn on automatic updates so the device keeps receiving fixes.

Then make two placement decisions. Keep speakers out of bedrooms and bathrooms unless you have a reason, and away from the television if false wakes bother you. Any device with a screen and camera should be treated as a camera first: use the physical shutter when you are not calling, because keeping a camera feed away from everybody else is a different problem from keeping a microphone quiet.

None of this makes a smart speaker private, and it is not meant to. It makes the trade explicit: hands free control in exchange for a microphone in the room and a request history in a company account. If that looks reasonable once the settings are right, keep it. If not, the same lights, plugs and sensors work fine from a phone app or a wall switch, and a smart home built around sensors and lighting rather than voice loses very little.

Common questions

Is my smart speaker recording everything I say?

No. The always on part is a small detector running on the device that checks a rolling couple of seconds for the wake word and then discards them. Audio is only sent once that pattern matches, which the indicator light shows. The real exposure is false wakes, where a few seconds of conversation is captured because the device thought it heard its name.

How do I delete my smart speaker recordings?

In the assistant app, open the voice or activity history, delete the existing entries, then change the retention setting so new ones are not kept. Those are two separate actions and doing only the first leaves the setting unchanged. Some assistants also accept a spoken command to delete what you just said or the last day of activity.

Can people at the company listen to my recordings?

A small sample can be reviewed by staff or contractors to improve speech recognition, and that has been true of every large assistant. It is now disclosed rather than hidden, and each of them offers an opt out, sometimes worded as helping improve the service. Switching it off costs you nothing noticeable and is worth doing during setup.

Does the mute button really work?

On mainstream speakers, yes. It is a hardware cut to the microphone circuit rather than a software toggle, which is why the device stops responding completely and shows a fixed red light. That makes it the one control that a software update cannot quietly change. Use it when a conversation genuinely should not be near a microphone.

Should I put a smart speaker in a bedroom?

That is a personal call, but two things should inform it. False wakes capture short clips from wherever the device sits, so an unguarded room raises the stakes of each mistake. And guests generally do not expect a microphone in a room they sleep in. If the only reason is an alarm or music, an ordinary speaker does both with nothing listening.