Apple announced on Wednesday that its two latest smartwatches come equipped with four opt-in audio intelligence tools powered by the microphones on the devices. These features include sound and music recognition, a conversation recap tool, and Live Rewind, which allows a user to view a transcription of environmental speech from the previous 15 seconds.
Seemingly aware that these capabilities could raise concerns about constant surveillance, Apple is emphasizing that all of them are built to prioritize privacy and security. The company is leveraging the extensive infrastructure it developed for its other Apple Intelligence and Siri services. As these types of artificial intelligence features continue to emerge across the technology landscape, they serve as a reminder of how ubiquitous AI-driven tools are becoming in modern computing and daily life.
“These features do not create or store audio recordings, and raw audio used for processing is completely inaccessible to the operating systems, apps, the user, or Apple,” the company stated in a report shared with researchers.
On-Device Processing and Secure Enclave
The new Apple Watch features are engineered to process as much data locally on the hardware as possible to prevent sensitive information from traveling to the cloud. Sound Recognition, for instance, notifies users about environmental noises such as doorbells, sirens, alarms, or a baby crying without transmitting any data off the watch. Tools that do send information to the cloud pre-process the data so it is not a raw audio file, and they utilize Apple’s Private Cloud Compute infrastructure.
Apple highlights that its on-device capabilities have expanded significantly due to the integration of new S11 chips. These chips feature a special memory-protected space known as the Secure Enclave. Designed specifically for sensor data, this isolated buffer remains inaccessible to the rest of the operating system, allowing the Apple Watch to hold and process audio data within a secure environment.
A new Shazam feature listens for music and generates a unique song signature stored within the Secure Enclave. If a user opens Shazam to identify the track, the tool transmits only that digital signature rather than an audio file to Shazam servers. As soon as the music is identified or the prompt is abandoned, the signature is immediately deleted from the watch.
Siri Recap and Conversation Summaries
The conversation recap feature, called Siri Recap, can operate continuously to summarize substantive discussions or be scheduled to listen during specific times of the day. A dedicated AI model determines when speech occurs without recording or transcribing any raw data. When a conversation is detected, audio enters a protected buffer inside the Secure Enclave on the Apple Watch, where it is encrypted and transmitted via secure Bluetooth pairing to the Secure Enclave on a paired iPhone. The audio is then deleted instantly from the Apple Watch.
The iPhone utilizes local speech recognition and language models to transcribe the audio and generate a condensed version that strips away nonessential elements like filler words and repetition. The raw audio is then deleted from the iPhone immediately. A final safety model on the iPhone screens the text to remove potentially harmful terms, according to Apple. The distilled conversation text is then encrypted and sent to Private Cloud Compute.
“Contextual information is also sent to improve summary quality,” Apple explains. “For example, Now Playing data helps the model understand that speech may have originated from music or a podcast. Calendar data can help generate accurate titles of summaries. High-level location labels such as home, work, and school, along with locality information such as city, state, and country, add context to the summarization. Point-of-interest categories like ‘grocery store’ or ‘park’ may also be included. Precise location and specific points-of-interest are not included.”
The condensed transcript passes through foundation models in Private Cloud Compute to generate a title and summary, which is encrypted and sent back to the paired devices. Apple notes that Siri Recap is engineered to strip away sensitive details such as financial information, identification numbers, and other personal identifiers from the final summaries.
Live Rewind and Additional Safeguards
The Live Rewind feature, which provides a text transcript of the last 15 seconds of ambient audio, utilizes the Secure Enclave buffer to hold sound on a rolling basis where older audio is continuously overwritten. To trigger a transcript, users must double-press the Apple Watch Digital Crown. Once activated, the watch sends the buffered audio to the paired iPhone. If the iPhone is unavailable, the transfer fails and the audio is automatically deleted.
If the transfer succeeds, the iPhone performs local speech recognition processing to generate the transcript, deletes the audio, and sends the text back to the watch. Users can then read the transcript and choose to discard it or save it within the Siri app. Despite these extensive safeguards, observers note that wearing such a device introduces potential privacy risks during private conversations or accidental slips of the tongue.
To mitigate potential privacy violations, Apple states that the watch emits an audible chime whenever a user initiates a Live Rewind transcript, alerting individuals nearby even if the device is muted or connected to headphones. Furthermore, the features are designed to filter out identifying information about speakers. However, the sheer scope of these audio tools highlights how the broader AI attack surface continues to expand as technology advances.


















