The new Apple Watch Series 12 and Ultra 4 come not just with better fitness tracking and upgraded noise reduction, but also a whole new way to listen.
Apple announced on Wednesday that the two smartwatches come with four opt-in “audio intelligence” tools that are powered by audio gathered by the watches’ microphones. They include sound and music recognition, a conversation recap feature, and “Live Rewind” so a user can see a transcription of anything that was said in their environment in the previous 15 seconds.
Seemingly wary that the features could make it feel like the walls have ears, Apple is emphasizing that all of them are built to prioritize privacy and security using the extensive infrastructure the company has developed to protect its other Apple Intelligence and Siri services. As more and more of these types of features debut, though, the audio intelligence announcements are a reminder that AI-powered tools are becoming increasingly universal in all areas of computing and life.
“These features do not create or store audio recordings, and raw audio used for processing is completely inaccessible to the operating systems, apps, the user, or Apple,” the company wrote in a report shared with WIRED.
The new Apple Watch features are designed to process as much data locally on devices as possible, so potentially sensitive data doesn’t go to the cloud. Sound Recognition, for example, alerts users about sounds in their environment—including doorbells, sirens, alarms, or, say, a baby crying—without sending any data off the watch. The tools that do send information out to the cloud preprocess the data so it isn’t raw audio files, then use Apple’s Private Cloud Compute infrastructure.
Apple emphasizes that its on-device capabilities for Apple watches have expanded thanks to the company’s new S11 chips. The chips include a special memory-protected space that Apple calls the Secure Exclave. Designed specifically for sensor data, the exclave is an isolated buffer that is inaccessible to the rest of the operating system where Apple Watch can hold and process audio data in a protected space.
A new Shazam feature “listens” for music and will generate a signature of a song that is stored in the Secure Exclave. If the user navigates to Shazam to find out what the music is, the tool will send that signature—not an audio file—to Shazam’s servers for identification. (Shazam is an Apple-owned service.) If the user doesn’t prompt to identify the music, or as soon as they do, the special signature is “immediately” deleted from the watch.
The recap feature, known as Siri Recap, can be set to be on all the time and generate recaps of any substantive conversation, or it can be given a schedule to listen in at certain times during the day. A dedicated AI model determines when speech is occurring without recording or transcribing any data. If it detects a conversation, audio goes into a protected buffer within the Secure Exclave on the Apple Watch, the watch encrypts it, and then transmits it using Apple’s special secure Bluetooth pairing to the Secure Exclave on a user’s iPhone. Then the audio is immediately deleted from the Apple Watch.
The iPhone uses local speech recognition and language models on-device to transcribe the audio and then generate a minimal version of the transcript that removes nonessential elements like extra words and repeated phrases. Then the raw audio is immediately deleted from the iPhone as well. The last step on the iPhone is a safety model that “screens the text to omit potentially harmful terms,” according to Apple. From there, the distillation of the conversation is encrypted, and the iPhone sends it out to Private Cloud Compute.
Leave a comment