Listening.

Design Lead · making seven audio platforms feel like one effortless way to listen — by voice, on your face.

Why it matters
Most-used
the listening experience became a core daily use case across 9M+ Ray-Ban Meta glasses.
7 partners
spanning different platform capabilities, unified into one consistent voice experience.
Connect '24
presented by Mark Zuckerberg in the keynote — a strategic, company-level bet.
The Challenge

From user signal to product strategy.

User Needs

The demand was already there.

Before voice listening existed, users were already asking the glasses to play music. It was a top-5 assistant request.

Because this behavior was not prompted by UI, the signal was clear: users were trying voice before the product could reliably answer.
Assistant query ranking showing play music as a top user request
Business Needs

Voice first

Voice was the primary interaction bet for glasses, not a secondary control.

Build habit

Listening was already a strong daily behavior, making it the clearest place to build a voice habit.

Audio ecosystem

Spotify proved the model. The goal was a broader audio ecosystem, so more partners kept signing on.

The goal
Bring voice to listening —
and scale it to 7 partners.
The team
1
designer
8
teams
30+
xfn
Timeline
2024.3 - 2024.9 from kickoff to all 7 partners live at Meta Connect
Spotify Apple Music Amazon Music iHeartRadio Audible Shazam Calm
The Outcome

shipped a voice-first listening system that let people play, discover, recognize, and route audio across 7 partners.

The Approach

Three decisions shaped voice, memory, and ecosystem complexity into one effortless listening system.

Decision 01 Why should listening use voice?

Use voice only where it is better than touch.

Ray-Ban Meta is not a phone replacement. 70% of glasses were sunglasses, and many listening moments happened outdoors, in motion, or in situations where pulling out a phone breaks the moment. Voice was strongest when it let people keep doing what they were already doing.

What makes voice valuable to users?

Quick and hands-free

Voice starts playback without pulling users out of the activity.

Stay in the moment

Users can express intent without stopping to search, scroll, or tap.

Beyond the phone

The glasses can use context a phone cannot easily capture.

Running

Mid-run, the song no longer fits. Reaching for a phone breaks the rhythm.

"Hey Meta, play my workout playlist"

Roadtrip

On the road, the user wants something new, but only has a mood in mind.

"Hey Meta, play something chill"

Hiking

At the viewpoint, the scene is beautiful. The user wants music that matches it.

"Hey Meta, look and play a song"
Define use cases

Based on those principles, the work defined the valuable listening intents and the voice commands the system needed to support.

Hands-free Control

"play · pause · skip · rewind"

Search

"play Taylor Swift"
"play my Daily Calm"

Personalization

"like this song"

Immersive

"look and play"

Music Recognition · Shazam

"what's this song?"

Discover

"play some chill music"
"play a podcast about history"

Sharing

"share this song with Kevin"

Decision 02 How do people remember what to say?

Make voice simple enough to remember and trust.

Voice has no visible button. If users cannot remember what to say, the feature disappears. Commands that did not reinforce the core habit were deprioritized, centering the MVP around one word: play.

The command stem users needed to remember — and the primary MVP language.
"play ..."
"what's the song?" Kept for Shazam recognition
"I like this song" Deprioritized
"look and play" Deprioritized
"share this song" Deprioritized
Teaching the habit

Make voice simple to say, hear, and trust.

01

Marketing

Launch surfaces showed the most useful “play” moments, so users could understand the voice concept before trying it.

Marketing surfaces introducing voice listening use cases
02

Education

In-product guidance made the voice language visible while the habit was still forming.

In-product education for voice listening
Follow-up education for voice listening commands
03

Connection flow

Partner setup became the first lesson in how listening by voice would work.

Spotify connected state in the partner setup flow
04

Make every response consistent enough to understand by ear.

The work defined both sides of the conversation: what the system should map from a user's phrasing, and how the assistant should respond when playback succeeds, fails, or needs setup. The goal was habit and trust, not just command coverage.

Accuracy

Map different user phrasing to a known playback action.

Clarity

Use consistent response patterns so users can understand the outcome quickly.

Control

Avoid free-form generation to save time, reduce cost, and keep quality predictable.

User intent
"Hey Meta, play a running mix"
Consistent response
"From [provider], here is [title]"

Keep experience quality in our hands, not leave it to the model.

Decision 03 What happens as more apps connect?

Turn partner ambiguity into a scalable system.

As the audio ecosystem grew from one partner to seven, provider choice became a system-level problem: the same “play” request could map to different services, content types, and setup states. With limited PM bandwidth, I stepped in to lead the resolution workstream across design, product, and engineering — defining the default logic, aligning the team, and helping launch the system at Meta Connect.

Partner coverage

The same voice request could point to different services.

Each partner covered a different mix of music, podcasts, audiobooks, and radio. Once users connected more than one service, “play” needed a quiet way to choose the right provider.

Partner coverage across music, podcast, audiobook, and radio categories

When a user says, “Hey Meta, play music,” and several connected providers can serve it,which one should answer?

Key assumption

Connection order can signal user preference.

Instead of asking users to configure defaults before they understood the ecosystem, I treated the first connected capable provider as an early preference signal. That gave the system a simple rule to resolve ambiguity sooner, while still leaving room for users to change it later.

The resolution layer

Each category defaults to the first connected provider that can serve it. Set silently, once.

Music, podcast, audiobook, and radio defaults are assigned as users connect partners. Later providers fill empty categories, while overlaps keep the original default unless the user explicitly changes it.

Resolution layer diagram showing how connection order assigns audio defaults by category
Invisible screen

Keep control available, but design for most users to never need it.

The defaults screen existed for transparency, recovery, and manual changes. But the main experience was designed to resolve ambiguity silently, before users hit a conflict.

Connected apps and assign defaults screens for audio provider preferences
Impact
01

Adoption

Voice listening saw strong early adoption and repeat use, showing it delivered value people came back to. Creators also started using music voice commands across a wide range of everyday scenarios.

02

Resolution

All 7 providers went live by Connect. For users with multiple providers connected, the resolution layer usually returned the right provider in the background, while only a small minority ever opened settings to change it.

Later launches

After MVP adoption, we launched follow-up features from the original use cases and extended listening onto AR glasses.

Like songs

Personalization came back as a lightweight voice action after the core listening habit was established.

Contextual listening

“Look and play” used what the glasses could see to match music to the moment.

AR music

The listening system later extended onto Ray-Ban Meta Display as a gesture-driven AR music UI.

The takeaway

The complexity was never a single screen. It was making seven platforms feel like one effortless way to listen.