Assertion certainty 4/5 debate potential 1/5

Staniszewski: ElevenLabs voice models operate on text and audio tokens simultaneously

Mati Staniszewski · The world of voice AI, with Mati Staniszewski of ElevenLabs · Apr 14, 2026 · at 5:45

ElevenLabs co-founder Mati Staniszewski explains to Stripe's John Collison how voice models represent speech tokens and maintain context during real-time streaming.

0:00 / 0:27exact quote · 27.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“In our models now, it's going to be a combination of not only operating on phoneme level, you also operate on the text level. You operate kind of, In both in sync, because when you are predicting the context, you need to understand how that sentence will get constructed, and especially if it's more of a streaming real-time use case in like a voice agent setting, you need both parts to work across. So it's similar to how you would operate on the token level on the tech side, we operate on the token level on the audio side.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mati Staniszewski

Prediction Not checkable as stated
Staniszewski: ElevenLabs hopes to pass voice Turing test within a year
“Our goal is to like pass the voice Turing test in all those cases, or the Turing test for all conversational agents outside of voice two. And I hope we will be there in the next year or so.”
Mati Staniszewski Apr 14, 2026 ▶ 20:44 The world of voice AI, with Mati Staniszewski of ElevenLabs
Disclosure
ElevenLabs bets research on cascaded voice architecture over speech-to-speech
“So today we are optimizing heavily on a cascaded approach... So that's like where we are betting a lot of the research work of how you can make that great, and we think we can make that great.”
Mati Staniszewski Apr 14, 2026 ▶ 28:33 The world of voice AI, with Mati Staniszewski of ElevenLabs
Assertion Not checkable as stated
Staniszewski: Speech-to-speech models are 'definitely dumber' due to size limits
“They are definitely dumber. You need smaller model, you cannot... If you are going speech to speech, usually you will use smaller models, so it's still quick.”
Mati Staniszewski Apr 14, 2026 ▶ 30:01 The world of voice AI, with Mati Staniszewski of ElevenLabs
Assertion Not checkable as stated
Staniszewski: Accents and emotion are emergent properties in ElevenLabs models
“And in our approach, effectively, you would give the model open-ended ability to select what those parameters should be. So it's not going to be British, Polish, Spanish, English speaker but the model will deduce them themselves. The same for other set of para…”
Mati Staniszewski Apr 14, 2026 ▶ 4:00 The world of voice AI, with Mati Staniszewski of ElevenLabs
Assertion Supported
ElevenLabs expressive mode lets voice agents detect and match human emotions
“So today, finally, you can have both speed generation or entire voice agent experience With what we call expressive mode where the agent knows the emotions on the other side. So if the person is stressed, it can react and be reassuring and that's generating a …”
Mati Staniszewski Apr 14, 2026 ▶ 26:57 The world of voice AI, with Mati Staniszewski of ElevenLabs
Insight
Staniszewski: Solving conversational voice AI inherently solves text agents too
“If you fix the voice agent, you'll have fixed text piece or like inherently as well.”
Mati Staniszewski Apr 14, 2026 ▶ 43:07 The world of voice AI, with Mati Staniszewski of ElevenLabs
Made with StarZero

Turn any episode into a week of clips.

This entire site, about 28 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.