Apr 14, 2026 · 1h 0m · cheeky-pint

The world of voice AI, with Mati Staniszewski of ElevenLabs

Mati Staniszewski · 38m spoken John Collison · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview, Stripe co-founder John Collison speaks with ElevenLabs co-founder and CEO Mati Staniszewski about the architecture, rapid scaling, and real-world impact of frontier voice AI. Staniszewski explains how proprietary data annotation, synchronized audio modeling, and an agile, AI-native organizational structure propelled ElevenLabs to over $450M in ARR.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. John holds 31.5% of the talking time here. How this is scored →

John as informed peer 5.7 Guest teaching 4.3 Guest disagreement 1.2 John pushing back 1.8
05100:0015:0030:0045:001:00:000:21–4:48 · John as informed peer 5/10 How Audio AI Models Work Under the Hood Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities.4:48–8:03 · John as informed peer 4/10 Audio Tokens, Phonemes, and Real-Time Synchronization Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development.8:04–10:44 · John as informed peer 5/10 Historical Parallels: Wolfgang von Kempelen's Mechanical Turk Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown.10:44–15:41 · John as informed peer 7/10 Platform Strategy and Horizontal Infrastructure vs. Vertical Apps Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances.15:41–17:50 · John as informed peer 5/10 Eleven Reader and Consumer Voice Applications Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks.17:51–20:58 · John as informed peer 6/10 The Conversational Voice Turing Test and Turn-Taking Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling.21:02–25:26 · John as informed peer 6/10 Stripe Link Sponsor Segment Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities.25:27–27:40 · John as informed peer 5/10 Controllability, Expressive Modes, and Speech Generation Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues.27:42–30:23 · John as informed peer 7/10 Cascaded Architectures vs. End-to-End Speech-to-Speech Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints.30:24–34:03 · John as informed peer 5/10 Behavioral Dynamics: Voice Interfaces vs. Static Web Forms Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games.34:03–40:11 · John as informed peer 7/10 Proactive Voice Agents and the Guinness Index Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models.40:12–44:08 · John as informed peer 6/10 Enterprise Conversational Agents in Sales and Support Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning.44:08–47:50 · John as informed peer 6/10 Scaling to $450M+ ARR and Small Autonomous Teams Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure.47:51–51:41 · John as informed peer 7/10 Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition.51:41–55:16 · John as informed peer 6/10 Organizational Design and Flatter Teams in the AI Era Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads.55:16–58:37 · John as informed peer 4/10 AI-Native Internal Tooling and Ukraine's Diia App Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents.58:38–1:00:17 · John as informed peer 6/10 Cultivating High Agency and Conclusion Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era.0:21–4:48 · Guest teaching 6/10 How Audio AI Models Work Under the Hood Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities.4:48–8:03 · Guest teaching 5/10 Audio Tokens, Phonemes, and Real-Time Synchronization Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development.8:04–10:44 · Guest teaching 3/10 Historical Parallels: Wolfgang von Kempelen's Mechanical Turk Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown.10:44–15:41 · Guest teaching 4/10 Platform Strategy and Horizontal Infrastructure vs. Vertical Apps Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances.15:41–17:50 · Guest teaching 4/10 Eleven Reader and Consumer Voice Applications Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks.17:51–20:58 · Guest teaching 4/10 The Conversational Voice Turing Test and Turn-Taking Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling.21:02–25:26 · Guest teaching 4/10 Stripe Link Sponsor Segment Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities.25:27–27:40 · Guest teaching 5/10 Controllability, Expressive Modes, and Speech Generation Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues.27:42–30:23 · Guest teaching 5/10 Cascaded Architectures vs. End-to-End Speech-to-Speech Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints.30:24–34:03 · Guest teaching 5/10 Behavioral Dynamics: Voice Interfaces vs. Static Web Forms Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games.34:03–40:11 · Guest teaching 4/10 Proactive Voice Agents and the Guinness Index Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models.40:12–44:08 · Guest teaching 4/10 Enterprise Conversational Agents in Sales and Support Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning.44:08–47:50 · Guest teaching 4/10 Scaling to $450M+ ARR and Small Autonomous Teams Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure.47:51–51:41 · Guest teaching 3/10 Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition.51:41–55:16 · Guest teaching 4/10 Organizational Design and Flatter Teams in the AI Era Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads.55:16–58:37 · Guest teaching 6/10 AI-Native Internal Tooling and Ukraine's Diia App Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents.58:38–1:00:17 · Guest teaching 3/10 Cultivating High Agency and Conclusion Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era.0:21–4:48 · Guest disagreement 1/10 How Audio AI Models Work Under the Hood Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities.4:48–8:03 · Guest disagreement 1/10 Audio Tokens, Phonemes, and Real-Time Synchronization Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development.8:04–10:44 · Guest disagreement 1/10 Historical Parallels: Wolfgang von Kempelen's Mechanical Turk Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown.10:44–15:41 · Guest disagreement 3/10 Platform Strategy and Horizontal Infrastructure vs. Vertical Apps Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances.15:41–17:50 · Guest disagreement 1/10 Eleven Reader and Consumer Voice Applications Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks.17:51–20:58 · Guest disagreement 2/10 The Conversational Voice Turing Test and Turn-Taking Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling.21:02–25:26 · Guest disagreement 1/10 Stripe Link Sponsor Segment Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities.25:27–27:40 · Guest disagreement 1/10 Controllability, Expressive Modes, and Speech Generation Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues.27:42–30:23 · Guest disagreement 1/10 Cascaded Architectures vs. End-to-End Speech-to-Speech Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints.30:24–34:03 · Guest disagreement 1/10 Behavioral Dynamics: Voice Interfaces vs. Static Web Forms Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games.34:03–40:11 · Guest disagreement 1/10 Proactive Voice Agents and the Guinness Index Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models.40:12–44:08 · Guest disagreement 1/10 Enterprise Conversational Agents in Sales and Support Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning.44:08–47:50 · Guest disagreement 1/10 Scaling to $450M+ ARR and Small Autonomous Teams Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure.47:51–51:41 · Guest disagreement 1/10 Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition.51:41–55:16 · Guest disagreement 2/10 Organizational Design and Flatter Teams in the AI Era Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads.55:16–58:37 · Guest disagreement 1/10 AI-Native Internal Tooling and Ukraine's Diia App Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents.58:38–1:00:17 · Guest disagreement 0/10 Cultivating High Agency and Conclusion Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era.0:21–4:48 · John pushing back 2/10 How Audio AI Models Work Under the Hood Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities.4:48–8:03 · John pushing back 1/10 Audio Tokens, Phonemes, and Real-Time Synchronization Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development.8:04–10:44 · John pushing back 1/10 Historical Parallels: Wolfgang von Kempelen's Mechanical Turk Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown.10:44–15:41 · John pushing back 5/10 Platform Strategy and Horizontal Infrastructure vs. Vertical Apps Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances.15:41–17:50 · John pushing back 2/10 Eleven Reader and Consumer Voice Applications Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks.17:51–20:58 · John pushing back 3/10 The Conversational Voice Turing Test and Turn-Taking Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling.21:02–25:26 · John pushing back 2/10 Stripe Link Sponsor Segment Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities.25:27–27:40 · John pushing back 1/10 Controllability, Expressive Modes, and Speech Generation Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues.27:42–30:23 · John pushing back 2/10 Cascaded Architectures vs. End-to-End Speech-to-Speech Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints.30:24–34:03 · John pushing back 1/10 Behavioral Dynamics: Voice Interfaces vs. Static Web Forms Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games.34:03–40:11 · John pushing back 2/10 Proactive Voice Agents and the Guinness Index Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models.40:12–44:08 · John pushing back 2/10 Enterprise Conversational Agents in Sales and Support Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning.44:08–47:50 · John pushing back 1/10 Scaling to $450M+ ARR and Small Autonomous Teams Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure.47:51–51:41 · John pushing back 2/10 Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition.51:41–55:16 · John pushing back 3/10 Organizational Design and Flatter Teams in the AI Era Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads.55:16–58:37 · John pushing back 1/10 AI-Native Internal Tooling and Ukraine's Diia App Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents.58:38–1:00:17 · John pushing back 0/10 Cultivating High Agency and Conclusion Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era.

speaking balance: gold is John, purple is the guest (3 minute bins)

0:00 · John 17.3% · guest 82.7%0:00 · John 17.3% · guest 82.7%3:00 · John 14.6% · guest 85.4%3:00 · John 14.6% · guest 85.4%6:00 · John 14.7% · guest 85.3%6:00 · John 14.7% · guest 85.3%9:00 · John 34.4% · guest 65.6%9:00 · John 34.4% · guest 65.6%12:00 · John 53.5% · guest 46.5%12:00 · John 53.5% · guest 46.5%15:00 · John 33.2% · guest 66.8%15:00 · John 33.2% · guest 66.8%18:00 · John 29.1% · guest 70.9%18:00 · John 29.1% · guest 70.9%21:00 · John 52.8% · guest 47.2%21:00 · John 52.8% · guest 47.2%24:00 · John 34.1% · guest 65.9%24:00 · John 34.1% · guest 65.9%27:00 · John 24.1% · guest 75.9%27:00 · John 24.1% · guest 75.9%30:00 · John 38.9% · guest 61.1%30:00 · John 38.9% · guest 61.1%33:00 · John 40.8% · guest 59.2%33:00 · John 40.8% · guest 59.2%36:00 · John 26.2% · guest 73.8%36:00 · John 26.2% · guest 73.8%39:00 · John 25.1% · guest 74.9%39:00 · John 25.1% · guest 74.9%42:00 · John 39% · guest 61%42:00 · John 39% · guest 61%45:00 · John 25.8% · guest 74.2%45:00 · John 25.8% · guest 74.2%48:00 · John 27.4% · guest 72.6%48:00 · John 27.4% · guest 72.6%51:00 · John 72.1% · guest 27.9%51:00 · John 72.1% · guest 27.9%54:00 · John 15.4% · guest 84.6%54:00 · John 15.4% · guest 84.6%57:00 · John 9.3% · guest 90.7%57:00 · John 9.3% · guest 90.7%1:00:00 · John 41.3% · guest 58.7%1:00:00 · John 41.3% · guest 58.7%
Sharpest disagreement ▶ 13:57 Rejection of the ten-year UX lag claim

Staniszewski politely but directly pushes back against Collison's framing that voice AI UX is stuck ten years behind modern LLMs.

Hardest push from John ▶ 53:45 Pushing back on extreme spans of control

Collison playfully challenges whether having 15+ direct reports per executive is genuine AI leverage or just early-stage founder management theory.

Biggest teaching moment ▶ 2:15 Explaining Mel Spectrogram and audio modeling

Staniszewski breaks down the core pipeline of text, Mel Spectrogram, and waveform synthesis, clarifying concepts Collison asked him to unpack.

John holds their own ▶ 44:46 Collison triangulates revenue trajectory

Collison uses quick mental math to convert Staniszewski's net new quarterly metrics into a sharp $450M+ run rate figure.

the scores for every segment, with the reasoning behind each
ChapterTopicJohn as informed peerGuest teachingGuest disagreementJohn pushing backWhy
How Audio AI Models Work Under the Hood 5612 Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities.
Audio Tokens, Phonemes, and Real-Time Synchronization 4511 Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development.
Historical Parallels: Wolfgang von Kempelen's Mechanical Turk 5311 Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown.
Platform Strategy and Horizontal Infrastructure vs. Vertical Apps 7435 Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances.
Eleven Reader and Consumer Voice Applications 5412 Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks.
The Conversational Voice Turing Test and Turn-Taking 6423 Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling.
Stripe Link Sponsor Segment 6412 Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities.
Controllability, Expressive Modes, and Speech Generation 5511 Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues.
Cascaded Architectures vs. End-to-End Speech-to-Speech 7512 Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints.
Behavioral Dynamics: Voice Interfaces vs. Static Web Forms 5511 Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games.
Proactive Voice Agents and the Guinness Index 7412 Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models.
Enterprise Conversational Agents in Sales and Support 6412 Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning.
Scaling to $450M+ ARR and Small Autonomous Teams 6411 Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure.
Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go 7312 Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition.
Organizational Design and Flatter Teams in the AI Era 6423 Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads.
AI-Native Internal Tooling and Ukraine's Diia App 4611 Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents.
Cultivating High Agency and Conclusion 6300 Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era.

Statements from this episode (33)

Assertion Not checkable as stated
Staniszewski: ElevenLabs applied transformer and diffusion concepts to voice
“And here credit to my co-founder, Piotr, who effectively came with that new idea of how you can now create voice models, which are both reliable, high quality, quick, where you would bring a lot of the ideas from transformer models, from diffusion models into …”
Mati Staniszewski Apr 14, 2026 ▶ 1:47
Assertion Not checkable as stated
Staniszewski: Accents and emotion are emergent properties in ElevenLabs models
“And in our approach, effectively, you would give the model open-ended ability to select what those parameters should be. So it's not going to be British, Polish, Spanish, English speaker but the model will deduce them themselves. The same for other set of para…”
Mati Staniszewski Apr 14, 2026 ▶ 4:00
Assertion Not checkable as stated
Staniszewski: ElevenLabs voice models operate on text and audio tokens simultaneously
“In our models now, it's going to be a combination of not only operating on phoneme level, you also operate on the text level. You operate kind of, In both in sync, because when you are predicting the context, you need to understand how that sentence will get c…”
Mati Staniszewski Apr 14, 2026 ▶ 5:45
Disclosure
ElevenLabs built internal speech-to-text models after market options fell short
“So speech-to-text model, Initially it was a model we did for ourselves because the models on the market just weren't good to annotate that data.”
Mati Staniszewski Apr 14, 2026 ▶ 7:16
Assertion Partly supported
Staniszewski says ElevenLabs speech-to-text models beat industry benchmarks across 100 languages
“Speech to text models that work over a hundred languages and happily beat others on benchmarks all the way through to conversational models of how you loop them together to music, to other domains of audio.”
Mati Staniszewski Apr 14, 2026 ▶ 9:43
Disclosure
Staniszewski says ElevenLabs strictly operates as a horizontal platform
“Today, we see ourselves as a platform where if you're building a horizontal use case in your business, a great place to come. If you have a lot of domain specificity, that's where I see a lot of kind of the application companies forming over time, where they w…”
Mati Staniszewski Apr 14, 2026 ▶ 11:06
Prediction Open · timeframe Dec 2029
Staniszewski: Advanced voice AI will reach cars via cloud in 2026
“I think this year it should be in the automotive Site two, or some of the applications. John Collison: Okay, so you think we'll start seeing kind of great voice models in cars this year? Mati Staniszewski: This year for their own cloud use cases, like on, on, …”
Mati Staniszewski Apr 14, 2026 ▶ 15:18
Assertion Supported
Staniszewski: Audible Blocked AI-Generated Audiobooks
“So Audible would, like, block AI content.”
Mati Staniszewski Apr 14, 2026 ▶ 16:15
Opinion
Collison: Text LLMs passed Turing test, but voice LLMs are nowhere near
“That's the simpler way of saying what I'm saying is that we have passed the Turing test with text LLMs a long time ago, and we're actually nowhere near that on voice LLMs.”
John Collison Apr 14, 2026 ▶ 19:38
Prediction Not checkable as stated
Staniszewski: ElevenLabs hopes to pass voice Turing test within a year
“Our goal is to like pass the voice Turing test in all those cases, or the Turing test for all conversational agents outside of voice two. And I hope we will be there in the next year or so.”
Mati Staniszewski Apr 14, 2026 ▶ 20:44
Disclosure
Staniszewski: ElevenLabs to release personalized voice transcription in coming months
“No, solvable. Like, we think we can roll it out in one of the next versions, which is, like, hopefully in the next months.”
Mati Staniszewski Apr 14, 2026 ▶ 24:19
Assertion Supported
ElevenLabs expressive mode lets voice agents detect and match human emotions
“So today, finally, you can have both speed generation or entire voice agent experience With what we call expressive mode where the agent knows the emotions on the other side. So if the person is stressed, it can react and be reassuring and that's generating a …”
Mati Staniszewski Apr 14, 2026 ▶ 26:57
Disclosure
ElevenLabs bets research on cascaded voice architecture over speech-to-speech
“So today we are optimizing heavily on a cascaded approach... So that's like where we are betting a lot of the research work of how you can make that great, and we think we can make that great.”
Mati Staniszewski Apr 14, 2026 ▶ 28:33
Prediction Not checkable as stated
Staniszewski: Speech-to-speech AI will flourish in companion apps
“And speech to speech, as you think about maybe more of like a companion version of the applications, that's where that will flourish because maybe the hallucinations aren't as important, but the latency is a little bit more, and maybe hallucinations are even a…”
Mati Staniszewski Apr 14, 2026 ▶ 29:08
Assertion Not checkable as stated
Staniszewski: Speech-to-speech models are 'definitely dumber' due to size limits
“They are definitely dumber. You need smaller model, you cannot... If you are going speech to speech, usually you will use smaller models, so it's still quick.”
Mati Staniszewski Apr 14, 2026 ▶ 30:01
Assertion Not checkable as stated
Staniszewski says voice lead forms capture richer customer data than text
“One, people were actually much more keen to leave the forms through speaking with the agent, so we would go through the form a lot easier. But second, they would be a lot more open-ended in terms of what the use case are. So they would start giving us informat…”
Mati Staniszewski Apr 14, 2026 ▶ 30:51
Assertion Partly supported
Staniszewski says leading voice models use under twenty billion parameters
“A few billion to low tens of billion parameter models.”
Mati Staniszewski Apr 14, 2026 ▶ 36:34
Disclosure
Staniszewski: ElevenLabs raised $500M at an $11B valuation
“We, of course, raised recently a half a billion at eleven billion valuation to, like.”
Mati Staniszewski Apr 14, 2026 ▶ 37:01
Disclosure
Staniszewski: ElevenLabs offers new models at cost to drive early customer adoption
“The way we usually do is like when we have a new model, we try to give it at cost to a lot of the customers so they can experience the best.”
Mati Staniszewski Apr 14, 2026 ▶ 38:09
Prediction Held up
Staniszewski: Fused voice-LLM models will reach hundreds of billions of parameters
“The thing that's, you know, like I hesitated on the question is, in a cascaded approach, you probably will not see like dramatic size changes. You inherently want the models to be quick and reliable. You want to orchestrate them in a smart way. In a fused appr…”
Mati Staniszewski Apr 14, 2026 ▶ 39:39
Insight
Staniszewski: Solving conversational voice AI inherently solves text agents too
“If you fix the voice agent, you'll have fixed text piece or like inherently as well.”
Mati Staniszewski Apr 14, 2026 ▶ 43:07
Disclosure
Staniszewski: ElevenLabs will not build for deep reasoning or financial analysis
“We wouldn't, for example, go into what I think will happen in a lot of those cases, like very deeply into reasoning version of those use cases where you maybe need to like the multi-touch a lot of complex, a lot of like, Financial analysis of, like, is, like, …”
Mati Staniszewski Apr 14, 2026 ▶ 43:50
Assertion Not checkable as stated
Staniszewski: ElevenLabs added $100M net new ARR in one quarter
“This quarter was kind of one of the best for enterprise growth, where we had the first quarter hit a hundred million in an additional ARR growth, which is crazy.”
Mati Staniszewski Apr 14, 2026 ▶ 44:33
Assertion Not checkable as stated
Staniszewski: Over 50% of ElevenLabs revenue is sales-led enterprise
“We are over 50% is now sales-led enterprise.”
Mati Staniszewski Apr 14, 2026 ▶ 45:30
Assertion Not checkable as stated
Staniszewski: ElevenLabs employs 470 people
“We are now 400 470 people as a company.”
Mati Staniszewski Apr 14, 2026 ▶ 46:54
Disclosure
Staniszewski: ElevenLabs keeps product and research teams under 10 people
“We have less than 10 people teams for each of the product or research initiatives, or even as you think about sharding some of our go-to-market strategy, those will be smaller teams, understanding the industry in depth, understanding the market in depth, and g…”
Mati Staniszewski Apr 14, 2026 ▶ 47:08
Prediction Not checkable as stated
Collison: Every AI product will need subscriptions plus overage payments
“I think every AI product will need, you know, they probably want to have some all you can eat, most of what you can eat, subscription with limits, and then the ability to pay for overages.”
John Collison Apr 14, 2026 ▶ 51:27
Disclosure
Staniszewski: ElevenLabs founders and leaders manage over 15 direct reports each
“Both me and my co-founder will have over 15 direct reports each that, that we'll work with, and most of those people will have that same scale of direct reports.”
Mati Staniszewski Apr 14, 2026 ▶ 53:15
Disclosure
Staniszewski: ElevenLabs embeds technical leads into non-technical teams like talent
“Even in non-technical teams, Having a technical resource. So, you know, we will have a person in ops or in talent that will, we have effectively a tech lead for that team. That helps them automate a lot of that, that, that work and helps up level the rest of t…”
Mati Staniszewski Apr 14, 2026 ▶ 54:26
Disclosure
ElevenLabs built an interactive voice agent to prep job candidates
“We created a voice agent that people can speak with and see what's the culture, but also get prepped for the interviews.”
Mati Staniszewski Apr 14, 2026 ▶ 56:46
Disclosure
ElevenLabs partnered with Ukraine to add voice capabilities to Diia
“We traveled to Kiev. We worked with them on bringing that and making that available for voice so everybody can access it.”
Mati Staniszewski Apr 14, 2026 ▶ 57:54
Opinion
Staniszewski calls Ukraine's government tech the most advanced he has seen
“So tacked forward, like the most advanced set of work we've seen.”
Mati Staniszewski Apr 14, 2026 ▶ 58:28
Prediction Not checkable as stated
Collison: High-Agency People Will Win From AI Advances, Low-Agency Will Lose
“My biggest takeaway from all this Has been that around agency where I feel like high agency people are the winners of the advances in AI and within organizations, low agency people will lose out.”
John Collison Apr 14, 2026 ▶ 59:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, about 28 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.