Dec 11, 2025 · 41m · no-priors

No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

Mati Staniszewski · 30m spoken Sarah Guo · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

ElevenLabs co-founder and CEO Mati Staniszewski discusses the company's rapid growth to a $300 million ARR run rate, detailing their frontier audio research, conversational agent architectures, and enterprise strategy. He outlines a transformative vision where real-time voice interfaces eliminate language barriers, enable screenless computing, and reshape customer engagement and education.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.2% of the talking time here. How this is scored →

The hosts as informed peer 5.8 Guest teaching 5.0 Guest disagreement 0.9 The hosts pushing back 1.1
05100:0015:0030:000:46–6:52 · The hosts as informed peer 5/10 Company Overview and Global Operating Scale Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice.6:52–10:11 · The hosts as informed peer 6/10 Sequencing Frontier Research and Product Platforms Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows.10:11–12:35 · The hosts as informed peer 5/10 Expanding Modalities and the Vision for Real-Time Dubbing Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts.12:36–17:54 · The hosts as informed peer 6/10 Voice Quality, Customization, and Benchmarking Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances.17:54–23:20 · The hosts as informed peer 5/10 Emerging Conversational Agent Platforms and Real-World Applications Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization.23:21–26:43 · The hosts as informed peer 7/10 Enterprise Go-to-Market and Competitive Architecture Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model.26:43–30:01 · The hosts as informed peer 6/10 Differentiating from Frontier Foundation Model Labs Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent.30:02–34:24 · The hosts as informed peer 7/10 Open Source Commoditization and Sustainable Moats Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems.34:25–36:52 · The hosts as informed peer 6/10 Next-Generation Model Architectures and Latency Optimization Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures.36:52–41:14 · The hosts as informed peer 5/10 The Future of AI Companions, Dictation, and AI Tutoring Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift.0:46–6:52 · Guest teaching 4/10 Company Overview and Global Operating Scale Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice.6:52–10:11 · Guest teaching 5/10 Sequencing Frontier Research and Product Platforms Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows.10:11–12:35 · Guest teaching 5/10 Expanding Modalities and the Vision for Real-Time Dubbing Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts.12:36–17:54 · Guest teaching 6/10 Voice Quality, Customization, and Benchmarking Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances.17:54–23:20 · Guest teaching 5/10 Emerging Conversational Agent Platforms and Real-World Applications Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization.23:21–26:43 · Guest teaching 5/10 Enterprise Go-to-Market and Competitive Architecture Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model.26:43–30:01 · Guest teaching 5/10 Differentiating from Frontier Foundation Model Labs Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent.30:02–34:24 · Guest teaching 5/10 Open Source Commoditization and Sustainable Moats Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems.34:25–36:52 · Guest teaching 6/10 Next-Generation Model Architectures and Latency Optimization Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures.36:52–41:14 · Guest teaching 4/10 The Future of AI Companions, Dictation, and AI Tutoring Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift.0:46–6:52 · Guest disagreement 1/10 Company Overview and Global Operating Scale Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice.6:52–10:11 · Guest disagreement 1/10 Sequencing Frontier Research and Product Platforms Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows.10:11–12:35 · Guest disagreement 1/10 Expanding Modalities and the Vision for Real-Time Dubbing Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts.12:36–17:54 · Guest disagreement 1/10 Voice Quality, Customization, and Benchmarking Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances.17:54–23:20 · Guest disagreement 0/10 Emerging Conversational Agent Platforms and Real-World Applications Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization.23:21–26:43 · Guest disagreement 1/10 Enterprise Go-to-Market and Competitive Architecture Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model.26:43–30:01 · Guest disagreement 2/10 Differentiating from Frontier Foundation Model Labs Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent.30:02–34:24 · Guest disagreement 1/10 Open Source Commoditization and Sustainable Moats Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems.34:25–36:52 · Guest disagreement 0/10 Next-Generation Model Architectures and Latency Optimization Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures.36:52–41:14 · Guest disagreement 1/10 The Future of AI Companions, Dictation, and AI Tutoring Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift.0:46–6:52 · The hosts pushing back 1/10 Company Overview and Global Operating Scale Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice.6:52–10:11 · The hosts pushing back 2/10 Sequencing Frontier Research and Product Platforms Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows.10:11–12:35 · The hosts pushing back 1/10 Expanding Modalities and the Vision for Real-Time Dubbing Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts.12:36–17:54 · The hosts pushing back 1/10 Voice Quality, Customization, and Benchmarking Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances.17:54–23:20 · The hosts pushing back 0/10 Emerging Conversational Agent Platforms and Real-World Applications Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization.23:21–26:43 · The hosts pushing back 2/10 Enterprise Go-to-Market and Competitive Architecture Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model.26:43–30:01 · The hosts pushing back 2/10 Differentiating from Frontier Foundation Model Labs Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent.30:02–34:24 · The hosts pushing back 1/10 Open Source Commoditization and Sustainable Moats Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems.34:25–36:52 · The hosts pushing back 0/10 Next-Generation Model Architectures and Latency Optimization Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures.36:52–41:14 · The hosts pushing back 1/10 The Future of AI Companions, Dictation, and AI Tutoring Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.7% · guest 64.3%0:00 · the hosts 35.7% · guest 64.3%3:00 · the hosts 20.3% · guest 79.7%3:00 · the hosts 20.3% · guest 79.7%6:00 · the hosts 24.3% · guest 75.7%6:00 · the hosts 24.3% · guest 75.7%9:00 · the hosts 10.9% · guest 89.1%9:00 · the hosts 10.9% · guest 89.1%12:00 · the hosts 23.2% · guest 76.8%12:00 · the hosts 23.2% · guest 76.8%15:00 · the hosts 15.2% · guest 84.8%15:00 · the hosts 15.2% · guest 84.8%18:00 · the hosts 7.6% · guest 92.4%18:00 · the hosts 7.6% · guest 92.4%21:00 · the hosts 28.2% · guest 71.8%21:00 · the hosts 28.2% · guest 71.8%24:00 · the hosts 22.9% · guest 77.1%24:00 · the hosts 22.9% · guest 77.1%27:00 · the hosts 15.3% · guest 84.7%27:00 · the hosts 15.3% · guest 84.7%30:00 · the hosts 24.5% · guest 75.5%30:00 · the hosts 24.5% · guest 75.5%33:00 · the hosts 7.9% · guest 92.1%33:00 · the hosts 7.9% · guest 92.1%36:00 · the hosts 21.6% · guest 78.4%36:00 · the hosts 21.6% · guest 78.4%39:00 · the hosts 27.9% · guest 72.1%39:00 · the hosts 27.9% · guest 72.1%
Sharpest disagreement ▶ 28:05 Rejecting brute-force scale assumptions in audio

Mati firmly rejects the premise that frontier foundation model labs will naturally dominate voice, arguing that audio performance relies on architectural innovation rather than compute scaling.

Hardest push from the hosts ▶ 23:25 Challenging vendor selection and platform overlaps

Sarah presses Mati on how enterprise buyers distinguish ElevenLabs from established consultancies like Palantir and focused agent vendors like Sierra.

Biggest teaching moment ▶ 17:05 Explaining the lack of qualitative audio benchmarks

Mati explains the fundamental gap in audio evaluation, demonstrating why standard text transcription labeling fails to capture prosody, emotion, and accent nuances.

The host holds their own ▶ 31:53 Dissecting transient tech moats and product velocity

Sarah articulates the investor reality that technical model advantages are fleeting windows rather than permanent moats, prompting Mati to validate her framework.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Company Overview and Global Operating Scale 5411 Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice.
Sequencing Frontier Research and Product Platforms 6512 Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows.
Expanding Modalities and the Vision for Real-Time Dubbing 5511 Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts.
Voice Quality, Customization, and Benchmarking 6611 Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances.
Emerging Conversational Agent Platforms and Real-World Applications 5500 Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization.
Enterprise Go-to-Market and Competitive Architecture 7512 Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model.
Differentiating from Frontier Foundation Model Labs 6522 Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent.
Open Source Commoditization and Sustainable Moats 7511 Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems.
Next-Generation Model Architectures and Latency Optimization 6600 Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures.
The Future of AI Companions, Dictation, and AI Tutoring 5411 Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift.

Statements from this episode (20)

Assertion Supported
Staniszewski: ElevenLabs has grown to 350 employees globally
“So we've grown to all, 350 people globally.”
Mati Staniszewski Dec 11, 2025 ▶ 1:58
Assertion Supported
ElevenLabs reached $300M ARR, split evenly between self-serve and enterprise
“We are at three hundred million in, in, in ARR, which is roughly fifty-fifty between SelfServe, so a lot of subscription and creators using our creative platform, and then Approaching 50% on the enterprise side using our agents platform work, and that's on the…”
Mati Staniszewski Dec 11, 2025 ▶ 2:13
Assertion Not checkable as stated
Staniszewski: ElevenLabs has over 5 million monthly active creative users
“And we serve more than five million monthly actives on that creative side of the work.”
Mati Staniszewski Dec 11, 2025 ▶ 2:34
Insight
Staniszewski: Voice agent success requires system integrations, not just orchestration
“There's the product layer that starts forming that it's not only the orchestration that matters. It's also the integrations of how you link up to the legacy systems, how you build functions around it, or how you deploy that in production and test, monitor, eva…”
Mati Staniszewski Dec 11, 2025 ▶ 9:57
Prediction Not checkable as stated
Staniszewski: Real-time audio dubbing will break global language barriers
“And I still think actually this market will be immense because it's not going to be only the static delivery in movies, but if you travel around the world and want to communicate in real time, like the full bubble fish idea from Hedgehacker's Guide to Galaxy, …”
Mati Staniszewski Dec 11, 2025 ▶ 11:34
Disclosure
Staniszewski: ElevenLabs employs 'voice sommelier' team for enterprise branding
“We have, like, a voice sommelier, effectively, with us. We work with enterprises. We deploy that person to Work with them and help them navigate. That person is like a voice coach, has an incredible voice themselves, and now we have, like, a team under that pe…”
Mati Staniszewski Dec 11, 2025 ▶ 13:36
Insight
Voice AI benchmarking remains unsolved due to subjective user preferences
“I think it's an unsolved problem still, where I think you have good benchmarks, of course, in LMS, I think in image space, they are pretty good. In voice space, you have, of course, the speech quality, but then so much of whether you like or not the speech dep…”
Mati Staniszewski Dec 11, 2025 ▶ 16:10
Disclosure
Legacy annotation firms failed at audio emotion labeling, forcing in-house creation
“The understanding of how you describe audio data is still lagging in the industry. Like, when we initially started, we, of course, went into the traditional players for them to help us label not only what was said, so, like, transcription, but also how it was …”
Mati Staniszewski Dec 11, 2025 ▶ 17:15
Assertion Supported
Staniszewski: ElevenLabs powered interactive Darth Vader voice in Fortnite
“We worked with them on bringing the voice of Darth Vader and Darth Vader into Fortnite, where millions of players could interact with Darth Vader live in the game, where you had like a full experience of Darth Vader in a, in, in, in a new way.”
Mati Staniszewski Dec 11, 2025 ▶ 20:14
Prediction Not checkable as stated
Staniszewski: Headphone-based personal tutors will be voice AI's biggest shift
“And then I think the one that's, that I'm most excited about for the world and for the shift is going to be education, where you will just be able to have, like, effectively a personal tutor on your headphone.”
Mati Staniszewski Dec 11, 2025 ▶ 20:38
Assertion Supported
Ukraine's government partnered with ElevenLabs to build agentic public services
“Recently I went to Ukraine where we are working with Ministry of Transformation, where they are effectively creating a first agent in government, and the crazy thing is, they have all of those.”
Mati Staniszewski Dec 11, 2025 ▶ 21:53
Insight
Audio AI progress requires architectural breakthroughs rather than raw compute scale
“The main part that I think is different in audio space is that you don't need the scale as much as you need the, Architectural breakthroughs, the model breakthroughs to really make a dent.”
Mati Staniszewski Dec 11, 2025 ▶ 28:52
Opinion
ElevenLabs employs roughly 10 of the world's top 100 audio researchers
“We think there's maybe 50 to a hundred researchers in audio space that could do it. We think we have probably 10 of them in the company that that are some of the best ones.”
Mati Staniszewski Dec 11, 2025 ▶ 29:13
Prediction Open · timeframe Dec 2027
Voice conversational AI will pass the Turing test within a year
“I think this is, is good for customer support or customer experience, but it's still a way away from like conversation like we have and like passing that Turing test. So I think this is still like a, at least a year, like within a year. And then you will have …”
Mati Staniszewski Dec 11, 2025 ▶ 31:29
Insight
Staniszewski: AI research is only a head start; ecosystems create lasting value
“The thing that will really give that long-term value is the ecosystem that you create around, whether that's the run and distribution, whether that's the collection of voices you can have, the collection of integrations you can build, the workflows that you ca…”
Mati Staniszewski Dec 11, 2025 ▶ 33:01
Disclosure
Staniszewski: ElevenLabs builds product workarounds if research takes over 3 months
“And rough rule of thumb is, like, three months. If we think it's going to be longer than three months, we would probably build it. If it's less than that, we probably won't.”
Mati Staniszewski Dec 11, 2025 ▶ 34:18
Assertion Supported
Staniszewski: ElevenLabs' Scribe v2 hits 93.5% accuracy under 150ms latency
“We just released our speech-to-text model, Scribe v.II, which is under a 150 milliseconds, 93.5% accuracy across the top 30 languages on, on Flourers. And it's only top 30 here because we serve so many others, but most of the people don't. So so it's so it's b…”
Mati Staniszewski Dec 11, 2025 ▶ 35:20
Disclosure
Staniszewski: ElevenLabs releasing low-latency emotional orchestration within months
“We are releasing, we'll be releasing over the next couple of months a new orchestration mechanism that will lower the end-to-end part we think in a great way. But second thing, which is what is so hard, is it's not going to only allow you to combine those piec…”
Mati Staniszewski Dec 11, 2025 ▶ 35:44
Prediction Not checkable as stated
Staniszewski: Cascaded voice architectures will dominate enterprise for 1–2 years
“And of course, depending on the use case, if you are, like, enterprise-reliable use case, the cascaded approach is the approach for the next year to...”
Mati Staniszewski Dec 11, 2025 ▶ 36:15
Prediction Not checkable as stated
Staniszewski: Future education will split between AI tutors and tech-free human learning
“You will have education, good percent of time spent with AI tutors, but then explicit percent of time spent without any technology, human to human, so you can kind of learn that part, too.”
Mati Staniszewski Dec 11, 2025 ▶ 38:14
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.