Dec 11, 2025 · 41m · no-priors
No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
ElevenLabs co-founder and CEO Mati Staniszewski discusses the company's rapid growth to a $300 million ARR run rate, detailing their frontier audio research, conversational agent architectures, and enterprise strategy. He outlines a transformative vision where real-time voice interfaces eliminate language barriers, enable screenless computing, and reshape customer engagement and education.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Mati firmly rejects the premise that frontier foundation model labs will naturally dominate voice, arguing that audio performance relies on architectural innovation rather than compute scaling.
Hardest push from the hosts ▶ 23:25 Challenging vendor selection and platform overlapsSarah presses Mati on how enterprise buyers distinguish ElevenLabs from established consultancies like Palantir and focused agent vendors like Sierra.
Biggest teaching moment ▶ 17:05 Explaining the lack of qualitative audio benchmarksMati explains the fundamental gap in audio evaluation, demonstrating why standard text transcription labeling fails to capture prosody, emotion, and accent nuances.
The host holds their own ▶ 31:53 Dissecting transient tech moats and product velocitySarah articulates the investor reality that technical model advantages are fleeting windows rather than permanent moats, prompting Mati to validate her framework.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Company Overview and Global Operating Scale | 5 | 4 | 1 | 1 | Sarah frames ElevenLabs within the generative creation cohort (comparing them to Midjourney and Suno) and questions early market sizing like dubbing. Mati explains the Polish voiceover backstory and details the company's evolution from static dubbing to interactive conversational voice. | |
| Sequencing Frontier Research and Product Platforms | 6 | 5 | 1 | 2 | Sarah challenges the conventional startup wisdom of focusing on one single product versus doing frontier research and multi-platform applications simultaneously. Mati details their internal lab structure, explaining how research on foundational voice models was paired immediately with applied product workflows. | |
| Expanding Modalities and the Vision for Real-Time Dubbing | 5 | 5 | 1 | 1 | Sarah queries whether ElevenLabs created new market demand or followed existing user requests. Mati outlines expansion into music and real-time translation, referencing universal translation concepts. | |
| Voice Quality, Customization, and Benchmarking | 6 | 6 | 1 | 1 | Sarah asks how non-technical enterprise buyers navigate voice selection given the lack of standardized audio evals. Mati outlines their specialized voice sommelier approach and details the difficulty of qualifying and labeling emotional voice nuances. | |
| Emerging Conversational Agent Platforms and Real-World Applications | 5 | 5 | 0 | 0 | Sarah asks about traction areas for conversational agents beyond basic support. Mati shares concrete deployments spanning proactive retail assistance, interactive gaming IP, and Ukrainian government digitization. | |
| Enterprise Go-to-Market and Competitive Architecture | 7 | 5 | 1 | 2 | Sarah probes how enterprise buyers should differentiate ElevenLabs from system integrators like Palantir or vertical agent providers like Sierra. Mati breaks down their modular platform architecture and forward-deployed engineering engagement model. | |
| Differentiating from Frontier Foundation Model Labs | 6 | 5 | 2 | 2 | Sarah brings up the persistent investor objection regarding frontier labs like OpenAI or Google eventually subsuming voice. Mati counters by emphasizing that audio requires architectural breakthroughs rather than brute compute scale, noting ElevenLabs holds key specialized talent. | |
| Open Source Commoditization and Sustainable Moats | 7 | 5 | 1 | 1 | Sarah and Mati discuss defensibility, agreeing that model advantages last only 6 to 12 months before commoditizing into open source. Mati argues that true sustainable value comes from surrounding product workflows and distribution ecosystems. | |
| Next-Generation Model Architectures and Latency Optimization | 6 | 6 | 0 | 0 | Sarah asks about active research initiatives and latency optimization. Mati breaks down the technical tradeoffs between cascaded enterprise-reliable pipelines and end-to-end speech-to-speech architectures. | |
| The Future of AI Companions, Dictation, and AI Tutoring | 5 | 4 | 1 | 1 | Sarah and Mati explore future paradigms including AI companions, voice dictation, and robotics. Mati highlights personalized 1-on-1 tutoring and interactive education as the most impactful upcoming shift. |