Apr 14, 2026 · 1h 0m · cheeky-pint
The world of voice AI, with Mati Staniszewski of ElevenLabs
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this in-depth interview, Stripe co-founder John Collison speaks with ElevenLabs co-founder and CEO Mati Staniszewski about the architecture, rapid scaling, and real-world impact of frontier voice AI. Staniszewski explains how proprietary data annotation, synchronized audio modeling, and an agile, AI-native organizational structure propelled ElevenLabs to over $450M in ARR.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. John holds 31.5% of the talking time here. How this is scored →
speaking balance: gold is John, purple is the guest (3 minute bins)
Staniszewski politely but directly pushes back against Collison's framing that voice AI UX is stuck ten years behind modern LLMs.
Hardest push from John ▶ 53:45 Pushing back on extreme spans of controlCollison playfully challenges whether having 15+ direct reports per executive is genuine AI leverage or just early-stage founder management theory.
Biggest teaching moment ▶ 2:15 Explaining Mel Spectrogram and audio modelingStaniszewski breaks down the core pipeline of text, Mel Spectrogram, and waveform synthesis, clarifying concepts Collison asked him to unpack.
John holds their own ▶ 44:46 Collison triangulates revenue trajectoryCollison uses quick mental math to convert Staniszewski's net new quarterly metrics into a sharp $450M+ run rate figure.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | John as informed peer | Guest teaching | Guest disagreement | John pushing back | Why |
|---|---|---|---|---|---|---|
| How Audio AI Models Work Under the Hood | 5 | 6 | 1 | 2 | Collison asks technical questions framing voice generation against LLMs. Staniszewski educates him on phonemes, spectrogram spaces, and emergent voice qualities. | |
| Audio Tokens, Phonemes, and Real-Time Synchronization | 4 | 5 | 1 | 1 | Collison inquires about audio token equivalents and human-sounding inflection. Staniszewski breaks down phoneme deconstruction, proprietary data labeling, and internal model development. | |
| Historical Parallels: Wolfgang von Kempelen's Mechanical Turk | 5 | 3 | 1 | 1 | Staniszewski shares historical context about Wolfgang von Kempelen's speaking machine and Mechanical Turk, which Collison quickly recognizes and connects to platform breakdown. | |
| Platform Strategy and Horizontal Infrastructure vs. Vertical Apps | 7 | 4 | 3 | 5 | Collison challenges why voice AI feels ten years behind consumer LLMs and probes intermediation risks. Staniszewski pushes back on the 10-year premise by citing deployment lag and rapid model advances. | |
| Eleven Reader and Consumer Voice Applications | 5 | 4 | 1 | 2 | Collison asks about consumer PDF reading apps and third-party mobile transcription. Staniszewski highlights Eleven Reader's creation in response to distributor blocks. | |
| The Conversational Voice Turing Test and Turn-Taking | 6 | 4 | 2 | 3 | Collison notes conversational voice still fails the Turing test compared to text. Staniszewski details the complex orchestration of turn-taking, latency, and tool-calling. | |
| Stripe Link Sponsor Segment | 6 | 4 | 1 | 2 | Collison reads a Stripe Link sponsor segment before discussing speaker-specific fine-tuning. Staniszewski explains keyword detection and diarization capabilities. | |
| Controllability, Expressive Modes, and Speech Generation | 5 | 5 | 1 | 1 | Collison pitches de-accenting and voice filters. Staniszewski describes ElevenLabs v3's expressive mode and emotional controllability cues. | |
| Cascaded Architectures vs. End-to-End Speech-to-Speech | 7 | 5 | 1 | 2 | Collison draws analogies to written language altering human cognition and asks if direct speech-to-speech models reason differently. Staniszewski clarifies latency trade-offs and model size constraints. | |
| Behavioral Dynamics: Voice Interfaces vs. Static Web Forms | 5 | 5 | 1 | 1 | Staniszewski explains how inbound voice agents elicit higher disclosure and ease compared to text forms. Collison compares the dynamic to open-ended adventure games. | |
| Proactive Voice Agents and the Guinness Index | 7 | 4 | 1 | 2 | Collison probes compute economics, training run costs, and model parameter counts. Staniszewski reveals voice models operate in the low billions of parameters and explains subsidizing new models. | |
| Enterprise Conversational Agents in Sales and Support | 6 | 4 | 1 | 2 | Collison compares conversational agents to Intercom's Finn and asks where voice fits in full-stack support and reasoning workflows. Staniszewski defines their boundary against deep reasoning. | |
| Scaling to $450M+ ARR and Small Autonomous Teams | 6 | 4 | 1 | 1 | Collison calculates ElevenLabs' run rate at $450M+ ARR after Staniszewski reveals $100M net new ARR in a single quarter, unpacking their small autonomous team structure. | |
| Product-Led Growth, Self-Serve Philosophy, and Pay-As-You-Go | 7 | 3 | 1 | 2 | Collison and Staniszewski align on the necessity of self-serve and pay-as-you-go billing, joking about ElevenLabs' feedback leading to Stripe's Metronome acquisition. | |
| Organizational Design and Flatter Teams in the AI Era | 6 | 4 | 2 | 3 | Collison questions whether having 15+ direct reports per leader is a startup gimmick or an AI-enabled reality. Staniszewski defends flat teams paired with embedded tech leads. | |
| AI-Native Internal Tooling and Ukraine's Diia App | 4 | 6 | 1 | 1 | Staniszewski shares how ElevenLabs worked with Ukraine's government on the Diia app, noting how each ministry embeds dedicated technical resources for AI agents. | |
| Cultivating High Agency and Conclusion | 6 | 3 | 0 | 0 | Collison and Staniszewski conclude by reflecting on high agency as the defining predictor of success and cultural scalability in the AI era. |