May 1, 2023 · 33m · mad
Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Data Driven NYC, AssemblyAI Founder and CEO Dylan Fox discusses the architecture, business model, and practical applications behind scaling speech recognition and generative AI for audio. He explains how AssemblyAI empowers product teams with specialized speech-to-text models and developer tools, contrasting dedicated AI platforms with in-house infrastructure and open-source alternatives.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 11.3% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Dylan playfully puts Matt on the spot by citing one of his tweets about AI market maps having shades of gray, prompting Matt to joke that all speakers must read his tweets beforehand.
Hardest push from Matt ▶ 11:11 Host interjecting on model categorizationMatt interrupts to clarify whether Conformer-1 is AssemblyAI's own LLM, probing the timeline and underlying Transformer architecture.
Biggest teaching moment ▶ 11:15 Correcting ASR versus LLM distinctionDylan politely corrects Matt's statement that Conformer-1 is an LLM, clarifying that it is a dedicated automatic speech recognition model.
Matt holds his own ▶ 21:00 Host's analysis of positioning and moatsMatt articulates a well-structured hypothesis on how AssemblyAI positions itself against cloud vendors by targeting last-mile application layers and no-code product workflows.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Core Speech Recognition and AssemblyAI's Processing Scale | 2 | 4 | 0 | 1 | Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms. | |
| Tailored Task-Specific Models and Audio Intelligence Features | 4 | 4 | 0 | 1 | Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit. | |
| Model Engineering and the Conformer-1 Architecture | 6 | 5 | 2 | 3 | Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps. | |
| Data Sourcing and the Build vs Buy Advantage | 6 | 5 | 1 | 3 | Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations. | |
| Emerging AI Trends and Technological Accessibility | 3 | 2 | 0 | 0 | Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing. | |
| Audience Q&A: Differentiating from Open-Source Whisper | 1 | 6 | 1 | 0 | An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation. | |
| Audience Q&A: Long-Term Competitive Moat and Distribution | 1 | 5 | 0 | 0 | An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio. |