May 1, 2023 · 33m · mad

Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox

Dylan Fox · 25m spoken Matt Turck · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Data Driven NYC, AssemblyAI Founder and CEO Dylan Fox discusses the architecture, business model, and practical applications behind scaling speech recognition and generative AI for audio. He explains how AssemblyAI empowers product teams with specialized speech-to-text models and developer tools, contrasting dedicated AI platforms with in-house infrastructure and open-source alternatives.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 11.3% of the talking time here. How this is scored →

Matt as informed peer 3.3 Guest teaching 4.4 Guest disagreement 0.6 Matt pushing back 1.1
05100:0010:0020:0030:000:09–4:12 · Matt as informed peer 2/10 Core Speech Recognition and AssemblyAI's Processing Scale Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms.4:12–8:45 · Matt as informed peer 4/10 Tailored Task-Specific Models and Audio Intelligence Features Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit.8:45–16:03 · Matt as informed peer 6/10 Model Engineering and the Conformer-1 Architecture Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps.16:03–23:36 · Matt as informed peer 6/10 Data Sourcing and the Build vs Buy Advantage Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations.23:36–25:37 · Matt as informed peer 3/10 Emerging AI Trends and Technological Accessibility Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing.25:37–28:22 · Matt as informed peer 1/10 Audience Q&A: Differentiating from Open-Source Whisper An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation.28:22–31:23 · Matt as informed peer 1/10 Audience Q&A: Long-Term Competitive Moat and Distribution An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio.0:09–4:12 · Guest teaching 4/10 Core Speech Recognition and AssemblyAI's Processing Scale Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms.4:12–8:45 · Guest teaching 4/10 Tailored Task-Specific Models and Audio Intelligence Features Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit.8:45–16:03 · Guest teaching 5/10 Model Engineering and the Conformer-1 Architecture Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps.16:03–23:36 · Guest teaching 5/10 Data Sourcing and the Build vs Buy Advantage Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations.23:36–25:37 · Guest teaching 2/10 Emerging AI Trends and Technological Accessibility Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing.25:37–28:22 · Guest teaching 6/10 Audience Q&A: Differentiating from Open-Source Whisper An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation.28:22–31:23 · Guest teaching 5/10 Audience Q&A: Long-Term Competitive Moat and Distribution An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio.0:09–4:12 · Guest disagreement 0/10 Core Speech Recognition and AssemblyAI's Processing Scale Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms.4:12–8:45 · Guest disagreement 0/10 Tailored Task-Specific Models and Audio Intelligence Features Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit.8:45–16:03 · Guest disagreement 2/10 Model Engineering and the Conformer-1 Architecture Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps.16:03–23:36 · Guest disagreement 1/10 Data Sourcing and the Build vs Buy Advantage Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations.23:36–25:37 · Guest disagreement 0/10 Emerging AI Trends and Technological Accessibility Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing.25:37–28:22 · Guest disagreement 1/10 Audience Q&A: Differentiating from Open-Source Whisper An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation.28:22–31:23 · Guest disagreement 0/10 Audience Q&A: Long-Term Competitive Moat and Distribution An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio.0:09–4:12 · Matt pushing back 1/10 Core Speech Recognition and AssemblyAI's Processing Scale Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms.4:12–8:45 · Matt pushing back 1/10 Tailored Task-Specific Models and Audio Intelligence Features Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit.8:45–16:03 · Matt pushing back 3/10 Model Engineering and the Conformer-1 Architecture Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps.16:03–23:36 · Matt pushing back 3/10 Data Sourcing and the Build vs Buy Advantage Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations.23:36–25:37 · Matt pushing back 0/10 Emerging AI Trends and Technological Accessibility Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing.25:37–28:22 · Matt pushing back 0/10 Audience Q&A: Differentiating from Open-Source Whisper An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation.28:22–31:23 · Matt pushing back 0/10 Audience Q&A: Long-Term Competitive Moat and Distribution An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 25.9% · guest 74.1%0:00 · Matt 25.9% · guest 74.1%3:00 · Matt 8.5% · guest 91.5%3:00 · Matt 8.5% · guest 91.5%6:00 · Matt 8.3% · guest 91.7%6:00 · Matt 8.3% · guest 91.7%9:00 · Matt 18.7% · guest 81.3%9:00 · Matt 18.7% · guest 81.3%12:00 · Matt 6.1% · guest 93.9%12:00 · Matt 6.1% · guest 93.9%15:00 · Matt 13.5% · guest 86.5%15:00 · Matt 13.5% · guest 86.5%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 34.9% · guest 65.1%21:00 · Matt 34.9% · guest 65.1%24:00 · Matt 5.3% · guest 94.7%24:00 · Matt 5.3% · guest 94.7%27:00 · Matt 0.8% · guest 99.2%27:00 · Matt 0.8% · guest 99.2%30:00 · Matt 3.9% · guest 96.1%30:00 · Matt 3.9% · guest 96.1%33:00 · Matt 1.4% · guest 98.6%33:00 · Matt 1.4% · guest 98.6%
Sharpest disagreement ▶ 9:06 Playful callout of host's tweet

Dylan playfully puts Matt on the spot by citing one of his tweets about AI market maps having shades of gray, prompting Matt to joke that all speakers must read his tweets beforehand.

Hardest push from Matt ▶ 11:11 Host interjecting on model categorization

Matt interrupts to clarify whether Conformer-1 is AssemblyAI's own LLM, probing the timeline and underlying Transformer architecture.

Biggest teaching moment ▶ 11:15 Correcting ASR versus LLM distinction

Dylan politely corrects Matt's statement that Conformer-1 is an LLM, clarifying that it is a dedicated automatic speech recognition model.

Matt holds his own ▶ 21:00 Host's analysis of positioning and moats

Matt articulates a well-structured hypothesis on how AssemblyAI positions itself against cloud vendors by targeting last-mile application layers and no-code product workflows.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Core Speech Recognition and AssemblyAI's Processing Scale 2401 Matt opens with standard introductory questions regarding speech-to-text fundamentals. Dylan outlines AssemblyAI's scale processing over 100 million files per month across virtual meetings, contact centers, and video platforms.
Tailored Task-Specific Models and Audio Intelligence Features 4401 Matt shows preparation by listing specific product capabilities like sentiment analysis, summarization, and content moderation. Dylan details how product teams use specialized, lightweight task models to establish product-market fit.
Model Engineering and the Conformer-1 Architecture 6523 Matt presses on model architectures and paper timelines, briefly mistaking Conformer-1 for an LLM before Dylan gently corrects him that it is an ASR model. Dylan also lightheartedly calls out Matt's previous tweet on AI market maps.
Data Sourcing and the Build vs Buy Advantage 6513 Matt probes into training data sourcing and offers an analytical breakdown of product personas versus developer sales motions. Dylan refines the perspective, noting that VP of Product buyers drive contract signatures while developers handle long-tail integrations.
Emerging AI Trends and Technological Accessibility 3200 Matt invites general thoughts on consumer AI breakthroughs, and Dylan highlights consumer novelties like the AI Drake song and accessible coding tools like ChatGPT. Matt compliments AssemblyAI's content marketing.
Audience Q&A: Differentiating from Open-Source Whisper 1610 An audience member asks how AssemblyAI competes with free open-source models like OpenAI's Whisper. Dylan explains why raw open-source models require production infrastructure, GPU optimization, and hallucination mitigation.
Audience Q&A: Long-Term Competitive Moat and Distribution 1500 An audience question asks about long-term competitive moats across models and APIs over a 5-year horizon. Dylan outlines distribution power and developer mindshare, comparing AssemblyAI's aspirational brand positioning to Twilio.

Statements from this episode (13)

Assertion Not checkable as stated
AssemblyAI has processed almost two billion audio files to date
“We've processed almost two billion audio files through our system.”
Dylan Fox May 1, 2023 ▶ 0:47
Disclosure
Fox: AssemblyAI serves over 1,000 customers and tens of thousands of monthly developers
“We've got over a thousand customers tens of thousands a month of developers that are building with the API.”
Dylan Fox May 1, 2023 ▶ 2:10
Insight
Most speech AI value comes from downstream workflows, not just transcription
“Where a lot of value is created is you're taking the transcription, and then you're using it as an input to do something else.”
Dylan Fox May 1, 2023 ▶ 3:15
Disclosure
Fox: AssemblyAI delays exploratory model investment until developers prove value
“Some of these other things that are more exploratory, like, we're not gonna put a ton of effort into those until we see that our customers the developers that use our API are actually able to find value and create value with those.”
Dylan Fox May 1, 2023 ▶ 7:42
Assertion Supported
AssemblyAI trained Conformer-1 on 650,000 hours of labeled audio data
“So we trained it on, like, 60 terabytes of audio data, like, labeled audio data. So it was, I think, something like 650,000 hours of audio data.”
Dylan Fox May 1, 2023 ▶ 13:59
Assertion Not checkable as stated
Most commercial speech models historically trained on roughly 50,000 hours of audio
“And our models prior, and most commercial speech recognition models trained on like, 50,000 hours”
Dylan Fox May 1, 2023 ▶ 14:09
Disclosure
AssemblyAI is training its next speech model on ~4 million hours of audio
“We're actually training conformer two or what might call it 1.5, but whatever this accessory will be is training right now. And that's something around four million hours of labeled audio data.”
Dylan Fox May 1, 2023 ▶ 14:19
Assertion Not checkable as stated
AssemblyAI processes over 100 million audio files monthly via API
“We've processed, ah, yeah, it's like over a hundred million audio files a month that are flowing through the API, and that's growing pretty quickly.”
Dylan Fox May 1, 2023 ▶ 16:30
Assertion Supported
State-of-the-art speech recognition models still carry a 15% error rate
“State-of-the-art automatic speech recognition still has, like, a 15% error rate on a lot of data sets”
Dylan Fox May 1, 2023 ▶ 20:05
Opinion
Major cloud providers are too big to ship good developer products
“I'm surprised the big cloud companies, and apologies if anyone here works there, I'm surprised they can't ship, you know, better developer products, but it's, I think they're maybe just too big at this point.”
Dylan Fox May 1, 2023 ▶ 23:20
Assertion Contradicted
AssemblyAI's JAX contribution sped up Whisper model training by 10x
“We actually, I think, published like, made a contribution to Jax to make it, like, 10 times faster to train Whisper.”
Dylan Fox May 1, 2023 ▶ 26:18
Disclosure
AssemblyAI employs about 40 full-time staff dedicated to improving speech models
“We got like 40 people full time working on this, you know, and you're gonna get all the benefits of that.”
Dylan Fox May 1, 2023 ▶ 27:44
Opinion
Large organizations remain confused about who should manage internal AI projects
“I think right now, larger organizations sometimes are, like, still confused, like, who's gonna manage this AI project, you know?”
Dylan Fox May 1, 2023 ▶ 33:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.