Jul 31, 2026 · 46m · sourcery

AssemblyAI Now Handles 4x YouTube's Daily Volume

Dylan Fox · 34m spoken Molly O'Shea · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Dylan Fox, founder and CEO of AssemblyAI, joins host Molly O'Shea on Sourcery to discuss how his company scaled to process four times YouTube's daily audio volume. Fox details breakthroughs in context-aware speech models, AssemblyAI's pure-play developer infrastructure approach, and the future transition toward ambient voice interfaces across hardware and robotics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Molly holds 18.2% of the talking time here. How this is scored →

Molly as informed peer 3.2 Guest teaching 2.9 Guest disagreement 0.5 Molly pushing back 0.3
05100:0015:0030:0045:001:06–3:33 · Molly as informed peer 4/10 AssemblyAI's Unprecedented Scale and Rapid Volume Inflection Molly opens by explaining her own use case of voice transcription for VC investment memos before prompting Dylan to quantify AssemblyAI's scale. Dylan explains the platform's massive growth metrics, including handling over four times YouTube's daily volume, while Molly reacts receptively.3:33–6:39 · Molly as informed peer 3/10 Y Combinator Origins and the Autonomous Vehicle Parallel Molly asks whether Dylan anticipated current voice adoption during AssemblyAI's early YC days. Dylan explains the YC AI batch origins and articulates his autonomous vehicle thesis, noting that market adoption unlocks progressively as core model thresholds are met.6:39–10:11 · Molly as informed peer 2/10 Three Macro Drivers Accelerating Voice AI Adoption Molly asks an open prompt about macro trends driving voice inflection. Dylan systematically educates the audience on three core drivers: model capability, adjacent AI tooling, and coding agents expanding the developer TAM.10:11–13:50 · Molly as informed peer 4/10 Pure-Play Voice Infrastructure and Scaling Engineering Molly inquires about competitive positioning against full-stack voice providers like ElevenLabs and Sierra. Dylan clarifies AssemblyAI's dedicated pure-play infrastructure focus across inference and orchestration layers rather than consumer application interfaces.13:50–17:05 · Molly as informed peer 3/10 Sponsor Spotlight: Brex Agentic Finance Platform Following an ad read for Brex, Molly asks how Dylan internally restructured AssemblyAI with AI tooling amid the SaaS shift. Dylan describes internal workflow automations, including personal executive assistants and autonomous web deployment.17:05–22:33 · Molly as informed peer 4/10 Sovereign AI and Enterprise Privacy Deployments Molly introduces an investor perspective suggesting mice and keyboards will soon be entirely obsolete. Dylan gently counters this maximalist view, pointing out that touchscreens did not replace keyboards and predicting voice will serve as an additive dimension rather than a wholesale replacement.22:33–25:14 · Molly as informed peer 4/10 Humanoid Robotics and Acoustic Disambiguation Challenges Molly inquires about humanoid robotics applications and suggests models distinguish speech via listening data rather than visual feeds. Dylan elaborates that even with visual sensors, acoustic disambiguation of overlapping speakers remains a core unsolved technical bottleneck.25:14–28:02 · Molly as informed peer 3/10 Navigating Cultural Nuance in Multilingual Voice AI Molly asks Dylan about the complexities of multilingual translation and whether AssemblyAI would acquire local specialized vendors. Dylan explains that localized performance depends heavily on cultural policy alignment and native nuance rather than pure underlying architecture.28:02–31:28 · Molly as informed peer 2/10 Sponsor Spotlight: MongoDB for AI Applications After ad reads for MongoDB and AssemblyAI, Molly playfully asks whether technology will enable animal translation, citing sci-fi. Dylan humorously debunks the voice premise, noting mind-reading interfaces or sub-vocal scans would be needed over acoustic translation.31:28–34:04 · Molly as informed peer 3/10 Training Regimes, Data Alignment, and Real-World Evals Molly asks how models are trained across millions of audio hours. Dylan details the reality of speech engineering, explaining that public benchmarks are easily gamed and real performance requires extensive custom evaluation suites and domain-specific background noise handling.34:04–36:48 · Molly as informed peer 3/10 AssemblyAI Team Structure, Founder Roots, and Infrastructure DNA Molly prompts Dylan to explore AssemblyAI's internal team breakdown and personal technical roots. Dylan recounts learning to code in college, participating in IRC communities, and defining the company's core identity around developer infrastructure.36:48–42:23 · Molly as informed peer 4/10 Future Outlook: On-Device Hardware and the Voice Agent UX Dilemma Dylan outlines the UX dilemma of voice agents attempting to deceptively mimic humans, while Molly counters that users will simply assume everything is AI. Dylan pushes back with a healthcare counterexample, illustrating why deceptive AI interaction induces user discomfort in sensitive settings.42:23–43:03 · Molly as informed peer 3/10 Final Technical Reflections: Context-Aware Voice Models Molly asks if any topics were missed, and Dylan highlights new context-aware models. Molly expresses surprise that voice systems have lacked environmental context until now, and Dylan confirms that AssemblyAI only recently engineered this breakthrough.1:06–3:33 · Guest teaching 2/10 AssemblyAI's Unprecedented Scale and Rapid Volume Inflection Molly opens by explaining her own use case of voice transcription for VC investment memos before prompting Dylan to quantify AssemblyAI's scale. Dylan explains the platform's massive growth metrics, including handling over four times YouTube's daily volume, while Molly reacts receptively.3:33–6:39 · Guest teaching 3/10 Y Combinator Origins and the Autonomous Vehicle Parallel Molly asks whether Dylan anticipated current voice adoption during AssemblyAI's early YC days. Dylan explains the YC AI batch origins and articulates his autonomous vehicle thesis, noting that market adoption unlocks progressively as core model thresholds are met.6:39–10:11 · Guest teaching 4/10 Three Macro Drivers Accelerating Voice AI Adoption Molly asks an open prompt about macro trends driving voice inflection. Dylan systematically educates the audience on three core drivers: model capability, adjacent AI tooling, and coding agents expanding the developer TAM.10:11–13:50 · Guest teaching 3/10 Pure-Play Voice Infrastructure and Scaling Engineering Molly inquires about competitive positioning against full-stack voice providers like ElevenLabs and Sierra. Dylan clarifies AssemblyAI's dedicated pure-play infrastructure focus across inference and orchestration layers rather than consumer application interfaces.13:50–17:05 · Guest teaching 1/10 Sponsor Spotlight: Brex Agentic Finance Platform Following an ad read for Brex, Molly asks how Dylan internally restructured AssemblyAI with AI tooling amid the SaaS shift. Dylan describes internal workflow automations, including personal executive assistants and autonomous web deployment.17:05–22:33 · Guest teaching 3/10 Sovereign AI and Enterprise Privacy Deployments Molly introduces an investor perspective suggesting mice and keyboards will soon be entirely obsolete. Dylan gently counters this maximalist view, pointing out that touchscreens did not replace keyboards and predicting voice will serve as an additive dimension rather than a wholesale replacement.22:33–25:14 · Guest teaching 4/10 Humanoid Robotics and Acoustic Disambiguation Challenges Molly inquires about humanoid robotics applications and suggests models distinguish speech via listening data rather than visual feeds. Dylan elaborates that even with visual sensors, acoustic disambiguation of overlapping speakers remains a core unsolved technical bottleneck.25:14–28:02 · Guest teaching 4/10 Navigating Cultural Nuance in Multilingual Voice AI Molly asks Dylan about the complexities of multilingual translation and whether AssemblyAI would acquire local specialized vendors. Dylan explains that localized performance depends heavily on cultural policy alignment and native nuance rather than pure underlying architecture.28:02–31:28 · Guest teaching 2/10 Sponsor Spotlight: MongoDB for AI Applications After ad reads for MongoDB and AssemblyAI, Molly playfully asks whether technology will enable animal translation, citing sci-fi. Dylan humorously debunks the voice premise, noting mind-reading interfaces or sub-vocal scans would be needed over acoustic translation.31:28–34:04 · Guest teaching 4/10 Training Regimes, Data Alignment, and Real-World Evals Molly asks how models are trained across millions of audio hours. Dylan details the reality of speech engineering, explaining that public benchmarks are easily gamed and real performance requires extensive custom evaluation suites and domain-specific background noise handling.34:04–36:48 · Guest teaching 1/10 AssemblyAI Team Structure, Founder Roots, and Infrastructure DNA Molly prompts Dylan to explore AssemblyAI's internal team breakdown and personal technical roots. Dylan recounts learning to code in college, participating in IRC communities, and defining the company's core identity around developer infrastructure.36:48–42:23 · Guest teaching 4/10 Future Outlook: On-Device Hardware and the Voice Agent UX Dilemma Dylan outlines the UX dilemma of voice agents attempting to deceptively mimic humans, while Molly counters that users will simply assume everything is AI. Dylan pushes back with a healthcare counterexample, illustrating why deceptive AI interaction induces user discomfort in sensitive settings.42:23–43:03 · Guest teaching 3/10 Final Technical Reflections: Context-Aware Voice Models Molly asks if any topics were missed, and Dylan highlights new context-aware models. Molly expresses surprise that voice systems have lacked environmental context until now, and Dylan confirms that AssemblyAI only recently engineered this breakthrough.1:06–3:33 · Guest disagreement 0/10 AssemblyAI's Unprecedented Scale and Rapid Volume Inflection Molly opens by explaining her own use case of voice transcription for VC investment memos before prompting Dylan to quantify AssemblyAI's scale. Dylan explains the platform's massive growth metrics, including handling over four times YouTube's daily volume, while Molly reacts receptively.3:33–6:39 · Guest disagreement 0/10 Y Combinator Origins and the Autonomous Vehicle Parallel Molly asks whether Dylan anticipated current voice adoption during AssemblyAI's early YC days. Dylan explains the YC AI batch origins and articulates his autonomous vehicle thesis, noting that market adoption unlocks progressively as core model thresholds are met.6:39–10:11 · Guest disagreement 0/10 Three Macro Drivers Accelerating Voice AI Adoption Molly asks an open prompt about macro trends driving voice inflection. Dylan systematically educates the audience on three core drivers: model capability, adjacent AI tooling, and coding agents expanding the developer TAM.10:11–13:50 · Guest disagreement 0/10 Pure-Play Voice Infrastructure and Scaling Engineering Molly inquires about competitive positioning against full-stack voice providers like ElevenLabs and Sierra. Dylan clarifies AssemblyAI's dedicated pure-play infrastructure focus across inference and orchestration layers rather than consumer application interfaces.13:50–17:05 · Guest disagreement 0/10 Sponsor Spotlight: Brex Agentic Finance Platform Following an ad read for Brex, Molly asks how Dylan internally restructured AssemblyAI with AI tooling amid the SaaS shift. Dylan describes internal workflow automations, including personal executive assistants and autonomous web deployment.17:05–22:33 · Guest disagreement 2/10 Sovereign AI and Enterprise Privacy Deployments Molly introduces an investor perspective suggesting mice and keyboards will soon be entirely obsolete. Dylan gently counters this maximalist view, pointing out that touchscreens did not replace keyboards and predicting voice will serve as an additive dimension rather than a wholesale replacement.22:33–25:14 · Guest disagreement 1/10 Humanoid Robotics and Acoustic Disambiguation Challenges Molly inquires about humanoid robotics applications and suggests models distinguish speech via listening data rather than visual feeds. Dylan elaborates that even with visual sensors, acoustic disambiguation of overlapping speakers remains a core unsolved technical bottleneck.25:14–28:02 · Guest disagreement 0/10 Navigating Cultural Nuance in Multilingual Voice AI Molly asks Dylan about the complexities of multilingual translation and whether AssemblyAI would acquire local specialized vendors. Dylan explains that localized performance depends heavily on cultural policy alignment and native nuance rather than pure underlying architecture.28:02–31:28 · Guest disagreement 1/10 Sponsor Spotlight: MongoDB for AI Applications After ad reads for MongoDB and AssemblyAI, Molly playfully asks whether technology will enable animal translation, citing sci-fi. Dylan humorously debunks the voice premise, noting mind-reading interfaces or sub-vocal scans would be needed over acoustic translation.31:28–34:04 · Guest disagreement 1/10 Training Regimes, Data Alignment, and Real-World Evals Molly asks how models are trained across millions of audio hours. Dylan details the reality of speech engineering, explaining that public benchmarks are easily gamed and real performance requires extensive custom evaluation suites and domain-specific background noise handling.34:04–36:48 · Guest disagreement 0/10 AssemblyAI Team Structure, Founder Roots, and Infrastructure DNA Molly prompts Dylan to explore AssemblyAI's internal team breakdown and personal technical roots. Dylan recounts learning to code in college, participating in IRC communities, and defining the company's core identity around developer infrastructure.36:48–42:23 · Guest disagreement 2/10 Future Outlook: On-Device Hardware and the Voice Agent UX Dilemma Dylan outlines the UX dilemma of voice agents attempting to deceptively mimic humans, while Molly counters that users will simply assume everything is AI. Dylan pushes back with a healthcare counterexample, illustrating why deceptive AI interaction induces user discomfort in sensitive settings.42:23–43:03 · Guest disagreement 0/10 Final Technical Reflections: Context-Aware Voice Models Molly asks if any topics were missed, and Dylan highlights new context-aware models. Molly expresses surprise that voice systems have lacked environmental context until now, and Dylan confirms that AssemblyAI only recently engineered this breakthrough.1:06–3:33 · Molly pushing back 0/10 AssemblyAI's Unprecedented Scale and Rapid Volume Inflection Molly opens by explaining her own use case of voice transcription for VC investment memos before prompting Dylan to quantify AssemblyAI's scale. Dylan explains the platform's massive growth metrics, including handling over four times YouTube's daily volume, while Molly reacts receptively.3:33–6:39 · Molly pushing back 0/10 Y Combinator Origins and the Autonomous Vehicle Parallel Molly asks whether Dylan anticipated current voice adoption during AssemblyAI's early YC days. Dylan explains the YC AI batch origins and articulates his autonomous vehicle thesis, noting that market adoption unlocks progressively as core model thresholds are met.6:39–10:11 · Molly pushing back 0/10 Three Macro Drivers Accelerating Voice AI Adoption Molly asks an open prompt about macro trends driving voice inflection. Dylan systematically educates the audience on three core drivers: model capability, adjacent AI tooling, and coding agents expanding the developer TAM.10:11–13:50 · Molly pushing back 0/10 Pure-Play Voice Infrastructure and Scaling Engineering Molly inquires about competitive positioning against full-stack voice providers like ElevenLabs and Sierra. Dylan clarifies AssemblyAI's dedicated pure-play infrastructure focus across inference and orchestration layers rather than consumer application interfaces.13:50–17:05 · Molly pushing back 0/10 Sponsor Spotlight: Brex Agentic Finance Platform Following an ad read for Brex, Molly asks how Dylan internally restructured AssemblyAI with AI tooling amid the SaaS shift. Dylan describes internal workflow automations, including personal executive assistants and autonomous web deployment.17:05–22:33 · Molly pushing back 1/10 Sovereign AI and Enterprise Privacy Deployments Molly introduces an investor perspective suggesting mice and keyboards will soon be entirely obsolete. Dylan gently counters this maximalist view, pointing out that touchscreens did not replace keyboards and predicting voice will serve as an additive dimension rather than a wholesale replacement.22:33–25:14 · Molly pushing back 0/10 Humanoid Robotics and Acoustic Disambiguation Challenges Molly inquires about humanoid robotics applications and suggests models distinguish speech via listening data rather than visual feeds. Dylan elaborates that even with visual sensors, acoustic disambiguation of overlapping speakers remains a core unsolved technical bottleneck.25:14–28:02 · Molly pushing back 0/10 Navigating Cultural Nuance in Multilingual Voice AI Molly asks Dylan about the complexities of multilingual translation and whether AssemblyAI would acquire local specialized vendors. Dylan explains that localized performance depends heavily on cultural policy alignment and native nuance rather than pure underlying architecture.28:02–31:28 · Molly pushing back 0/10 Sponsor Spotlight: MongoDB for AI Applications After ad reads for MongoDB and AssemblyAI, Molly playfully asks whether technology will enable animal translation, citing sci-fi. Dylan humorously debunks the voice premise, noting mind-reading interfaces or sub-vocal scans would be needed over acoustic translation.31:28–34:04 · Molly pushing back 0/10 Training Regimes, Data Alignment, and Real-World Evals Molly asks how models are trained across millions of audio hours. Dylan details the reality of speech engineering, explaining that public benchmarks are easily gamed and real performance requires extensive custom evaluation suites and domain-specific background noise handling.34:04–36:48 · Molly pushing back 0/10 AssemblyAI Team Structure, Founder Roots, and Infrastructure DNA Molly prompts Dylan to explore AssemblyAI's internal team breakdown and personal technical roots. Dylan recounts learning to code in college, participating in IRC communities, and defining the company's core identity around developer infrastructure.36:48–42:23 · Molly pushing back 2/10 Future Outlook: On-Device Hardware and the Voice Agent UX Dilemma Dylan outlines the UX dilemma of voice agents attempting to deceptively mimic humans, while Molly counters that users will simply assume everything is AI. Dylan pushes back with a healthcare counterexample, illustrating why deceptive AI interaction induces user discomfort in sensitive settings.42:23–43:03 · Molly pushing back 1/10 Final Technical Reflections: Context-Aware Voice Models Molly asks if any topics were missed, and Dylan highlights new context-aware models. Molly expresses surprise that voice systems have lacked environmental context until now, and Dylan confirms that AssemblyAI only recently engineered this breakthrough.

speaking balance: gold is Molly, purple is the guest (3 minute bins)

0:00 · Molly 36.1% · guest 63.9%0:00 · Molly 36.1% · guest 63.9%3:00 · Molly 7.3% · guest 92.7%3:00 · Molly 7.3% · guest 92.7%6:00 · Molly 1% · guest 99%6:00 · Molly 1% · guest 99%9:00 · Molly 8% · guest 92%9:00 · Molly 8% · guest 92%12:00 · Molly 40.9% · guest 59.1%12:00 · Molly 40.9% · guest 59.1%15:00 · Molly 15.8% · guest 84.2%15:00 · Molly 15.8% · guest 84.2%18:00 · Molly 11.4% · guest 88.6%18:00 · Molly 11.4% · guest 88.6%21:00 · Molly 4.5% · guest 95.5%21:00 · Molly 4.5% · guest 95.5%24:00 · Molly 2.7% · guest 97.3%24:00 · Molly 2.7% · guest 97.3%27:00 · Molly 57.5% · guest 42.5%27:00 · Molly 57.5% · guest 42.5%30:00 · Molly 27.1% · guest 72.9%30:00 · Molly 27.1% · guest 72.9%33:00 · Molly 3.5% · guest 96.5%33:00 · Molly 3.5% · guest 96.5%36:00 · Molly 6.2% · guest 93.8%36:00 · Molly 6.2% · guest 93.8%39:00 · Molly 2.5% · guest 97.5%39:00 · Molly 2.5% · guest 97.5%42:00 · Molly 21.8% · guest 78.2%42:00 · Molly 21.8% · guest 78.2%45:00 · Molly 79.9% · guest 20.1%45:00 · Molly 79.9% · guest 20.1%
Sharpest disagreement ▶ 41:28 Dylan rejects assuming everything is AI

When Molly suggests users should just assume all calls are automated AI, Dylan directly counters with a medical nurse helpline scenario to demonstrate why deceptive anthropomorphic agents generate significant distrust.

Hardest push from Molly ▶ 42:46 Molly questions the novelty of contextual voice models

Molly challenges Dylan's presentation of environmental context as a novel feature, questioning how existing commercial drive-through voice systems could possibly lack basic environmental awareness.

Biggest teaching moment ▶ 33:20 Dylan exposes deceptive open-source speech benchmarks

Dylan educates Molly on model evaluation, explaining that public benchmarks are easily optimized for marketing vanity, whereas enterprise utility requires fine-grained filtering like distinguishing between ordering customers and backseat screaming.

Molly holds their own ▶ 1:09 Molly demonstrates operational voice workflow expertise

Molly illustrates her practical domain fluency right at the outset, detailing how she captures audio data to synthesize automated investor calls and open-source memos.

the scores for every segment, with the reasoning behind each
ChapterTopicMolly as informed peerGuest teachingGuest disagreementMolly pushing backWhy
AssemblyAI's Unprecedented Scale and Rapid Volume Inflection 4200 Molly opens by explaining her own use case of voice transcription for VC investment memos before prompting Dylan to quantify AssemblyAI's scale. Dylan explains the platform's massive growth metrics, including handling over four times YouTube's daily volume, while Molly reacts receptively.
Y Combinator Origins and the Autonomous Vehicle Parallel 3300 Molly asks whether Dylan anticipated current voice adoption during AssemblyAI's early YC days. Dylan explains the YC AI batch origins and articulates his autonomous vehicle thesis, noting that market adoption unlocks progressively as core model thresholds are met.
Three Macro Drivers Accelerating Voice AI Adoption 2400 Molly asks an open prompt about macro trends driving voice inflection. Dylan systematically educates the audience on three core drivers: model capability, adjacent AI tooling, and coding agents expanding the developer TAM.
Pure-Play Voice Infrastructure and Scaling Engineering 4300 Molly inquires about competitive positioning against full-stack voice providers like ElevenLabs and Sierra. Dylan clarifies AssemblyAI's dedicated pure-play infrastructure focus across inference and orchestration layers rather than consumer application interfaces.
Sponsor Spotlight: Brex Agentic Finance Platform 3100 Following an ad read for Brex, Molly asks how Dylan internally restructured AssemblyAI with AI tooling amid the SaaS shift. Dylan describes internal workflow automations, including personal executive assistants and autonomous web deployment.
Sovereign AI and Enterprise Privacy Deployments 4321 Molly introduces an investor perspective suggesting mice and keyboards will soon be entirely obsolete. Dylan gently counters this maximalist view, pointing out that touchscreens did not replace keyboards and predicting voice will serve as an additive dimension rather than a wholesale replacement.
Humanoid Robotics and Acoustic Disambiguation Challenges 4410 Molly inquires about humanoid robotics applications and suggests models distinguish speech via listening data rather than visual feeds. Dylan elaborates that even with visual sensors, acoustic disambiguation of overlapping speakers remains a core unsolved technical bottleneck.
Navigating Cultural Nuance in Multilingual Voice AI 3400 Molly asks Dylan about the complexities of multilingual translation and whether AssemblyAI would acquire local specialized vendors. Dylan explains that localized performance depends heavily on cultural policy alignment and native nuance rather than pure underlying architecture.
Sponsor Spotlight: MongoDB for AI Applications 2210 After ad reads for MongoDB and AssemblyAI, Molly playfully asks whether technology will enable animal translation, citing sci-fi. Dylan humorously debunks the voice premise, noting mind-reading interfaces or sub-vocal scans would be needed over acoustic translation.
Training Regimes, Data Alignment, and Real-World Evals 3410 Molly asks how models are trained across millions of audio hours. Dylan details the reality of speech engineering, explaining that public benchmarks are easily gamed and real performance requires extensive custom evaluation suites and domain-specific background noise handling.
AssemblyAI Team Structure, Founder Roots, and Infrastructure DNA 3100 Molly prompts Dylan to explore AssemblyAI's internal team breakdown and personal technical roots. Dylan recounts learning to code in college, participating in IRC communities, and defining the company's core identity around developer infrastructure.
Future Outlook: On-Device Hardware and the Voice Agent UX Dilemma 4422 Dylan outlines the UX dilemma of voice agents attempting to deceptively mimic humans, while Molly counters that users will simply assume everything is AI. Dylan pushes back with a healthcare counterexample, illustrating why deceptive AI interaction induces user discomfort in sensitive settings.
Final Technical Reflections: Context-Aware Voice Models 3301 Molly asks if any topics were missed, and Dylan highlights new context-aware models. Molly expresses surprise that voice systems have lacked environmental context until now, and Dylan confirms that AssemblyAI only recently engineered this breakthrough.

Statements from this episode (19)

Assertion Not checkable as stated
Fox: AssemblyAI weekly conversation volume grew 800 percent over three years
“One stat I was just looking at before I came over here was like the amount of weekly conversations that Assembly handles through our APIs every week is up over 800% over the last three years.”
Dylan Fox Jul 31, 2026 ▶ 2:21
Assertion Not checkable as stated
Fox: AssemblyAI processes four times YouTube's daily audio volume on peak weeks
“And so now, you know, on a given week, on a peak week, there'll be something like over a hundred and twenty million conversations, voice conversations going through our platform, over two million hours of voice, which as of December of this past year was four,…”
Dylan Fox Jul 31, 2026 ▶ 2:33
Assertion Not checkable as stated
Fox: AssemblyAI handles nearly 100 million daily API calls from developers
“There's, you know, almost a hundred million API calls a day coming against our API, about a million developers, a little over a million developers on the platform now.”
Dylan Fox Jul 31, 2026 ▶ 3:01
Insight
Fox: Voice AI Follows Self-Driving Cars in Tech-Threshold Adoption
“I always viewed this journey as similar to self-driving cars, where You know, it wasn't a question of like, oh, is the product market fit going to come? It was like, no, the technology just sucks. And as it gets better and better, you're going to continuously …”
Dylan Fox Jul 31, 2026 ▶ 5:47
Disclosure
Fox: Lawn care chains use Lovable and Cursor to build on AssemblyAI
“We saw this, like, small business, like, lawn care chain sign up, and, you know, we're like, what do they do with the API? And a lot of these small businesses are automating parts of their back office using Lovable, or Cloud Code, or Cursor, or whatever, Repli…”
Dylan Fox Jul 31, 2026 ▶ 8:53
Opinion
Fox: Coding Agents Have Expanded AssemblyAI's TAM by 100x
“So I think about this as like our TAM has just increased by a hundred X because we're not just selling to engineering teams within product companies. It's now like any, anyone.”
Dylan Fox Jul 31, 2026 ▶ 9:23
Assertion Not checkable as stated
Fox: Most AI note-takers use AssemblyAI as voice infrastructure
“And so the reason, you know, most AI note takers are using assembly as the voice infrastructure is because our models are the most scalable, right?”
Dylan Fox Jul 31, 2026 ▶ 12:08
Opinion
Fox: Big tech AI labs build models disconnected from actual customer needs
“We have an amazing team of engineers, of researchers that just operate so closely to customers that they really understand, like, how this stuff is being deployed. I think that's the biggest difference between us and, like, a lab at a bigger company. So, you k…”
Dylan Fox Jul 31, 2026 ▶ 12:52
Assertion Not checkable as stated
Fox: AssemblyAI used Claude to rebuild and migrate its entire website
“Like, you know, the most recent example is we had Cloud rebuild our whole website off of Webflow and just deployed on Vercel.”
Dylan Fox Jul 31, 2026 ▶ 16:43
Assertion Not checkable as stated
Fox: Customers abandon fine-tuned open-source voice models due to maintenance burdens
“We've seen some customers maybe six months ago or a year ago, they'll take an open source model and fine tune it or something. And then, you know, they're now like, okay, this is like, shit, this is like really behind and it's like nightmare to maintain and I …”
Dylan Fox Jul 31, 2026 ▶ 17:45
Assertion Not checkable as stated
Fox: Voice AI became reliable data capture only in the past year
“Like for the first time probably ever in the past year, it's like a reliable form of data capture.”
Dylan Fox Jul 31, 2026 ▶ 19:12
Prediction Not checkable as stated
Fox: Voice will augment, not replace, screens over the next couple years
“And so I don't think that means that computers are just going to be like a screen and you just talk to it, because actually we'd get tired of talking, but I think it's going to be a dimension that's added to everything, and so that over the next couple years w…”
Dylan Fox Jul 31, 2026 ▶ 21:30
Assertion Not checkable as stated
Fox: Multi-speaker acoustic disambiguation remains a major hurdle for humanoid robots
“One of the main problems the humanoid robots face today, because a lot of them are using our APIs If you have three people standing next to the robot, it doesn't know who to listen to. And it has a hard time disambiguating who's saying what.”
Dylan Fox Jul 31, 2026 ▶ 23:12
Prediction Not checkable as stated
Fox: In five years, kids will expect all devices to understand voice
“Across humanoid robots, consumer electronics, like, you're gonna see voice as this dimension that you just, like, expect And what, the example I think about is, you know, when you see pictures of, like, kids, like, trying to, like, swipe on TVs or something, b…”
Dylan Fox Jul 31, 2026 ▶ 24:42
Insight
Fox: Local voice AI vendors beat global generalists through cultural nuance
“If you look at, like, local voice AI vendors in, you know, certain countries, like, they actually are typically the best, because it's, you know, they speak the language, they understand it, and they can more quickly identify, like, which data is good and stuf…”
Dylan Fox Jul 31, 2026 ▶ 26:24
Insight
Fox: Training data drives roughly 75 percent of AI model performance
“It's really, probably, like, 75% of it is, like, the data that you're training on. Like, I would say for any AI model, it's like, there's always, like, you know, like, these, like, step functions and, like, algorithms and architectures and stuff, but the data …”
Dylan Fox Jul 31, 2026 ▶ 31:46
Prediction Not checkable as stated
Fox: Consumer voice hardware and software applications will surge within a year
“I think over the next year, we'll see a lot of, a lot more consumer applications, hardware and software, where voice is a core dimension. So, toys, games consumer electronics”
Dylan Fox Jul 31, 2026 ▶ 37:07
Disclosure
Fox: AssemblyAI is developing on-device models for low-power hardware
“We're, for example, working on on-device models, so models that can run, like, on a phone or on a, you know, really low-powered piece of hardware you know, like a remote control for a TV or something”
Dylan Fox Jul 31, 2026 ▶ 38:14
Assertion Not checkable as stated
Fox: Callers hang up immediately if voice agents disclose they are AI
“If our customers, if they're building a voice agent and you disclose up front that you're an AI, people just hang up versus if you don't, people continue.”
Dylan Fox Jul 31, 2026 ▶ 39:55
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 160 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.