Insight certainty 4/5 debate potential 2/5

Most speech AI value comes from downstream workflows, not just transcription

Dylan Fox · Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox · May 1, 2023 · at 3:15

Dylan Fox is the Founder and CEO of AssemblyAI. He discusses how enterprise product teams derive value from speech AI models.

0:00 / 0:07exact quote · 7.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Where a lot of value is created is you're taking the transcription, and then you're using it as an input to do something else.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dylan Fox

Opinion
Major cloud providers are too big to ship good developer products
“I'm surprised the big cloud companies, and apologies if anyone here works there, I'm surprised they can't ship, you know, better developer products, but it's, I think they're maybe just too big at this point.”
Dylan Fox May 1, 2023 ▶ 23:20 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Disclosure
Fox: AssemblyAI delays exploratory model investment until developers prove value
“Some of these other things that are more exploratory, like, we're not gonna put a ton of effort into those until we see that our customers the developers that use our API are actually able to find value and create value with those.”
Dylan Fox May 1, 2023 ▶ 7:42 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Assertion Not checkable as stated
Most commercial speech models historically trained on roughly 50,000 hours of audio
“And our models prior, and most commercial speech recognition models trained on like, 50,000 hours”
Dylan Fox May 1, 2023 ▶ 14:09 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Disclosure
AssemblyAI is training its next speech model on ~4 million hours of audio
“We're actually training conformer two or what might call it 1.5, but whatever this accessory will be is training right now. And that's something around four million hours of labeled audio data.”
Dylan Fox May 1, 2023 ▶ 14:19 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Assertion Not checkable as stated
AssemblyAI processes over 100 million audio files monthly via API
“We've processed, ah, yeah, it's like over a hundred million audio files a month that are flowing through the API, and that's growing pretty quickly.”
Dylan Fox May 1, 2023 ▶ 16:30 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Assertion Supported
State-of-the-art speech recognition models still carry a 15% error rate
“State-of-the-art automatic speech recognition still has, like, a 15% error rate on a lot of data sets”
Dylan Fox May 1, 2023 ▶ 20:05 Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.