“We're actually training conformer two or what might call it 1.5, but whatever this accessory will be is training right now. And that's something around four million hours of labeled audio data.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dylan Fox
Opinion
Major cloud providers are too big to ship good developer products
“I'm surprised the big cloud companies, and apologies if anyone here works there, I'm surprised they can't ship, you know, better developer products, but it's, I think they're maybe just too big at this point.”
Dylan FoxMay 1, 2023▶ 23:20Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Insight
Most speech AI value comes from downstream workflows, not just transcription
“Where a lot of value is created is you're taking the transcription, and then you're using it as an input to do something else.”
Dylan FoxMay 1, 2023▶ 3:15Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Disclosure
Fox: AssemblyAI delays exploratory model investment until developers prove value
“Some of these other things that are more exploratory, like, we're not gonna put a ton of effort into those until we see that our customers the developers that use our API are actually able to find value and create value with those.”
Dylan FoxMay 1, 2023▶ 7:42Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
AssertionNot checkable as stated
Most commercial speech models historically trained on roughly 50,000 hours of audio
“And our models prior, and most commercial speech recognition models trained on like, 50,000 hours”
Dylan FoxMay 1, 2023▶ 14:09Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
AssertionNot checkable as stated
AssemblyAI processes over 100 million audio files monthly via API
“We've processed, ah, yeah, it's like over a hundred million audio files a month that are flowing through the API, and that's growing pretty quickly.”
Dylan FoxMay 1, 2023▶ 16:30Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
AssertionSupported
State-of-the-art speech recognition models still carry a 15% error rate
“State-of-the-art automatic speech recognition still has, like, a 15% error rate on a lot of data sets”
Dylan FoxMay 1, 2023▶ 20:05Generative AI for Speech Recognition | AssemblyAI Founder & CEO, Dylan Fox
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.