Aug 11, 2017 · 30m · y-combinator

Baidu's AI Lab Director on Advancing Speech Recognition and Simulation · Y Combinator

Adam Coates · 23m spoken Craig Cannon · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Y Combinator interview, Adam Coates, Director of Baidu's Silicon Valley AI Lab, discusses bridging basic deep learning research with scalable commercial products like Deep Speech, TalkType, and SwiftScribe. He shares technical insights on voice AI, latency optimization, media literacy, and the rising demand for full-stack machine learning engineers.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 2.1 Guest teaching 3.6 Guest disagreement 0.1 The partners pushing back 0.1
05100:0010:0020:0030:000:44–3:09 · The partners as informed peer 1/10 Mission and Strategy of Baidu's AI Lab Host Craig Cannon asks standard introductory questions regarding how Baidu's AI Lab balances basic research against shipping real products. Guest Adam Coates describes their mission-driven threshold of impacting at least 100 million people without friction.3:09–6:27 · The partners as informed peer 2/10 Deep Speech: Scaling Superhuman Speech Recognition Cannon asks how Deep Speech was built and how much training data would be required to support new languages like German. Coates educates Cannon on the scale hypothesis and notes their systems consume 10,000 to 20,000 hours of labeled audio.6:27–11:52 · The partners as informed peer 3/10 Voice Simulation and Unsupervised Learning Techniques Cannon brings up Lyrebird's low-data voice simulation claims and contrasts audio transcription with captioning. Coates breaks down supervised versus unsupervised learning and explains how deep learning eliminates hand-engineered acoustic feature pipelines.11:52–20:17 · The partners as informed peer 3/10 Deep Voice and the Deep Learning Pipeline Cannon references OpenAI's language prediction research on Amazon reviews to guess how streaming audio models operate. Coates clarifies how online speech architectures adjust contextual output to minimize millisecond latencies for user experience.20:17–23:02 · The partners as informed peer 3/10 Next Frontier in Speech AI: SwiftScribe and Environmental Noise Cannon questions whether humans must adapt their phrasing for voice assistants based on personal travel experiences. Coates gently rejects the premise, projecting speech recognition will be human-level and a solved problem while outlining acoustic frontiers like reverberation and crosstalk.23:02–28:40 · The partners as informed peer 2/10 Social Implications and Assistive Tech in AI Cannon asks about deepfake risks, technological unemployment, and career guidance for machine learning candidates. Coates offers a grounded perspective focusing on adaptive media literacy, assistive tech benefits, and developing full-stack ML engineers.28:40–30:28 · The partners as informed peer 1/10 Startup Mindset for AI Leaders and Conclusion Cannon asks Coates for personal recommendations and closing thoughts. Coates shares how startup literature shaped his engineering philosophy around identifying core unknowns and maintaining rapid learning loops.0:44–3:09 · Guest teaching 2/10 Mission and Strategy of Baidu's AI Lab Host Craig Cannon asks standard introductory questions regarding how Baidu's AI Lab balances basic research against shipping real products. Guest Adam Coates describes their mission-driven threshold of impacting at least 100 million people without friction.3:09–6:27 · Guest teaching 4/10 Deep Speech: Scaling Superhuman Speech Recognition Cannon asks how Deep Speech was built and how much training data would be required to support new languages like German. Coates educates Cannon on the scale hypothesis and notes their systems consume 10,000 to 20,000 hours of labeled audio.6:27–11:52 · Guest teaching 5/10 Voice Simulation and Unsupervised Learning Techniques Cannon brings up Lyrebird's low-data voice simulation claims and contrasts audio transcription with captioning. Coates breaks down supervised versus unsupervised learning and explains how deep learning eliminates hand-engineered acoustic feature pipelines.11:52–20:17 · Guest teaching 4/10 Deep Voice and the Deep Learning Pipeline Cannon references OpenAI's language prediction research on Amazon reviews to guess how streaming audio models operate. Coates clarifies how online speech architectures adjust contextual output to minimize millisecond latencies for user experience.20:17–23:02 · Guest teaching 4/10 Next Frontier in Speech AI: SwiftScribe and Environmental Noise Cannon questions whether humans must adapt their phrasing for voice assistants based on personal travel experiences. Coates gently rejects the premise, projecting speech recognition will be human-level and a solved problem while outlining acoustic frontiers like reverberation and crosstalk.23:02–28:40 · Guest teaching 4/10 Social Implications and Assistive Tech in AI Cannon asks about deepfake risks, technological unemployment, and career guidance for machine learning candidates. Coates offers a grounded perspective focusing on adaptive media literacy, assistive tech benefits, and developing full-stack ML engineers.28:40–30:28 · Guest teaching 2/10 Startup Mindset for AI Leaders and Conclusion Cannon asks Coates for personal recommendations and closing thoughts. Coates shares how startup literature shaped his engineering philosophy around identifying core unknowns and maintaining rapid learning loops.0:44–3:09 · Guest disagreement 0/10 Mission and Strategy of Baidu's AI Lab Host Craig Cannon asks standard introductory questions regarding how Baidu's AI Lab balances basic research against shipping real products. Guest Adam Coates describes their mission-driven threshold of impacting at least 100 million people without friction.3:09–6:27 · Guest disagreement 0/10 Deep Speech: Scaling Superhuman Speech Recognition Cannon asks how Deep Speech was built and how much training data would be required to support new languages like German. Coates educates Cannon on the scale hypothesis and notes their systems consume 10,000 to 20,000 hours of labeled audio.6:27–11:52 · Guest disagreement 0/10 Voice Simulation and Unsupervised Learning Techniques Cannon brings up Lyrebird's low-data voice simulation claims and contrasts audio transcription with captioning. Coates breaks down supervised versus unsupervised learning and explains how deep learning eliminates hand-engineered acoustic feature pipelines.11:52–20:17 · Guest disagreement 0/10 Deep Voice and the Deep Learning Pipeline Cannon references OpenAI's language prediction research on Amazon reviews to guess how streaming audio models operate. Coates clarifies how online speech architectures adjust contextual output to minimize millisecond latencies for user experience.20:17–23:02 · Guest disagreement 1/10 Next Frontier in Speech AI: SwiftScribe and Environmental Noise Cannon questions whether humans must adapt their phrasing for voice assistants based on personal travel experiences. Coates gently rejects the premise, projecting speech recognition will be human-level and a solved problem while outlining acoustic frontiers like reverberation and crosstalk.23:02–28:40 · Guest disagreement 0/10 Social Implications and Assistive Tech in AI Cannon asks about deepfake risks, technological unemployment, and career guidance for machine learning candidates. Coates offers a grounded perspective focusing on adaptive media literacy, assistive tech benefits, and developing full-stack ML engineers.28:40–30:28 · Guest disagreement 0/10 Startup Mindset for AI Leaders and Conclusion Cannon asks Coates for personal recommendations and closing thoughts. Coates shares how startup literature shaped his engineering philosophy around identifying core unknowns and maintaining rapid learning loops.0:44–3:09 · The partners pushing back 0/10 Mission and Strategy of Baidu's AI Lab Host Craig Cannon asks standard introductory questions regarding how Baidu's AI Lab balances basic research against shipping real products. Guest Adam Coates describes their mission-driven threshold of impacting at least 100 million people without friction.3:09–6:27 · The partners pushing back 0/10 Deep Speech: Scaling Superhuman Speech Recognition Cannon asks how Deep Speech was built and how much training data would be required to support new languages like German. Coates educates Cannon on the scale hypothesis and notes their systems consume 10,000 to 20,000 hours of labeled audio.6:27–11:52 · The partners pushing back 0/10 Voice Simulation and Unsupervised Learning Techniques Cannon brings up Lyrebird's low-data voice simulation claims and contrasts audio transcription with captioning. Coates breaks down supervised versus unsupervised learning and explains how deep learning eliminates hand-engineered acoustic feature pipelines.11:52–20:17 · The partners pushing back 0/10 Deep Voice and the Deep Learning Pipeline Cannon references OpenAI's language prediction research on Amazon reviews to guess how streaming audio models operate. Coates clarifies how online speech architectures adjust contextual output to minimize millisecond latencies for user experience.20:17–23:02 · The partners pushing back 1/10 Next Frontier in Speech AI: SwiftScribe and Environmental Noise Cannon questions whether humans must adapt their phrasing for voice assistants based on personal travel experiences. Coates gently rejects the premise, projecting speech recognition will be human-level and a solved problem while outlining acoustic frontiers like reverberation and crosstalk.23:02–28:40 · The partners pushing back 0/10 Social Implications and Assistive Tech in AI Cannon asks about deepfake risks, technological unemployment, and career guidance for machine learning candidates. Coates offers a grounded perspective focusing on adaptive media literacy, assistive tech benefits, and developing full-stack ML engineers.28:40–30:28 · The partners pushing back 0/10 Startup Mindset for AI Leaders and Conclusion Cannon asks Coates for personal recommendations and closing thoughts. Coates shares how startup literature shaped his engineering philosophy around identifying core unknowns and maintaining rapid learning loops.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 20:54 Reframing machine adaptation versus human speech habits

Coates politely pushes back on Cannon's premise that humans must adjust phrasing for machines, stating he expects speech recognition to match full human capability.

Hardest push from the partners ▶ 6:27 Inquiring about low-data voice simulation claims

Cannon presses Coates on how emerging startups like Lyrebird claim to emulate voices with minimal audio compared to Baidu's vast data requirements.

Biggest teaching moment ▶ 9:18 Explaining the shift away from hand-crafted phoneme pipelines

Coates explains the conceptual leap in modern speech AI, detailing how end-to-end deep learning bypassed decades of hand-engineered linguistic and acoustic decoders.

The partners hold their own ▶ 19:15 Connecting speech latency models to OpenAI review prediction research

Cannon actively demonstrates domain awareness by linking Coates' latency explanation to OpenAI's contemporary work predicting text in Amazon reviews.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Mission and Strategy of Baidu's AI Lab 1200 Host Craig Cannon asks standard introductory questions regarding how Baidu's AI Lab balances basic research against shipping real products. Guest Adam Coates describes their mission-driven threshold of impacting at least 100 million people without friction.
Deep Speech: Scaling Superhuman Speech Recognition 2400 Cannon asks how Deep Speech was built and how much training data would be required to support new languages like German. Coates educates Cannon on the scale hypothesis and notes their systems consume 10,000 to 20,000 hours of labeled audio.
Voice Simulation and Unsupervised Learning Techniques 3500 Cannon brings up Lyrebird's low-data voice simulation claims and contrasts audio transcription with captioning. Coates breaks down supervised versus unsupervised learning and explains how deep learning eliminates hand-engineered acoustic feature pipelines.
Deep Voice and the Deep Learning Pipeline 3400 Cannon references OpenAI's language prediction research on Amazon reviews to guess how streaming audio models operate. Coates clarifies how online speech architectures adjust contextual output to minimize millisecond latencies for user experience.
Next Frontier in Speech AI: SwiftScribe and Environmental Noise 3411 Cannon questions whether humans must adapt their phrasing for voice assistants based on personal travel experiences. Coates gently rejects the premise, projecting speech recognition will be human-level and a solved problem while outlining acoustic frontiers like reverberation and crosstalk.
Social Implications and Assistive Tech in AI 2400 Cannon asks about deepfake risks, technological unemployment, and career guidance for machine learning candidates. Coates offers a grounded perspective focusing on adaptive media literacy, assistive tech benefits, and developing full-stack ML engineers.
Startup Mindset for AI Leaders and Conclusion 1200 Cannon asks Coates for personal recommendations and closing thoughts. Coates shares how startup literature shaped his engineering philosophy around identifying core unknowns and maintaining rapid learning loops.

Statements from this episode (12)

Assertion Supported
Baidu's Deep Speech engine achieves superhuman accuracy on short queries
“The speech engine that we've built at Baidu called Deep Speech it's actually superhuman for these short queries.”
Adam Coates Aug 11, 2017 ▶ 3:49
Disclosure
Baidu trains English speech AI on 10,000 to 20,000 audio hours
“Our English system uses like 10 to 20,000 hours of audio. The Mandarin systems are using even more for top end products.”
Adam Coates Aug 11, 2017 ▶ 5:48
Insight
Pre-training on thousands of voices enables low-data voice cloning
“If I learn to mimic lots of different voices and then you give me the 1001st voice you'd hope that the first thousand taught you virtually everything you need to know about language and that what's left is really some idiosyncratic change. That you could learn…”
Adam Coates Aug 11, 2017 ▶ 6:59
Insight
Deep learning and massive data eliminate hand-engineered speech recognition pipelines
“And in the past, it always looked like there was some fundamental problem that maybe we could never escape this need for these hand engineered representations. But it turns out that once you have enough data, all of those things go away.”
Adam Coates Aug 11, 2017 ▶ 10:23
Disclosure
Baidu crowdsources cheap audiobook recordings to acquire English training data
“We actually do a lot of clever tricks in English where we don't have a lot of a large number of English language products. So for example, it turns out that if you go onto, say, a crowdsourcing service, you can hire people very cheaply to just read books to yo…”
Adam Coates Aug 11, 2017 ▶ 10:41
Insight
Neural networks learn English spelling and exceptions purely from data
“But now it's all data-driven, so if I hear enough of these unusual words, You see these neural networks actually learn to spell on their own, even considering all the weird exceptions of English.”
Adam Coates Aug 11, 2017 ▶ 11:29
Prediction Not checkable as stated
Next-generation AI will transition from bolted-on features to immersive products
“So as we have the next wave of AI products I think we're going to move from these sort of bolted on AI features to really immersive AI products.”
Adam Coates Aug 11, 2017 ▶ 13:54
Insight
Voice AI latency of 200ms versus 100ms is clearly perceptible to users
“That the difference between 50 or a hundred milliseconds of latency and 200 milliseconds of latency is actually quite perceptible. And it really anything we can do to bring that down actually affects user experience quite a bit.”
Adam Coates Aug 11, 2017 ▶ 18:04
Prediction Not checkable as stated
Speech recognition could be a solved problem within a few years
“I sincerely think there's a chance that over the next few years we're going to regard speech recognition as a solved problem.”
Adam Coates Aug 11, 2017 ▶ 21:05
Insight
Society must apply text-style skepticism to generative AI media
“I think culturally we're all going to have to exercise a lot of critical thinking. We, we've always had this problem in some sense that I can read an article that has someone's name on it and notwithstanding understanding writing style I don't know for sure wh…”
Adam Coates Aug 11, 2017 ▶ 23:30
Opinion
Workers are not currently at risk of AI taking their jobs
“I don't think we're at risk of robots taking our jobs right now.”
Adam Coates Aug 11, 2017 ▶ 25:51
Insight
Rapid AI progress demands flexible full-stack machine learning engineers
“Because the field's moving so quickly we also need a different kind of person now. We also need people who are sort of chameleons, who are these highly flexible types that can understand and even contribute to a research project but can also simultaneously Shi…”
Adam Coates Aug 11, 2017 ▶ 26:21
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.