Jul 5, 2024 · 2h 18m · latent-space

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka

Yi Tay · 1h 33m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview, former Google Brain researcher and Reka co-founder Yi Tay provides insider perspectives on transformer architectures, scaling laws, startup compute hurdles, and the evolving research metagame of frontier artificial intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 6.0 Guest teaching 5.5 Guest disagreement 2.4 The hosts pushing back 2.4
05100:0020:0040:001:00:001:20:001:40:002:00:001:39–7:58 · The hosts as informed peer 6/10 Evolution of the AI Research Metagame The host provides a well-informed recap of Yi Tay's career shift from academic NLP to foundational models. Yi explains how the research meta shifted from task-specific architectures to universal foundation models.7:59–13:20 · The hosts as informed peer 5/10 Foundational Research Breakthroughs at Google Brain The host inquires about the origins of major breakthrough projects like UL2 and Flan. Yi explains the dynamics inside Google Brain between bottom-up holiday tinker projects and top-down org initiatives.13:23–18:30 · The hosts as informed peer 5/10 Co-Leading PaLM 2 and Scaling Career Opportunities The host presses on the career mechanics of becoming a co-lead on PaLM 2, asking if it was a deliberate strategy. Yi downplays intentionality, attributing it to organic visibility from UL2 and situational luck.18:31–22:26 · The hosts as informed peer 7/10 Emergent Abilities and the Mirage Paper Controversy The host passionately defends the concept of emergent abilities against the Mirage paper, displaying deep familiarity with the evaluation debate. Yi agrees that emergent capabilities are real while expressing annoyance at academic best paper hype.22:27–36:02 · The hosts as informed peer 5/10 Mentorship, Collaborations, and Research Taste at Google Brain Yi shares key insights learned from Quoc Le, Jason Wei, and Hyung Won Chung, emphasizing PR, taste, and engineering discipline. The host adds his own framework on picking up what mentors leave behind.36:03–41:09 · The hosts as informed peer 5/10 Essential Qualities and Technical Habits of AI Researchers Yi highlights the extreme dedication and 4 AM debugging mindset necessary for LLM researchers while cautioning against burnout. The host asks for specific tech stack recommendations, but Yi advocates framework agnosticism.41:10–46:09 · The hosts as informed peer 5/10 Founding Reka and Leaving Big Tech The host asks why Yi chose to co-found Reka instead of joining established startups like Mistral or Inflection. Yi explains his desire for a genuine co-founding learning experience rather than repeating big-tech dynamics.46:09–59:44 · The hosts as informed peer 7/10 Compute Infrastructure Challenges and GPU Node Instability Yi details the brutal operational reality of bad GPU nodes killing full cluster runs and risk-sharing issues with compute providers. The host pushes back technically, questioning why frequent checkpointing and standard orchestration don't mitigate the loss.59:44–1:05:29 · The hosts as informed peer 5/10 Training Reka Flash, Core, and Edge with a Lean Team The host asks how a 3-5 person pre-training team had the confidence to beat frontier baselines. Yi explains de-risking via 4B ablation runs and rapid iterative hill-climbing rather than predictive certainty.1:05:29–1:20:03 · The hosts as informed peer 6/10 Analyzing Transformer Architectures: Noam Baselines and Encoder-Decoder Designs Yi delivers a detailed technical breakdown of Noam Shazeer's architectural contributions and the intrinsic 2x FLOP parameter efficiency of encoder-decoder models. The host acknowledges learning a major new architectural perspective.1:20:04–1:29:45 · The hosts as informed peer 7/10 Evaluating Open LLMs, Benchmark Saturation, and Incentive Mismatches Yi and the host critique benchmark saturation and contamination on MMLU and GSM8K, discussing academic publication incentives. The host proposes formulaic, seed-generated parametric evaluations to prevent benchmark cheating.1:29:46–1:35:29 · The hosts as informed peer 6/10 Multimodal Architecture Paradigms: Early Fusion and Screen Intelligence The conversation covers multimodal early versus late fusion and Adept's screen-first thesis. Yi argues that early fusion will become universal and models must master both natural imagery and interface screens concurrently.1:35:29–1:44:31 · The hosts as informed peer 7/10 Pushing Past Chinchilla Scaling and Long Context vs. RAG The host challenges the superiority of long context over RAG in production environments due to inference costs. Yi counters that complex reasoning and whole-document synthesis fundamentally break under RAG chunking limitations.1:44:32–1:58:22 · The hosts as informed peer 8/10 The Reality of Efficiency Research and Mixture-of-Experts The host cites external analyses on why most toy efficiency research fails to scale. Yi agrees forcefully, explaining how unoptimized kernels and non-throughput-matched parameter reductions deceive academic reviewers.1:58:22–2:02:20 · The hosts as informed peer 6/10 Open Source vs. Closed Source AI Dynamics Yi rejects romanticized narratives of decentralized open-source development, arguing that grassroots fine-tuning waves rely on distilling proprietary models and gaming leaderboards rather than true frontier training.2:02:20–2:08:23 · The hosts as informed peer 6/10 Personal Productivity, Research Execution, and Writing Habits Yi shares his productivity routine of pre-drafting paper narratives in Overleaf before experiments finish, while the host outlines his own automated newsletter information funnel.2:08:24–2:18:39 · The hosts as informed peer 6/10 Global AI Talent, Research Culture in Singapore, and the US Ecosystem The host and Yi explore the cultural differences between US impact-driven research and publishing-focused Asian academia, as well as global talent drain and national AI policy.1:39–7:58 · Guest teaching 4/10 Evolution of the AI Research Metagame The host provides a well-informed recap of Yi Tay's career shift from academic NLP to foundational models. Yi explains how the research meta shifted from task-specific architectures to universal foundation models.7:59–13:20 · Guest teaching 5/10 Foundational Research Breakthroughs at Google Brain The host inquires about the origins of major breakthrough projects like UL2 and Flan. Yi explains the dynamics inside Google Brain between bottom-up holiday tinker projects and top-down org initiatives.13:23–18:30 · Guest teaching 4/10 Co-Leading PaLM 2 and Scaling Career Opportunities The host presses on the career mechanics of becoming a co-lead on PaLM 2, asking if it was a deliberate strategy. Yi downplays intentionality, attributing it to organic visibility from UL2 and situational luck.18:31–22:26 · Guest teaching 3/10 Emergent Abilities and the Mirage Paper Controversy The host passionately defends the concept of emergent abilities against the Mirage paper, displaying deep familiarity with the evaluation debate. Yi agrees that emergent capabilities are real while expressing annoyance at academic best paper hype.22:27–36:02 · Guest teaching 6/10 Mentorship, Collaborations, and Research Taste at Google Brain Yi shares key insights learned from Quoc Le, Jason Wei, and Hyung Won Chung, emphasizing PR, taste, and engineering discipline. The host adds his own framework on picking up what mentors leave behind.36:03–41:09 · Guest teaching 6/10 Essential Qualities and Technical Habits of AI Researchers Yi highlights the extreme dedication and 4 AM debugging mindset necessary for LLM researchers while cautioning against burnout. The host asks for specific tech stack recommendations, but Yi advocates framework agnosticism.41:10–46:09 · Guest teaching 4/10 Founding Reka and Leaving Big Tech The host asks why Yi chose to co-found Reka instead of joining established startups like Mistral or Inflection. Yi explains his desire for a genuine co-founding learning experience rather than repeating big-tech dynamics.46:09–59:44 · Guest teaching 6/10 Compute Infrastructure Challenges and GPU Node Instability Yi details the brutal operational reality of bad GPU nodes killing full cluster runs and risk-sharing issues with compute providers. The host pushes back technically, questioning why frequent checkpointing and standard orchestration don't mitigate the loss.59:44–1:05:29 · Guest teaching 5/10 Training Reka Flash, Core, and Edge with a Lean Team The host asks how a 3-5 person pre-training team had the confidence to beat frontier baselines. Yi explains de-risking via 4B ablation runs and rapid iterative hill-climbing rather than predictive certainty.1:05:29–1:20:03 · Guest teaching 8/10 Analyzing Transformer Architectures: Noam Baselines and Encoder-Decoder Designs Yi delivers a detailed technical breakdown of Noam Shazeer's architectural contributions and the intrinsic 2x FLOP parameter efficiency of encoder-decoder models. The host acknowledges learning a major new architectural perspective.1:20:04–1:29:45 · Guest teaching 6/10 Evaluating Open LLMs, Benchmark Saturation, and Incentive Mismatches Yi and the host critique benchmark saturation and contamination on MMLU and GSM8K, discussing academic publication incentives. The host proposes formulaic, seed-generated parametric evaluations to prevent benchmark cheating.1:29:46–1:35:29 · Guest teaching 6/10 Multimodal Architecture Paradigms: Early Fusion and Screen Intelligence The conversation covers multimodal early versus late fusion and Adept's screen-first thesis. Yi argues that early fusion will become universal and models must master both natural imagery and interface screens concurrently.1:35:29–1:44:31 · Guest teaching 6/10 Pushing Past Chinchilla Scaling and Long Context vs. RAG The host challenges the superiority of long context over RAG in production environments due to inference costs. Yi counters that complex reasoning and whole-document synthesis fundamentally break under RAG chunking limitations.1:44:32–1:58:22 · Guest teaching 7/10 The Reality of Efficiency Research and Mixture-of-Experts The host cites external analyses on why most toy efficiency research fails to scale. Yi agrees forcefully, explaining how unoptimized kernels and non-throughput-matched parameter reductions deceive academic reviewers.1:58:22–2:02:20 · Guest teaching 6/10 Open Source vs. Closed Source AI Dynamics Yi rejects romanticized narratives of decentralized open-source development, arguing that grassroots fine-tuning waves rely on distilling proprietary models and gaming leaderboards rather than true frontier training.2:02:20–2:08:23 · Guest teaching 5/10 Personal Productivity, Research Execution, and Writing Habits Yi shares his productivity routine of pre-drafting paper narratives in Overleaf before experiments finish, while the host outlines his own automated newsletter information funnel.2:08:24–2:18:39 · Guest teaching 6/10 Global AI Talent, Research Culture in Singapore, and the US Ecosystem The host and Yi explore the cultural differences between US impact-driven research and publishing-focused Asian academia, as well as global talent drain and national AI policy.1:39–7:58 · Guest disagreement 1/10 Evolution of the AI Research Metagame The host provides a well-informed recap of Yi Tay's career shift from academic NLP to foundational models. Yi explains how the research meta shifted from task-specific architectures to universal foundation models.7:59–13:20 · Guest disagreement 1/10 Foundational Research Breakthroughs at Google Brain The host inquires about the origins of major breakthrough projects like UL2 and Flan. Yi explains the dynamics inside Google Brain between bottom-up holiday tinker projects and top-down org initiatives.13:23–18:30 · Guest disagreement 2/10 Co-Leading PaLM 2 and Scaling Career Opportunities The host presses on the career mechanics of becoming a co-lead on PaLM 2, asking if it was a deliberate strategy. Yi downplays intentionality, attributing it to organic visibility from UL2 and situational luck.18:31–22:26 · Guest disagreement 3/10 Emergent Abilities and the Mirage Paper Controversy The host passionately defends the concept of emergent abilities against the Mirage paper, displaying deep familiarity with the evaluation debate. Yi agrees that emergent capabilities are real while expressing annoyance at academic best paper hype.22:27–36:02 · Guest disagreement 1/10 Mentorship, Collaborations, and Research Taste at Google Brain Yi shares key insights learned from Quoc Le, Jason Wei, and Hyung Won Chung, emphasizing PR, taste, and engineering discipline. The host adds his own framework on picking up what mentors leave behind.36:03–41:09 · Guest disagreement 2/10 Essential Qualities and Technical Habits of AI Researchers Yi highlights the extreme dedication and 4 AM debugging mindset necessary for LLM researchers while cautioning against burnout. The host asks for specific tech stack recommendations, but Yi advocates framework agnosticism.41:10–46:09 · Guest disagreement 2/10 Founding Reka and Leaving Big Tech The host asks why Yi chose to co-found Reka instead of joining established startups like Mistral or Inflection. Yi explains his desire for a genuine co-founding learning experience rather than repeating big-tech dynamics.46:09–59:44 · Guest disagreement 3/10 Compute Infrastructure Challenges and GPU Node Instability Yi details the brutal operational reality of bad GPU nodes killing full cluster runs and risk-sharing issues with compute providers. The host pushes back technically, questioning why frequent checkpointing and standard orchestration don't mitigate the loss.59:44–1:05:29 · Guest disagreement 1/10 Training Reka Flash, Core, and Edge with a Lean Team The host asks how a 3-5 person pre-training team had the confidence to beat frontier baselines. Yi explains de-risking via 4B ablation runs and rapid iterative hill-climbing rather than predictive certainty.1:05:29–1:20:03 · Guest disagreement 2/10 Analyzing Transformer Architectures: Noam Baselines and Encoder-Decoder Designs Yi delivers a detailed technical breakdown of Noam Shazeer's architectural contributions and the intrinsic 2x FLOP parameter efficiency of encoder-decoder models. The host acknowledges learning a major new architectural perspective.1:20:04–1:29:45 · Guest disagreement 3/10 Evaluating Open LLMs, Benchmark Saturation, and Incentive Mismatches Yi and the host critique benchmark saturation and contamination on MMLU and GSM8K, discussing academic publication incentives. The host proposes formulaic, seed-generated parametric evaluations to prevent benchmark cheating.1:29:46–1:35:29 · Guest disagreement 2/10 Multimodal Architecture Paradigms: Early Fusion and Screen Intelligence The conversation covers multimodal early versus late fusion and Adept's screen-first thesis. Yi argues that early fusion will become universal and models must master both natural imagery and interface screens concurrently.1:35:29–1:44:31 · Guest disagreement 3/10 Pushing Past Chinchilla Scaling and Long Context vs. RAG The host challenges the superiority of long context over RAG in production environments due to inference costs. Yi counters that complex reasoning and whole-document synthesis fundamentally break under RAG chunking limitations.1:44:32–1:58:22 · Guest disagreement 4/10 The Reality of Efficiency Research and Mixture-of-Experts The host cites external analyses on why most toy efficiency research fails to scale. Yi agrees forcefully, explaining how unoptimized kernels and non-throughput-matched parameter reductions deceive academic reviewers.1:58:22–2:02:20 · Guest disagreement 5/10 Open Source vs. Closed Source AI Dynamics Yi rejects romanticized narratives of decentralized open-source development, arguing that grassroots fine-tuning waves rely on distilling proprietary models and gaming leaderboards rather than true frontier training.2:02:20–2:08:23 · Guest disagreement 2/10 Personal Productivity, Research Execution, and Writing Habits Yi shares his productivity routine of pre-drafting paper narratives in Overleaf before experiments finish, while the host outlines his own automated newsletter information funnel.2:08:24–2:18:39 · Guest disagreement 3/10 Global AI Talent, Research Culture in Singapore, and the US Ecosystem The host and Yi explore the cultural differences between US impact-driven research and publishing-focused Asian academia, as well as global talent drain and national AI policy.1:39–7:58 · The hosts pushing back 2/10 Evolution of the AI Research Metagame The host provides a well-informed recap of Yi Tay's career shift from academic NLP to foundational models. Yi explains how the research meta shifted from task-specific architectures to universal foundation models.7:59–13:20 · The hosts pushing back 1/10 Foundational Research Breakthroughs at Google Brain The host inquires about the origins of major breakthrough projects like UL2 and Flan. Yi explains the dynamics inside Google Brain between bottom-up holiday tinker projects and top-down org initiatives.13:23–18:30 · The hosts pushing back 3/10 Co-Leading PaLM 2 and Scaling Career Opportunities The host presses on the career mechanics of becoming a co-lead on PaLM 2, asking if it was a deliberate strategy. Yi downplays intentionality, attributing it to organic visibility from UL2 and situational luck.18:31–22:26 · The hosts pushing back 2/10 Emergent Abilities and the Mirage Paper Controversy The host passionately defends the concept of emergent abilities against the Mirage paper, displaying deep familiarity with the evaluation debate. Yi agrees that emergent capabilities are real while expressing annoyance at academic best paper hype.22:27–36:02 · The hosts pushing back 1/10 Mentorship, Collaborations, and Research Taste at Google Brain Yi shares key insights learned from Quoc Le, Jason Wei, and Hyung Won Chung, emphasizing PR, taste, and engineering discipline. The host adds his own framework on picking up what mentors leave behind.36:03–41:09 · The hosts pushing back 2/10 Essential Qualities and Technical Habits of AI Researchers Yi highlights the extreme dedication and 4 AM debugging mindset necessary for LLM researchers while cautioning against burnout. The host asks for specific tech stack recommendations, but Yi advocates framework agnosticism.41:10–46:09 · The hosts pushing back 2/10 Founding Reka and Leaving Big Tech The host asks why Yi chose to co-found Reka instead of joining established startups like Mistral or Inflection. Yi explains his desire for a genuine co-founding learning experience rather than repeating big-tech dynamics.46:09–59:44 · The hosts pushing back 5/10 Compute Infrastructure Challenges and GPU Node Instability Yi details the brutal operational reality of bad GPU nodes killing full cluster runs and risk-sharing issues with compute providers. The host pushes back technically, questioning why frequent checkpointing and standard orchestration don't mitigate the loss.59:44–1:05:29 · The hosts pushing back 2/10 Training Reka Flash, Core, and Edge with a Lean Team The host asks how a 3-5 person pre-training team had the confidence to beat frontier baselines. Yi explains de-risking via 4B ablation runs and rapid iterative hill-climbing rather than predictive certainty.1:05:29–1:20:03 · The hosts pushing back 2/10 Analyzing Transformer Architectures: Noam Baselines and Encoder-Decoder Designs Yi delivers a detailed technical breakdown of Noam Shazeer's architectural contributions and the intrinsic 2x FLOP parameter efficiency of encoder-decoder models. The host acknowledges learning a major new architectural perspective.1:20:04–1:29:45 · The hosts pushing back 2/10 Evaluating Open LLMs, Benchmark Saturation, and Incentive Mismatches Yi and the host critique benchmark saturation and contamination on MMLU and GSM8K, discussing academic publication incentives. The host proposes formulaic, seed-generated parametric evaluations to prevent benchmark cheating.1:29:46–1:35:29 · The hosts pushing back 2/10 Multimodal Architecture Paradigms: Early Fusion and Screen Intelligence The conversation covers multimodal early versus late fusion and Adept's screen-first thesis. Yi argues that early fusion will become universal and models must master both natural imagery and interface screens concurrently.1:35:29–1:44:31 · The hosts pushing back 5/10 Pushing Past Chinchilla Scaling and Long Context vs. RAG The host challenges the superiority of long context over RAG in production environments due to inference costs. Yi counters that complex reasoning and whole-document synthesis fundamentally break under RAG chunking limitations.1:44:32–1:58:22 · The hosts pushing back 3/10 The Reality of Efficiency Research and Mixture-of-Experts The host cites external analyses on why most toy efficiency research fails to scale. Yi agrees forcefully, explaining how unoptimized kernels and non-throughput-matched parameter reductions deceive academic reviewers.1:58:22–2:02:20 · The hosts pushing back 2/10 Open Source vs. Closed Source AI Dynamics Yi rejects romanticized narratives of decentralized open-source development, arguing that grassroots fine-tuning waves rely on distilling proprietary models and gaming leaderboards rather than true frontier training.2:02:20–2:08:23 · The hosts pushing back 2/10 Personal Productivity, Research Execution, and Writing Habits Yi shares his productivity routine of pre-drafting paper narratives in Overleaf before experiments finish, while the host outlines his own automated newsletter information funnel.2:08:24–2:18:39 · The hosts pushing back 3/10 Global AI Talent, Research Culture in Singapore, and the US Ecosystem The host and Yi explore the cultural differences between US impact-driven research and publishing-focused Asian academia, as well as global talent drain and national AI policy.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:57:00 · the hosts 0% · guest 100%1:57:00 · the hosts 0% · guest 100%2:00:00 · the hosts 0% · guest 100%2:00:00 · the hosts 0% · guest 100%2:03:00 · the hosts 0% · guest 100%2:03:00 · the hosts 0% · guest 100%2:06:00 · the hosts 0% · guest 100%2:06:00 · the hosts 0% · guest 100%2:09:00 · the hosts 0% · guest 100%2:09:00 · the hosts 0% · guest 100%2:12:00 · the hosts 0% · guest 100%2:12:00 · the hosts 0% · guest 100%2:15:00 · the hosts 0% · guest 100%2:15:00 · the hosts 0% · guest 100%2:18:00 · the hosts 0% · guest 100%2:18:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:59:54 Dismantling the open-source grassroots narrative

Yi forcefully rejects the romanticized view of bottom-up open source AI progress, pointing out that many developers merely distill GPT-4, rename models, and exploit leaderboards.

Hardest push from the hosts ▶ 56:24 Pushing back on GPU cluster fragility and checkpointing

The host refuses to accept that node failures should be catastrophic, pressing Yi on why file I/O, checkpoint cadences, and standard orchestration haven't resolved the issue.

Biggest teaching moment ▶ 1:13:50 Masterclass on encoder-decoder intrinsic sparsity

Yi breaks down the mathematical FLOP-to-parameter advantage of encoder-decoder architectures, revealing insights that the host admits are completely new and compelling to him.

The host holds their own ▶ 1:45:09 Citing TR Texas on the illusion of toy efficiency papers

The host quotes detailed technical critiques demonstrating that most academic efficiency papers rely on small-scale experiments that fail completely when scaled to production models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Evolution of the AI Research Metagame 6412 The host provides a well-informed recap of Yi Tay's career shift from academic NLP to foundational models. Yi explains how the research meta shifted from task-specific architectures to universal foundation models.
Foundational Research Breakthroughs at Google Brain 5511 The host inquires about the origins of major breakthrough projects like UL2 and Flan. Yi explains the dynamics inside Google Brain between bottom-up holiday tinker projects and top-down org initiatives.
Co-Leading PaLM 2 and Scaling Career Opportunities 5423 The host presses on the career mechanics of becoming a co-lead on PaLM 2, asking if it was a deliberate strategy. Yi downplays intentionality, attributing it to organic visibility from UL2 and situational luck.
Emergent Abilities and the Mirage Paper Controversy 7332 The host passionately defends the concept of emergent abilities against the Mirage paper, displaying deep familiarity with the evaluation debate. Yi agrees that emergent capabilities are real while expressing annoyance at academic best paper hype.
Mentorship, Collaborations, and Research Taste at Google Brain 5611 Yi shares key insights learned from Quoc Le, Jason Wei, and Hyung Won Chung, emphasizing PR, taste, and engineering discipline. The host adds his own framework on picking up what mentors leave behind.
Essential Qualities and Technical Habits of AI Researchers 5622 Yi highlights the extreme dedication and 4 AM debugging mindset necessary for LLM researchers while cautioning against burnout. The host asks for specific tech stack recommendations, but Yi advocates framework agnosticism.
Founding Reka and Leaving Big Tech 5422 The host asks why Yi chose to co-found Reka instead of joining established startups like Mistral or Inflection. Yi explains his desire for a genuine co-founding learning experience rather than repeating big-tech dynamics.
Compute Infrastructure Challenges and GPU Node Instability 7635 Yi details the brutal operational reality of bad GPU nodes killing full cluster runs and risk-sharing issues with compute providers. The host pushes back technically, questioning why frequent checkpointing and standard orchestration don't mitigate the loss.
Training Reka Flash, Core, and Edge with a Lean Team 5512 The host asks how a 3-5 person pre-training team had the confidence to beat frontier baselines. Yi explains de-risking via 4B ablation runs and rapid iterative hill-climbing rather than predictive certainty.
Analyzing Transformer Architectures: Noam Baselines and Encoder-Decoder Designs 6822 Yi delivers a detailed technical breakdown of Noam Shazeer's architectural contributions and the intrinsic 2x FLOP parameter efficiency of encoder-decoder models. The host acknowledges learning a major new architectural perspective.
Evaluating Open LLMs, Benchmark Saturation, and Incentive Mismatches 7632 Yi and the host critique benchmark saturation and contamination on MMLU and GSM8K, discussing academic publication incentives. The host proposes formulaic, seed-generated parametric evaluations to prevent benchmark cheating.
Multimodal Architecture Paradigms: Early Fusion and Screen Intelligence 6622 The conversation covers multimodal early versus late fusion and Adept's screen-first thesis. Yi argues that early fusion will become universal and models must master both natural imagery and interface screens concurrently.
Pushing Past Chinchilla Scaling and Long Context vs. RAG 7635 The host challenges the superiority of long context over RAG in production environments due to inference costs. Yi counters that complex reasoning and whole-document synthesis fundamentally break under RAG chunking limitations.
The Reality of Efficiency Research and Mixture-of-Experts 8743 The host cites external analyses on why most toy efficiency research fails to scale. Yi agrees forcefully, explaining how unoptimized kernels and non-throughput-matched parameter reductions deceive academic reviewers.
Open Source vs. Closed Source AI Dynamics 6652 Yi rejects romanticized narratives of decentralized open-source development, arguing that grassroots fine-tuning waves rely on distilling proprietary models and gaming leaderboards rather than true frontier training.
Personal Productivity, Research Execution, and Writing Habits 6522 Yi shares his productivity routine of pre-drafting paper narratives in Overleaf before experiments finish, while the host outlines his own automated newsletter information funnel.
Global AI Talent, Research Culture in Singapore, and the US Ecosystem 6633 The host and Yi explore the cultural differences between US impact-driven research and publishing-focused Asian academia, as well as global talent drain and national AI policy.

Statements from this episode (34)

Opinion
Tay: Underlying AI research principles have not changed much beyond compute scale
“Fundamentally, I don't think, like, the, like, the stuff has actually, like, the underlying principles of research hasn't really changed that much, except for, like, compute.”
Yi Tay Jul 5, 2024 ▶ 4:52
Opinion
Tay: ChatGPT's release made task-specific academic NLP research obsolete
“The big thing about the ChatGPT moment of, like, twenty-twenty-two, the thing that changed drastically is, like, it completely, like, it was, like, this sharp, like, make all this work, like, kind of, like, obsolete”
Yi Tay Jul 5, 2024 ▶ 6:26
Assertion Not checkable as stated
Tay: Google and OpenAI built general models three years before academia
“Places like Google and Meta, OpenAI, we will be working on things, like, Three years ahead of everybody else, and then suddenly, like, then Academia would be, like, still working on, like, these task-specific things.”
Yi Tay Jul 5, 2024 ▶ 7:07
Assertion Supported
Yi Tay: UL2 is a 20B encoder-decoder model using T5 pre-training data
“So I think UL-II is an encoder, decoder, 20 B model. I think when we got it approved, it was, like kind of, you know, it was released as, like, kind of, like, the big brother of T-Five, you know, kind of like, okay, we updated T-Five with, like, a new objectiv…”
Yi Tay Jul 5, 2024 ▶ 14:54
Assertion Not checkable as stated
Yi Tay: Jason Wei originated the Emergent Abilities paper thesis
“That was mostly Jason's thesis, and I have to really say that, like Jason has really good ideas, and I was more of, like, a support role for that paper, yeah.”
Yi Tay Jul 5, 2024 ▶ 20:05
Opinion
Tay: Academic best paper awards at AI conferences are completely meaningless
“Does best paper awards, like, mean anything? Actually, it doesn't mean anything, right, like, but like, I think that was more of, like, my, Like, where my angst was coming from, right?”
Yi Tay Jul 5, 2024 ▶ 22:03
Assertion Contradicted
Tay: Most first-author papers by Jason Wei reach 1,000 annual citations
“Like, every single, so every single first author paper that, that, like, Jason writes in the last, has like, 1000 citations in one year. Like, no, I mean, not every, but like, most of it that he leads.”
Yi Tay Jul 5, 2024 ▶ 30:34
Insight
Tay: Keyboard shortcuts on one monitor beat multi-monitor head turning
“Keyboard is more optimal than moving your head because If you can switch your screen fast enough, it's faster than your head, like, moving to different screens, and stuff like that.”
Yi Tay Jul 5, 2024 ▶ 35:25
Insight
Tay: Frontier AI researchers cannot maintain standard nine-to-five work-life balance
“You cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, you know, like, eight pm, I'm not going to, my…”
Yi Tay Jul 5, 2024 ▶ 38:49
Disclosure
Tay discloses he does not know how to write CUDA kernels
“Should I know CUDA kernels? I don't know CUDA kernels.”
Yi Tay Jul 5, 2024 ▶ 39:56
Insight
Yi Tay: AI researchers should stay framework-agnostic instead of mastering one stack
“I don't think there's any specific thing. In fact, I will try to be as, like, agnostic, like, I don't, like, I don't really say, like, okay, you need to learn JAX, you need to learn this, by the time you finish learning, there's a new framework out, anyway. So…”
Yi Tay Jul 5, 2024 ▶ 40:34
Opinion
Tay: Inflection AI is completely gone and effectively defunct
“I wouldn't have left, like, for inflection, or something like that. Like, I mean, inflection is gone now. RIP.”
Yi Tay Jul 5, 2024 ▶ 45:07
Disclosure
Tay: Reka relied on 500 A100s during major H100 delivery delays
“For a long period of time, we had, 500 A-One hundreds, because we made a commitment, like and they were constantly being delayed, I think, because of H-One hundred, supply demand, whatever, like, reasons that and it was also very hard to get, like, a lot of co…”
Yi Tay Jul 5, 2024 ▶ 47:12
Insight
Tay: Newly provisioned GPU clusters are highly unreliable and require node filtering
“Usually when, like, a provider, like, provisions new notes, or they would, like give us... Yeah, it's usually, like, bad, like, dog shit, like, at the start. And then it gets, like, better as you go through the process of, like, returning notes, like, and, you…”
Yi Tay Jul 5, 2024 ▶ 51:07
Insight
Tay: A single bad GPU node kills an entire distributed LLM training job
“If there's one bad note, it kills the entire job, right? So, like, the fact of, like, the game became, like, just eliminating bad notes from the thing, right?”
Yi Tay Jul 5, 2024 ▶ 51:36
Insight
Tay: The biggest green flag for GPU providers is sharing node failure costs
“If you do it like a, like, compute startup or anything, the biggest green flag would be to share the cause of node failures with, ah, with, ah, your customers, right”
Yi Tay Jul 5, 2024 ▶ 54:08
Insight
Tay: Complex architecture modifications fail due to an implementation lottery
“A lot of architecture changes, right, the moment they are, like, tedious to implement, like, nobody, like, SuiGuru is a simple thing, right? [4306] Yi Tay: Just split it and then get it. [4307] Yi Tay: It's a very simple thing to implement. [4309] Yi Tay: Mayb…”
Yi Tay Jul 5, 2024 ▶ 1:11:30
Insight
Tay: Encoder-decoders provide 2x parameter capacity at matched FLOPs
“The only big benefit of encoder decoders [4425] Yi Tay: Is that it has this thing called, like, I mean, what I like to call intrinsic sparsity. [4430] Podcast Host: Ok. [4431] Yi Tay: So basically, an encoder decoder with, like, n baramps is, like, basically, …”
Yi Tay Jul 5, 2024 ▶ 1:13:41
Opinion
Yi Tay: Llama 3 shows Meta may have caught up to Google
“So I think I don't really follow, like, fine much, but I think that, like, Lama Tree actually shows that, like, kind of, like, Meta got a pretty, like, a good stack around training these models you know, like, oh, and I've even started to feel like, oh, they a…”
Yi Tay Jul 5, 2024 ▶ 1:21:59
Insight
Tay: Serious AI labs should never release their good evaluation benchmarks
“Serious LMS that create their own evals, and they, a good eval set is one that you don't release. A good eval set is the one that you, like, ok, you release some of it, but, like, it's like, you don't, like, you know, let it be contaminated by the community.”
Yi Tay Jul 5, 2024 ▶ 1:23:05
Opinion
Yi Tay: GSM8K and HumanEval are saturated, contaminated, and uninformative
“I mean, like, you know, the things like GSMK human eval, the coding human eval, they're all, like Contaminated. Like, not, not, I wouldn't say, they're all, like, saturated, contaminated, you know, like, you know, GSMK, whether you're a 92, 91, like, no one ca…”
Yi Tay Jul 5, 2024 ▶ 1:23:30
Prediction Not checkable as stated
Tay: Multimodal AI architectures will eventually move completely to early fusion
“As early fusion models get more traction, I think the themes will start to get more and more, like, it's a bit like how all the tasks like unify, like from Like, two zero one nine to, like, now it's like all the tasks are unifying, now it's like all the modali…”
Yi Tay Jul 5, 2024 ▶ 1:31:55
Prediction Not checkable as stated
Tay: Vision models will unify screen intelligence and natural imagery without bifurcating
“I think at the end of the day, like, the models would become, like, I don't see that there will be, like, screen agents and, like, natural images. Humans, like, you can read what's on a screen, you can go out and appreciate the scenery, right? You're not, like…”
Yi Tay Jul 5, 2024 ▶ 1:34:02
Opinion
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
Yi Tay Jul 5, 2024 ▶ 1:40:05
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Yi Tay Jul 5, 2024 ▶ 1:46:43
Opinion
Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, lik…”
Yi Tay Jul 5, 2024 ▶ 1:51:09
Insight
Tay: Training MoE models from scratch is ideal over sparse upcycling
“I think in the ideal case, you do MOE from scratch.”
Yi Tay Jul 5, 2024 ▶ 1:52:36
Insight
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Yi Tay Jul 5, 2024 ▶ 1:59:19
Insight
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Yi Tay Jul 5, 2024 ▶ 2:01:00
Opinion
Tay: Hugging Face's Open LLM Leaderboard is a major problem
“The open LM leaderboard is, like, probably, like, the, a big, like, Problem, to be honest.”
Yi Tay Jul 5, 2024 ▶ 2:01:39
Insight
Tay: Researchers can rely on the Twitter algorithm to surface important papers
“You actually don't have to follow anything. If the paper is important enough, the Twitter algorithm will give it to you.”
Yi Tay Jul 5, 2024 ▶ 2:06:14
Insight
Yi Tay: AI researchers should draft papers and titles before running experiments
“I usually start writing the thing while working on that thing itself. Like, so, even, like, let's say, like, if you want to launch something, like, then the end goal is, like, a blog post, or shipping something, everything, right? I like, or not really a launc…”
Yi Tay Jul 5, 2024 ▶ 2:06:22
Opinion
Tay: Singapore AI research prioritizes paper counts over real-world impact
“I think, to be honest, the research here is, like, in Singapore is just basically, like, they just care about publishing papers and stuff like that and then it's not, like, impact-driven. I think, at U.S., it's mostly focused on impact-driven, and the thing ne…”
Yi Tay Jul 5, 2024 ▶ 2:11:38
Insight
Yi Tay: Governments cannot artificially manufacture top AI talent ecosystems
“I don't think there's actually much, like, the government can do to, like, influence, like, this kind of thing is, like, a natural, like, organic, natural thing, right? The worst thing to do is probably, like, to create, like, create a lot of artificial things…”
Yi Tay Jul 5, 2024 ▶ 2:15:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.