Jan 23, 2026 · 1h 32m · latent-space

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay

Yi Tay · 48m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this conversation, Google DeepMind Senior Research Scientist Yi Tay breaks down the algorithmic breakthroughs behind Gemini's IMO Gold victory, the mechanics of on-policy reinforcement learning, and the establishment of GDM Singapore. He offers deep insights into Transformer longevity, data efficiency, AI-augmented engineering workflows, and the enduring necessity of algorithmic innovation on the path to AGI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.6 Guest teaching 4.2 Guest disagreement 1.6 The hosts pushing back 2.3
05100:0020:0040:001:00:001:20:000:00–4:53 · The hosts as informed peer 6/10 Emergent Utility in AI Coding and Image Generation The host opens by connecting Yi's previous research on UL2 and T5 to modern RL modeling, while probing into Google DeepMind team structure.4:54–8:00 · The hosts as informed peer 5/10 Contrasting On-Policy Reinforcement Learning with Supervised Imitation Yi breaks down the technical distinction between on-policy reinforcement learning and off-policy supervised fine-tuning using human sports and learning analogies.8:00–10:46 · The hosts as informed peer 7/10 Translating Machine Learning Optimization Principles to Human Learning The host drives the thesis that ML optimization dynamics, specifically learning rate adjustments and Bayesian updates, map directly onto human cognitive shifts.10:47–16:32 · The hosts as informed peer 6/10 Parallel Trajectories, Self-Consistency, and Chain of Thought Reasoning Host and guest discuss the nuances between parallel test-time sampling, chain of thought, and self-consistency mechanisms in frontier models.16:32–20:22 · The hosts as informed peer 6/10 Subsuming Specialized Symbolic Engines into General-Purpose Parameters Yi argues that general connectionist neural parameters will inevitably subsume specialized symbolic engines like Lean or external calculators.20:22–24:10 · The hosts as informed peer 4/10 Multi-Timezone Captaincy and Live Olympiad Verification Dynamics Yi details the collaborative mechanics of co-captaining the IMO run across multiple international timezones while monitoring live human participant benchmarks.24:11–26:23 · The hosts as informed peer 5/10 Reflecting on Five Years of AI Progress and the IMOCAT Run Discussion reflects on the pace of AI advancement over five years and clarifies the origin of the IMOCAT experiment config name.26:23–29:55 · The hosts as informed peer 6/10 Evaluating Long-Horizon Agent Planning Through Pokemon Benchmarks Host and guest evaluate Pokemon as a contamination-resistant long-horizon benchmark testing spatial reasoning, state understanding, and deep research.29:56–32:23 · The hosts as informed peer 6/10 Novel Knowledge Discovery vs. Synthesis in the AI Scientist Paradigm The host challenges whether guide-following is true intelligence versus inventing novel concepts from scratch like the Transformer architecture.32:24–35:10 · The hosts as informed peer 6/10 Demystifying Reasoning: Discrete Chains of Thought vs. Latent Thinking Yi demystifies reasoning into post-training RL elicitation and discusses latent thinking representations versus discrete token generation.35:10–41:02 · The hosts as informed peer 6/10 Web Contamination, Synthetic Reasoning Loops, and Code Generalization The host questions web data contamination with synthetic reasoning tokens, leading into Yi sharing how AI coding tools accelerated his ML workflows.41:02–43:25 · The hosts as informed peer 5/10 AI as an Augmentation Aura and Managing Model Laziness Yi characterizes AI as an augmentation aura that saves time across an engineering organization, while acknowledging occasional model laziness.43:25–48:18 · The hosts as informed peer 6/10 Compounding Datasets and the Collective Nature of AI Scaling Discussion centers on whether Transformer attention remains sufficient for AGI or if fundamental architecture replacements will emerge.48:18–51:41 · The hosts as informed peer 7/10 Scaling Context Lengths, Continual Learning, and Local Minima The host challenges the sequence-to-sequence assumption for massive 200M context windows, prompting Yi to discuss gradient learning paradigms and local minima traps.51:41–54:05 · The hosts as informed peer 5/10 The Sweet Lesson: Research Ideas and the Expanding Closed-Lab Gap Yi proposes the Sweet Lesson to highlight that research ideas and algorithmic innovations remain decisive, widening the gap between closed labs and open-source.54:06–59:37 · The hosts as informed peer 6/10 Memory vs. Compute Constraints in Next-Generation Hardware Host outlines memory bandwidth and networking bottlenecks across hardware generations, while Yi clarifies his focus stays strictly on algorithmic modeling.59:37–1:06:08 · The hosts as informed peer 7/10 Categorizing World Models and Optimizing Compute Allocation Per Token The host synthesizes three paradigms of world models, while Yi reframes learning efficiency into spending higher FLOPs per token.1:06:08–1:13:22 · The hosts as informed peer 6/10 The Economics and Market Demand for Specialized RL Environments Host probes the market valuations of proprietary RL simulation environments before transitioning into DSI and recommendation systems.1:13:22–1:18:23 · The hosts as informed peer 6/10 Comparing Modeling Dynamics in Information Retrieval and Language Tasks Yi candidly describes the modeling dynamics of IR and recommendation systems as counter-intuitive and disconnected from pure ML progress.1:18:27–1:21:19 · The hosts as informed peer 4/10 Launching the GDM Singapore Symposium and Meeting Government Leadership Yi recaps organizing the DeepMind Singapore symposium with Jeff Dean and Quoc Le, including their briefing with Singapore government leadership.1:21:20–1:24:45 · The hosts as informed peer 5/10 Establishing Frontier AI Research in Singapore vs. Silicon Valley Discussion contrasts building frontier research teams in Singapore versus the hyper-saturated monoculture of the San Francisco Bay Area.1:24:46–1:28:57 · The hosts as informed peer 5/10 Hiring Philosophy at GDM Singapore: High-Stat Talent and Research Taste Yi details hiring criteria focused on high general cognitive stats and independent research taste over formal RL specialization.1:28:57–1:31:54 · The hosts as informed peer 4/10 Physical Biohacking, Health Metrics, and Sustaining Research Energy Yi shares his health transformation metrics and how physical conditioning sustains stamina for intense frontier AI research.0:00–4:53 · Guest teaching 3/10 Emergent Utility in AI Coding and Image Generation The host opens by connecting Yi's previous research on UL2 and T5 to modern RL modeling, while probing into Google DeepMind team structure.4:54–8:00 · Guest teaching 6/10 Contrasting On-Policy Reinforcement Learning with Supervised Imitation Yi breaks down the technical distinction between on-policy reinforcement learning and off-policy supervised fine-tuning using human sports and learning analogies.8:00–10:46 · Guest teaching 2/10 Translating Machine Learning Optimization Principles to Human Learning The host drives the thesis that ML optimization dynamics, specifically learning rate adjustments and Bayesian updates, map directly onto human cognitive shifts.10:47–16:32 · Guest teaching 4/10 Parallel Trajectories, Self-Consistency, and Chain of Thought Reasoning Host and guest discuss the nuances between parallel test-time sampling, chain of thought, and self-consistency mechanisms in frontier models.16:32–20:22 · Guest teaching 5/10 Subsuming Specialized Symbolic Engines into General-Purpose Parameters Yi argues that general connectionist neural parameters will inevitably subsume specialized symbolic engines like Lean or external calculators.20:22–24:10 · Guest teaching 6/10 Multi-Timezone Captaincy and Live Olympiad Verification Dynamics Yi details the collaborative mechanics of co-captaining the IMO run across multiple international timezones while monitoring live human participant benchmarks.24:11–26:23 · Guest teaching 4/10 Reflecting on Five Years of AI Progress and the IMOCAT Run Discussion reflects on the pace of AI advancement over five years and clarifies the origin of the IMOCAT experiment config name.26:23–29:55 · Guest teaching 4/10 Evaluating Long-Horizon Agent Planning Through Pokemon Benchmarks Host and guest evaluate Pokemon as a contamination-resistant long-horizon benchmark testing spatial reasoning, state understanding, and deep research.29:56–32:23 · Guest teaching 4/10 Novel Knowledge Discovery vs. Synthesis in the AI Scientist Paradigm The host challenges whether guide-following is true intelligence versus inventing novel concepts from scratch like the Transformer architecture.32:24–35:10 · Guest teaching 4/10 Demystifying Reasoning: Discrete Chains of Thought vs. Latent Thinking Yi demystifies reasoning into post-training RL elicitation and discusses latent thinking representations versus discrete token generation.35:10–41:02 · Guest teaching 3/10 Web Contamination, Synthetic Reasoning Loops, and Code Generalization The host questions web data contamination with synthetic reasoning tokens, leading into Yi sharing how AI coding tools accelerated his ML workflows.41:02–43:25 · Guest teaching 4/10 AI as an Augmentation Aura and Managing Model Laziness Yi characterizes AI as an augmentation aura that saves time across an engineering organization, while acknowledging occasional model laziness.43:25–48:18 · Guest teaching 4/10 Compounding Datasets and the Collective Nature of AI Scaling Discussion centers on whether Transformer attention remains sufficient for AGI or if fundamental architecture replacements will emerge.48:18–51:41 · Guest teaching 4/10 Scaling Context Lengths, Continual Learning, and Local Minima The host challenges the sequence-to-sequence assumption for massive 200M context windows, prompting Yi to discuss gradient learning paradigms and local minima traps.51:41–54:05 · Guest teaching 5/10 The Sweet Lesson: Research Ideas and the Expanding Closed-Lab Gap Yi proposes the Sweet Lesson to highlight that research ideas and algorithmic innovations remain decisive, widening the gap between closed labs and open-source.54:06–59:37 · Guest teaching 3/10 Memory vs. Compute Constraints in Next-Generation Hardware Host outlines memory bandwidth and networking bottlenecks across hardware generations, while Yi clarifies his focus stays strictly on algorithmic modeling.59:37–1:06:08 · Guest teaching 4/10 Categorizing World Models and Optimizing Compute Allocation Per Token The host synthesizes three paradigms of world models, while Yi reframes learning efficiency into spending higher FLOPs per token.1:06:08–1:13:22 · Guest teaching 3/10 The Economics and Market Demand for Specialized RL Environments Host probes the market valuations of proprietary RL simulation environments before transitioning into DSI and recommendation systems.1:13:22–1:18:23 · Guest teaching 6/10 Comparing Modeling Dynamics in Information Retrieval and Language Tasks Yi candidly describes the modeling dynamics of IR and recommendation systems as counter-intuitive and disconnected from pure ML progress.1:18:27–1:21:19 · Guest teaching 5/10 Launching the GDM Singapore Symposium and Meeting Government Leadership Yi recaps organizing the DeepMind Singapore symposium with Jeff Dean and Quoc Le, including their briefing with Singapore government leadership.1:21:20–1:24:45 · Guest teaching 4/10 Establishing Frontier AI Research in Singapore vs. Silicon Valley Discussion contrasts building frontier research teams in Singapore versus the hyper-saturated monoculture of the San Francisco Bay Area.1:24:46–1:28:57 · Guest teaching 5/10 Hiring Philosophy at GDM Singapore: High-Stat Talent and Research Taste Yi details hiring criteria focused on high general cognitive stats and independent research taste over formal RL specialization.1:28:57–1:31:54 · Guest teaching 4/10 Physical Biohacking, Health Metrics, and Sustaining Research Energy Yi shares his health transformation metrics and how physical conditioning sustains stamina for intense frontier AI research.0:00–4:53 · Guest disagreement 1/10 Emergent Utility in AI Coding and Image Generation The host opens by connecting Yi's previous research on UL2 and T5 to modern RL modeling, while probing into Google DeepMind team structure.4:54–8:00 · Guest disagreement 2/10 Contrasting On-Policy Reinforcement Learning with Supervised Imitation Yi breaks down the technical distinction between on-policy reinforcement learning and off-policy supervised fine-tuning using human sports and learning analogies.8:00–10:46 · Guest disagreement 2/10 Translating Machine Learning Optimization Principles to Human Learning The host drives the thesis that ML optimization dynamics, specifically learning rate adjustments and Bayesian updates, map directly onto human cognitive shifts.10:47–16:32 · Guest disagreement 1/10 Parallel Trajectories, Self-Consistency, and Chain of Thought Reasoning Host and guest discuss the nuances between parallel test-time sampling, chain of thought, and self-consistency mechanisms in frontier models.16:32–20:22 · Guest disagreement 2/10 Subsuming Specialized Symbolic Engines into General-Purpose Parameters Yi argues that general connectionist neural parameters will inevitably subsume specialized symbolic engines like Lean or external calculators.20:22–24:10 · Guest disagreement 1/10 Multi-Timezone Captaincy and Live Olympiad Verification Dynamics Yi details the collaborative mechanics of co-captaining the IMO run across multiple international timezones while monitoring live human participant benchmarks.24:11–26:23 · Guest disagreement 1/10 Reflecting on Five Years of AI Progress and the IMOCAT Run Discussion reflects on the pace of AI advancement over five years and clarifies the origin of the IMOCAT experiment config name.26:23–29:55 · Guest disagreement 1/10 Evaluating Long-Horizon Agent Planning Through Pokemon Benchmarks Host and guest evaluate Pokemon as a contamination-resistant long-horizon benchmark testing spatial reasoning, state understanding, and deep research.29:56–32:23 · Guest disagreement 3/10 Novel Knowledge Discovery vs. Synthesis in the AI Scientist Paradigm The host challenges whether guide-following is true intelligence versus inventing novel concepts from scratch like the Transformer architecture.32:24–35:10 · Guest disagreement 2/10 Demystifying Reasoning: Discrete Chains of Thought vs. Latent Thinking Yi demystifies reasoning into post-training RL elicitation and discusses latent thinking representations versus discrete token generation.35:10–41:02 · Guest disagreement 2/10 Web Contamination, Synthetic Reasoning Loops, and Code Generalization The host questions web data contamination with synthetic reasoning tokens, leading into Yi sharing how AI coding tools accelerated his ML workflows.41:02–43:25 · Guest disagreement 1/10 AI as an Augmentation Aura and Managing Model Laziness Yi characterizes AI as an augmentation aura that saves time across an engineering organization, while acknowledging occasional model laziness.43:25–48:18 · Guest disagreement 1/10 Compounding Datasets and the Collective Nature of AI Scaling Discussion centers on whether Transformer attention remains sufficient for AGI or if fundamental architecture replacements will emerge.48:18–51:41 · Guest disagreement 2/10 Scaling Context Lengths, Continual Learning, and Local Minima The host challenges the sequence-to-sequence assumption for massive 200M context windows, prompting Yi to discuss gradient learning paradigms and local minima traps.51:41–54:05 · Guest disagreement 2/10 The Sweet Lesson: Research Ideas and the Expanding Closed-Lab Gap Yi proposes the Sweet Lesson to highlight that research ideas and algorithmic innovations remain decisive, widening the gap between closed labs and open-source.54:06–59:37 · Guest disagreement 2/10 Memory vs. Compute Constraints in Next-Generation Hardware Host outlines memory bandwidth and networking bottlenecks across hardware generations, while Yi clarifies his focus stays strictly on algorithmic modeling.59:37–1:06:08 · Guest disagreement 2/10 Categorizing World Models and Optimizing Compute Allocation Per Token The host synthesizes three paradigms of world models, while Yi reframes learning efficiency into spending higher FLOPs per token.1:06:08–1:13:22 · Guest disagreement 2/10 The Economics and Market Demand for Specialized RL Environments Host probes the market valuations of proprietary RL simulation environments before transitioning into DSI and recommendation systems.1:13:22–1:18:23 · Guest disagreement 3/10 Comparing Modeling Dynamics in Information Retrieval and Language Tasks Yi candidly describes the modeling dynamics of IR and recommendation systems as counter-intuitive and disconnected from pure ML progress.1:18:27–1:21:19 · Guest disagreement 1/10 Launching the GDM Singapore Symposium and Meeting Government Leadership Yi recaps organizing the DeepMind Singapore symposium with Jeff Dean and Quoc Le, including their briefing with Singapore government leadership.1:21:20–1:24:45 · Guest disagreement 1/10 Establishing Frontier AI Research in Singapore vs. Silicon Valley Discussion contrasts building frontier research teams in Singapore versus the hyper-saturated monoculture of the San Francisco Bay Area.1:24:46–1:28:57 · Guest disagreement 1/10 Hiring Philosophy at GDM Singapore: High-Stat Talent and Research Taste Yi details hiring criteria focused on high general cognitive stats and independent research taste over formal RL specialization.1:28:57–1:31:54 · Guest disagreement 2/10 Physical Biohacking, Health Metrics, and Sustaining Research Energy Yi shares his health transformation metrics and how physical conditioning sustains stamina for intense frontier AI research.0:00–4:53 · The hosts pushing back 2/10 Emergent Utility in AI Coding and Image Generation The host opens by connecting Yi's previous research on UL2 and T5 to modern RL modeling, while probing into Google DeepMind team structure.4:54–8:00 · The hosts pushing back 2/10 Contrasting On-Policy Reinforcement Learning with Supervised Imitation Yi breaks down the technical distinction between on-policy reinforcement learning and off-policy supervised fine-tuning using human sports and learning analogies.8:00–10:46 · The hosts pushing back 3/10 Translating Machine Learning Optimization Principles to Human Learning The host drives the thesis that ML optimization dynamics, specifically learning rate adjustments and Bayesian updates, map directly onto human cognitive shifts.10:47–16:32 · The hosts pushing back 2/10 Parallel Trajectories, Self-Consistency, and Chain of Thought Reasoning Host and guest discuss the nuances between parallel test-time sampling, chain of thought, and self-consistency mechanisms in frontier models.16:32–20:22 · The hosts pushing back 3/10 Subsuming Specialized Symbolic Engines into General-Purpose Parameters Yi argues that general connectionist neural parameters will inevitably subsume specialized symbolic engines like Lean or external calculators.20:22–24:10 · The hosts pushing back 1/10 Multi-Timezone Captaincy and Live Olympiad Verification Dynamics Yi details the collaborative mechanics of co-captaining the IMO run across multiple international timezones while monitoring live human participant benchmarks.24:11–26:23 · The hosts pushing back 2/10 Reflecting on Five Years of AI Progress and the IMOCAT Run Discussion reflects on the pace of AI advancement over five years and clarifies the origin of the IMOCAT experiment config name.26:23–29:55 · The hosts pushing back 2/10 Evaluating Long-Horizon Agent Planning Through Pokemon Benchmarks Host and guest evaluate Pokemon as a contamination-resistant long-horizon benchmark testing spatial reasoning, state understanding, and deep research.29:56–32:23 · The hosts pushing back 4/10 Novel Knowledge Discovery vs. Synthesis in the AI Scientist Paradigm The host challenges whether guide-following is true intelligence versus inventing novel concepts from scratch like the Transformer architecture.32:24–35:10 · The hosts pushing back 2/10 Demystifying Reasoning: Discrete Chains of Thought vs. Latent Thinking Yi demystifies reasoning into post-training RL elicitation and discusses latent thinking representations versus discrete token generation.35:10–41:02 · The hosts pushing back 2/10 Web Contamination, Synthetic Reasoning Loops, and Code Generalization The host questions web data contamination with synthetic reasoning tokens, leading into Yi sharing how AI coding tools accelerated his ML workflows.41:02–43:25 · The hosts pushing back 2/10 AI as an Augmentation Aura and Managing Model Laziness Yi characterizes AI as an augmentation aura that saves time across an engineering organization, while acknowledging occasional model laziness.43:25–48:18 · The hosts pushing back 2/10 Compounding Datasets and the Collective Nature of AI Scaling Discussion centers on whether Transformer attention remains sufficient for AGI or if fundamental architecture replacements will emerge.48:18–51:41 · The hosts pushing back 4/10 Scaling Context Lengths, Continual Learning, and Local Minima The host challenges the sequence-to-sequence assumption for massive 200M context windows, prompting Yi to discuss gradient learning paradigms and local minima traps.51:41–54:05 · The hosts pushing back 2/10 The Sweet Lesson: Research Ideas and the Expanding Closed-Lab Gap Yi proposes the Sweet Lesson to highlight that research ideas and algorithmic innovations remain decisive, widening the gap between closed labs and open-source.54:06–59:37 · The hosts pushing back 3/10 Memory vs. Compute Constraints in Next-Generation Hardware Host outlines memory bandwidth and networking bottlenecks across hardware generations, while Yi clarifies his focus stays strictly on algorithmic modeling.59:37–1:06:08 · The hosts pushing back 3/10 Categorizing World Models and Optimizing Compute Allocation Per Token The host synthesizes three paradigms of world models, while Yi reframes learning efficiency into spending higher FLOPs per token.1:06:08–1:13:22 · The hosts pushing back 3/10 The Economics and Market Demand for Specialized RL Environments Host probes the market valuations of proprietary RL simulation environments before transitioning into DSI and recommendation systems.1:13:22–1:18:23 · The hosts pushing back 2/10 Comparing Modeling Dynamics in Information Retrieval and Language Tasks Yi candidly describes the modeling dynamics of IR and recommendation systems as counter-intuitive and disconnected from pure ML progress.1:18:27–1:21:19 · The hosts pushing back 1/10 Launching the GDM Singapore Symposium and Meeting Government Leadership Yi recaps organizing the DeepMind Singapore symposium with Jeff Dean and Quoc Le, including their briefing with Singapore government leadership.1:21:20–1:24:45 · The hosts pushing back 2/10 Establishing Frontier AI Research in Singapore vs. Silicon Valley Discussion contrasts building frontier research teams in Singapore versus the hyper-saturated monoculture of the San Francisco Bay Area.1:24:46–1:28:57 · The hosts pushing back 2/10 Hiring Philosophy at GDM Singapore: High-Stat Talent and Research Taste Yi details hiring criteria focused on high general cognitive stats and independent research taste over formal RL specialization.1:28:57–1:31:54 · The hosts pushing back 2/10 Physical Biohacking, Health Metrics, and Sustaining Research Energy Yi shares his health transformation metrics and how physical conditioning sustains stamina for intense frontier AI research.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:16:00 Ranting on recommendation systems and IR conferences

Yi strongly dismisses the IR and recommender systems subfield, calling its modeling dynamics weird, rude, and fundamentally lagging behind mainstream ML.

Hardest push from the hosts ▶ 48:59 Refusing the sequence-to-sequence paradigm for 200M contexts

The host rejects the guest's assumption that attention models are all we need, arguing that massive context scaling requires continual learning rather than standard seq2seq inference.

Biggest teaching moment ▶ 5:26 Explaining on-policy RL distillation mechanics

Yi gives a precise conceptual breakdown of how on-policy LM reinforcement learning differs fundamentally from off-policy supervised imitation learning.

The host holds their own ▶ 8:30 Host maps learning rate theory to Bayesian belief updating

The host demonstrates deep analytical synthesis by explaining how Bayesian priors fail when single counter-examples require aggressive learning rate adjustments.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Emergent Utility in AI Coding and Image Generation 6312 The host opens by connecting Yi's previous research on UL2 and T5 to modern RL modeling, while probing into Google DeepMind team structure.
Contrasting On-Policy Reinforcement Learning with Supervised Imitation 5622 Yi breaks down the technical distinction between on-policy reinforcement learning and off-policy supervised fine-tuning using human sports and learning analogies.
Translating Machine Learning Optimization Principles to Human Learning 7223 The host drives the thesis that ML optimization dynamics, specifically learning rate adjustments and Bayesian updates, map directly onto human cognitive shifts.
Parallel Trajectories, Self-Consistency, and Chain of Thought Reasoning 6412 Host and guest discuss the nuances between parallel test-time sampling, chain of thought, and self-consistency mechanisms in frontier models.
Subsuming Specialized Symbolic Engines into General-Purpose Parameters 6523 Yi argues that general connectionist neural parameters will inevitably subsume specialized symbolic engines like Lean or external calculators.
Multi-Timezone Captaincy and Live Olympiad Verification Dynamics 4611 Yi details the collaborative mechanics of co-captaining the IMO run across multiple international timezones while monitoring live human participant benchmarks.
Reflecting on Five Years of AI Progress and the IMOCAT Run 5412 Discussion reflects on the pace of AI advancement over five years and clarifies the origin of the IMOCAT experiment config name.
Evaluating Long-Horizon Agent Planning Through Pokemon Benchmarks 6412 Host and guest evaluate Pokemon as a contamination-resistant long-horizon benchmark testing spatial reasoning, state understanding, and deep research.
Novel Knowledge Discovery vs. Synthesis in the AI Scientist Paradigm 6434 The host challenges whether guide-following is true intelligence versus inventing novel concepts from scratch like the Transformer architecture.
Demystifying Reasoning: Discrete Chains of Thought vs. Latent Thinking 6422 Yi demystifies reasoning into post-training RL elicitation and discusses latent thinking representations versus discrete token generation.
Web Contamination, Synthetic Reasoning Loops, and Code Generalization 6322 The host questions web data contamination with synthetic reasoning tokens, leading into Yi sharing how AI coding tools accelerated his ML workflows.
AI as an Augmentation Aura and Managing Model Laziness 5412 Yi characterizes AI as an augmentation aura that saves time across an engineering organization, while acknowledging occasional model laziness.
Compounding Datasets and the Collective Nature of AI Scaling 6412 Discussion centers on whether Transformer attention remains sufficient for AGI or if fundamental architecture replacements will emerge.
Scaling Context Lengths, Continual Learning, and Local Minima 7424 The host challenges the sequence-to-sequence assumption for massive 200M context windows, prompting Yi to discuss gradient learning paradigms and local minima traps.
The Sweet Lesson: Research Ideas and the Expanding Closed-Lab Gap 5522 Yi proposes the Sweet Lesson to highlight that research ideas and algorithmic innovations remain decisive, widening the gap between closed labs and open-source.
Memory vs. Compute Constraints in Next-Generation Hardware 6323 Host outlines memory bandwidth and networking bottlenecks across hardware generations, while Yi clarifies his focus stays strictly on algorithmic modeling.
Categorizing World Models and Optimizing Compute Allocation Per Token 7423 The host synthesizes three paradigms of world models, while Yi reframes learning efficiency into spending higher FLOPs per token.
The Economics and Market Demand for Specialized RL Environments 6323 Host probes the market valuations of proprietary RL simulation environments before transitioning into DSI and recommendation systems.
Comparing Modeling Dynamics in Information Retrieval and Language Tasks 6632 Yi candidly describes the modeling dynamics of IR and recommendation systems as counter-intuitive and disconnected from pure ML progress.
Launching the GDM Singapore Symposium and Meeting Government Leadership 4511 Yi recaps organizing the DeepMind Singapore symposium with Jeff Dean and Quoc Le, including their briefing with Singapore government leadership.
Establishing Frontier AI Research in Singapore vs. Silicon Valley 5412 Discussion contrasts building frontier research teams in Singapore versus the hyper-saturated monoculture of the San Francisco Bay Area.
Hiring Philosophy at GDM Singapore: High-Stat Talent and Research Taste 5512 Yi details hiring criteria focused on high general cognitive stats and independent research taste over formal RL specialization.
Physical Biohacking, Health Metrics, and Sustaining Research Energy 4422 Yi shares his health transformation metrics and how physical conditioning sustains stamina for intense frontier AI research.

Statements from this episode (34)

Disclosure
DeepMind Singapore Explicitly Adds AGI to Job Postings
“I think that, like, one reason why we work on these models is that we want to get to AGI, and, like, this was a right thing that we added AGI to the job posting, yeah. There is no, like, formal name of the team yet, but it's basically the Gemini theme Singapor…”
Yi Tay Jan 23, 2026 ▶ 1:10
Disclosure
Yi Tay: I had almost no RL background before returning to DeepMind
“I spent a lot of my past life, I call it the past art, working on like architectures and pre-training, but I think now I more, I have like transitioned more into RL. I'm not like old school RL, but the games RL and the old school RL, and to be honest, I had al…”
Yi Tay Jan 23, 2026 ▶ 3:30
Opinion
Yi Tay: Reinforcement learning is the primary AI modeling toolset today
“So I think RL is basically the main modeling tool set that we play around with these days.”
Yi Tay Jan 23, 2026 ▶ 4:10
Insight
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”
Yi Tay Jan 23, 2026 ▶ 5:59
Disclosure
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Yi Tay Jan 23, 2026 ▶ 13:38
Assertion Not checkable as stated
Gemini's IMO Model Checkpoint Required Only One Week of Training
“The training process of this IMO model itself was, like, maybe a week or so.”
Yi Tay Jan 23, 2026 ▶ 16:16
Disclosure
DeepMind Shipped Full Gemini IMO Config Only to Select Mathematicians
“So the inference time config was like the one serve to most people is different, but And the full IMO, like, inference config was presented, like, shipped to some mathematicians just because of the inference cost, right? But that was good enough to be a genera…”
Yi Tay Jan 23, 2026 ▶ 19:04
Prediction Not checkable as stated
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Yi Tay Jan 23, 2026 ▶ 19:31
Disclosure
Yi Tay: Four captains across three locations trained DeepMind's IMO model
“So I think there were four captains for the IMO, two from London. Jonathan was from Mountain View. I was from Singapore. So I think four of us basically trained this model together.”
Yi Tay Jan 23, 2026 ▶ 22:08
What-if
Tay: Present AI Milestones Would Have Been Viewed as AGI Five Years Ago
“If you just look at the AI progress now and five years ago, I think people would think that we already reached like AGI.”
Yi Tay Jan 23, 2026 ▶ 24:53
Disclosure
Tay: DeepMind did not optimize Gemini specifically for Pokémon
“There's actually nothing specifically done for Pokemon.”
Yi Tay Jan 23, 2026 ▶ 27:10
Opinion
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Yi Tay Jan 23, 2026 ▶ 31:59
Insight
Yi Tay: 'Reasoning' Technically Just Means Post-Training RL with Thinking Trajectories
“So I think the actual, like, technical definition of reasoning is making models better with thinking and post-training. Ok? Yeah. So basically, like, RL-ing the model to think better.”
Yi Tay Jan 23, 2026 ▶ 33:24
Opinion
Yi Tay: AI model thoughts do not need to resemble human thoughts
“Generally, I'm not really, I don't really believe that model thoughts have to be the same with human thoughts. I'm actually like, generally in ML, I'm more of the school of thought of let the model do whatever it wants.”
Yi Tay Jan 23, 2026 ▶ 34:28
Disclosure
Yi Tay Fixes ML Bugs Automatically Using AI Without Reading Error Traces
“I think AI coding has started to become the point where I run a job, I get a bug. I almost don't look at the bug. I paste it into, like, anti-gravity, and, like, I throw it, that will fix the bug for me. And then I relaunched the job. And, like, beyond, like, …”
Yi Tay Jan 23, 2026 ▶ 37:29
Insight
Yi Tay: AI Tools Act as a Team Productivity Aura Rather Than Replacing Engineers
“These things are not like going to replace one person as it is, but more like a passive aura that buffs everybody.”
Yi Tay Jan 23, 2026 ▶ 41:54
Prediction Not checkable as stated
Yi Tay: AI Model Laziness and Edge Flaws Will Disappear via General Scaling
“I don't think there's anything that to be done to specifically like focus fire. These things is more like general capability improvements. The models just get better over time and then these things will just like go away.”
Yi Tay Jan 23, 2026 ▶ 43:04
Insight
Yi Tay: AI progress is driven by compounding small incremental changes
“I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that, yeah, that push. I think that's accurate. That's true. There's also, it also feels that there's also a lot of like small, like seemingly mi…”
Yi Tay Jan 23, 2026 ▶ 43:54
Prediction Not checkable as stated
Yi Tay: The Architecture That Achieves AGI Will Still Be a Transformer
“It will be a transformer, I think. Like people, it depends on what you call it, but I think unless the paradigm shifts completely, which is, I mean, as a scientist, you cannot like completely say no to like that, this would never happen. But my feeling is that…”
Yi Tay Jan 23, 2026 ▶ 46:26
Insight
Yi Tay: Efficient attention research consistently failed to eliminate self-attention
“There was this whole big era, which I was also involved in this era, where people try to like undermine the attention as much as possible. Like they try to remove it, simplify it, make it efficient, like this whole Like, efficient attention era. At the end of …”
Yi Tay Jan 23, 2026 ▶ 47:28
Insight
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Yi Tay Jan 23, 2026 ▶ 49:07
Opinion
Yi Tay: The AI Industry Is Stuck in a Transformer Local Minimum
“So now we are like in this local minima of like transformers, everything, everything, right? Maybe it's not easy to like get totally out Of this, because also a lot of people's investment optimization have been done. So the things that play well needs to play …”
Yi Tay Jan 23, 2026 ▶ 50:43
Insight
Yi Tay: The 'Bitter Lesson' Is Overapplied; Architectural Ideas Fundamentally Matter
“The bitter lesson gets used too much in, like, too conveniently used around, but actually there's also a little bit of a, not a bit, there's also a sweet lesson where it's like, ideas matter.”
Yi Tay Jan 23, 2026 ▶ 52:19
Opinion
Yi Tay: AI Research Has Not Entered Diminishing Returns on Ideas
“The number of ideas that actually work is not decreasing compared to the last, like, we're not in the era of diminishing returns yet.”
Yi Tay Jan 23, 2026 ▶ 52:53
Opinion
Yi Tay: Gap Between Closed AI Labs and Open-Source Is Increasing
“I think the gap is definitely increasing.”
Yi Tay Jan 23, 2026 ▶ 53:49
Assertion Not checkable as stated
Yi Tay: AI Labs Prioritize Data Efficiency Because the World Lacks Tokens
“I think in general, the, like learning more, like extracting more from varied data points is definitely valuable, but I think that's what related to the fact that we're like running out of Tokens in the world.”
Yi Tay Jan 23, 2026 ▶ 57:07
Opinion
Yi Tay: The term 'world models' is not well-defined in AI
“I don't think about world models that often. I think because world models are just not really well defined in the first place.”
Yi Tay Jan 23, 2026 ▶ 1:04:03
Insight
Yi Tay: Data-Bound AI Must Spend More Compute Per Token to Scale
“So if you are, you come to a point where you are Very data bound, but not compute bound at all. You just find algorithms that spend a lot of compute on every token.”
Yi Tay Jan 23, 2026 ▶ 1:05:09
Opinion
Yi Tay: IR and RecSys research lags significantly behind NeurIPS and ICML
“Also the IR community and the retrieval community is also like always behind the mainstream. And then now it's just probably gotten even more worse because of ILM and stuff. So, okay, I'm getting into Hottick territory, but it's just, like, certain conferences…”
Yi Tay Jan 23, 2026 ▶ 1:17:23
Opinion
Yi Tay: Top Southeast Asian AI talent requires prominent leadership to unlock
“I feel like the talent we can get from the region is really, really good, but it's only It's only because it's us, we can unlock this talent. Otherwise, might join some other place.”
Yi Tay Jan 23, 2026 ▶ 1:23:10
Opinion
Yi Tay: AI research benefits from escaping the Bay Area monoculture
“I do think that to some extent, if you want to do research in, you need a little bit of peace and quiet somewhere, right? So this island may be good for that, but then you can, you're still, like, able to, like, be connected, right?”
Yi Tay Jan 23, 2026 ▶ 1:24:26
Disclosure
DeepMind Singapore is keeping its Gemini RL team small to maximize compute per capita
“We're hiring, like, my team will work on like RL and reasoning for Gemini and Gemini deep thing. I think we care more about like talent density now. So we're not like also like growing that big, this small team first, just because compute per capita is probabl…”
Yi Tay Jan 23, 2026 ▶ 1:24:52
Insight
Yi Tay: ML and RL Knowledge Can Be Learned Easily by Engineers
“ML. ML can be learned easily. Our knowledge can be learned easily.”
Yi Tay Jan 23, 2026 ▶ 1:26:12
Insight
Yi Tay: Independent research taste is a stronger hiring signal than execution
“If somebody comes up with something and then does something that you feel that is very tasteful and it aligns with what Like researchers in the labs, like one, and they come up with that independently, you know that the function that is good, right? Like if yo…”
Yi Tay Jan 23, 2026 ▶ 1:27:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.