Dec 18, 2025 · 55m · mad

”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI

Sebastien Bourgeau · 36m spoken Matt Turck · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, Google DeepMind Pre-Training Lead Sebastien Bourgeau discusses the technical, architectural, and organizational shifts driving Gemini 3 and frontier AI systems. He explores topics including data-limited regimes, native multimodality, test-time reasoning, vertical hardware integration, and the future of AI-driven research.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 28.5% of the talking time here. How this is scored →

Matt as informed peer 3.3 Guest teaching 4.0 Guest disagreement 0.8 Matt pushing back 1.0
05100:0015:0030:0045:000:56–3:51 · Matt as informed peer 3/10 Deconstructing the 'Secret' Behind Gemini 3 The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering.3:51–6:59 · Matt as informed peer 3/10 Internal Utility and Realistic AI Trajectories The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering.6:59–9:34 · Matt as informed peer 3/10 AI in Research Workflows and Paradigm Continuity Across Labs The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength.9:34–12:09 · Matt as informed peer 4/10 Explorative Research and Post-Transformer Paradigms The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack.12:09–15:35 · Matt as informed peer 2/10 Inside the Pre-Training Lead Role The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams.15:35–18:10 · Matt as informed peer 1/10 Joining DeepMind and Shifting to Real-World Data The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets.18:10–20:29 · Matt as informed peer 4/10 DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism.20:29–23:00 · Matt as informed peer 4/10 Defining 'Research Taste' and Managing Complexity The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability.23:00–25:37 · Matt as informed peer 3/10 Balancing Exploration, Execution, and Product Pressure The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups.25:37–29:01 · Matt as informed peer 3/10 DeepMind's Organizational Structure across Pre- and Post-Training The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size.29:01–32:11 · Matt as informed peer 4/10 Multimodal Computational Costs and Optimization The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations.32:11–35:48 · Matt as informed peer 4/10 Data Mixes, RL Scaling, and Data Shortages The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits.35:48–38:41 · Matt as informed peer 3/10 Human vs. Machine Data Efficiency The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data.38:41–42:28 · Matt as informed peer 5/10 Large Context Windows vs. Retrieval-Augmented Generation The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap.42:28–44:41 · Matt as informed peer 4/10 AI Model Alignment and Harmful Pre-Training Data The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts.44:41–48:19 · Matt as informed peer 3/10 Agentic Workflows, Screen Understanding, and Model Vibes The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling.48:19–51:34 · Matt as informed peer 4/10 Finite Data Regimes and Inference Cost Optimization The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research.51:34–53:35 · Matt as informed peer 3/10 Guidance for Startups Facing Rapidly Advancing Base Models The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers.0:56–3:51 · Guest teaching 4/10 Deconstructing the 'Secret' Behind Gemini 3 The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering.3:51–6:59 · Guest teaching 4/10 Internal Utility and Realistic AI Trajectories The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering.6:59–9:34 · Guest teaching 4/10 AI in Research Workflows and Paradigm Continuity Across Labs The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength.9:34–12:09 · Guest teaching 3/10 Explorative Research and Post-Transformer Paradigms The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack.12:09–15:35 · Guest teaching 2/10 Inside the Pre-Training Lead Role The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams.15:35–18:10 · Guest teaching 3/10 Joining DeepMind and Shifting to Real-World Data The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets.18:10–20:29 · Guest teaching 5/10 DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism.20:29–23:00 · Guest teaching 5/10 Defining 'Research Taste' and Managing Complexity The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability.23:00–25:37 · Guest teaching 3/10 Balancing Exploration, Execution, and Product Pressure The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups.25:37–29:01 · Guest teaching 4/10 DeepMind's Organizational Structure across Pre- and Post-Training The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size.29:01–32:11 · Guest teaching 4/10 Multimodal Computational Costs and Optimization The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations.32:11–35:48 · Guest teaching 4/10 Data Mixes, RL Scaling, and Data Shortages The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits.35:48–38:41 · Guest teaching 5/10 Human vs. Machine Data Efficiency The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data.38:41–42:28 · Guest teaching 5/10 Large Context Windows vs. Retrieval-Augmented Generation The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap.42:28–44:41 · Guest teaching 4/10 AI Model Alignment and Harmful Pre-Training Data The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts.44:41–48:19 · Guest teaching 4/10 Agentic Workflows, Screen Understanding, and Model Vibes The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling.48:19–51:34 · Guest teaching 5/10 Finite Data Regimes and Inference Cost Optimization The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research.51:34–53:35 · Guest teaching 4/10 Guidance for Startups Facing Rapidly Advancing Base Models The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers.0:56–3:51 · Guest disagreement 1/10 Deconstructing the 'Secret' Behind Gemini 3 The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering.3:51–6:59 · Guest disagreement 1/10 Internal Utility and Realistic AI Trajectories The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering.6:59–9:34 · Guest disagreement 1/10 AI in Research Workflows and Paradigm Continuity Across Labs The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength.9:34–12:09 · Guest disagreement 1/10 Explorative Research and Post-Transformer Paradigms The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack.12:09–15:35 · Guest disagreement 0/10 Inside the Pre-Training Lead Role The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams.15:35–18:10 · Guest disagreement 0/10 Joining DeepMind and Shifting to Real-World Data The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets.18:10–20:29 · Guest disagreement 0/10 DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism.20:29–23:00 · Guest disagreement 0/10 Defining 'Research Taste' and Managing Complexity The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability.23:00–25:37 · Guest disagreement 0/10 Balancing Exploration, Execution, and Product Pressure The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups.25:37–29:01 · Guest disagreement 0/10 DeepMind's Organizational Structure across Pre- and Post-Training The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size.29:01–32:11 · Guest disagreement 1/10 Multimodal Computational Costs and Optimization The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations.32:11–35:48 · Guest disagreement 3/10 Data Mixes, RL Scaling, and Data Shortages The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits.35:48–38:41 · Guest disagreement 1/10 Human vs. Machine Data Efficiency The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data.38:41–42:28 · Guest disagreement 1/10 Large Context Windows vs. Retrieval-Augmented Generation The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap.42:28–44:41 · Guest disagreement 3/10 AI Model Alignment and Harmful Pre-Training Data The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts.44:41–48:19 · Guest disagreement 1/10 Agentic Workflows, Screen Understanding, and Model Vibes The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling.48:19–51:34 · Guest disagreement 0/10 Finite Data Regimes and Inference Cost Optimization The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research.51:34–53:35 · Guest disagreement 0/10 Guidance for Startups Facing Rapidly Advancing Base Models The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers.0:56–3:51 · Matt pushing back 1/10 Deconstructing the 'Secret' Behind Gemini 3 The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering.3:51–6:59 · Matt pushing back 1/10 Internal Utility and Realistic AI Trajectories The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering.6:59–9:34 · Matt pushing back 1/10 AI in Research Workflows and Paradigm Continuity Across Labs The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength.9:34–12:09 · Matt pushing back 2/10 Explorative Research and Post-Transformer Paradigms The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack.12:09–15:35 · Matt pushing back 0/10 Inside the Pre-Training Lead Role The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams.15:35–18:10 · Matt pushing back 0/10 Joining DeepMind and Shifting to Real-World Data The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets.18:10–20:29 · Matt pushing back 0/10 DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism.20:29–23:00 · Matt pushing back 1/10 Defining 'Research Taste' and Managing Complexity The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability.23:00–25:37 · Matt pushing back 1/10 Balancing Exploration, Execution, and Product Pressure The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups.25:37–29:01 · Matt pushing back 0/10 DeepMind's Organizational Structure across Pre- and Post-Training The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size.29:01–32:11 · Matt pushing back 2/10 Multimodal Computational Costs and Optimization The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations.32:11–35:48 · Matt pushing back 3/10 Data Mixes, RL Scaling, and Data Shortages The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits.35:48–38:41 · Matt pushing back 1/10 Human vs. Machine Data Efficiency The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data.38:41–42:28 · Matt pushing back 1/10 Large Context Windows vs. Retrieval-Augmented Generation The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap.42:28–44:41 · Matt pushing back 2/10 AI Model Alignment and Harmful Pre-Training Data The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts.44:41–48:19 · Matt pushing back 1/10 Agentic Workflows, Screen Understanding, and Model Vibes The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling.48:19–51:34 · Matt pushing back 0/10 Finite Data Regimes and Inference Cost Optimization The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research.51:34–53:35 · Matt pushing back 1/10 Guidance for Startups Facing Rapidly Advancing Base Models The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 44.5% · guest 55.5%0:00 · Matt 44.5% · guest 55.5%3:00 · Matt 34.8% · guest 65.2%3:00 · Matt 34.8% · guest 65.2%6:00 · Matt 43.8% · guest 56.2%6:00 · Matt 43.8% · guest 56.2%9:00 · Matt 27.3% · guest 72.7%9:00 · Matt 27.3% · guest 72.7%12:00 · Matt 28.2% · guest 71.8%12:00 · Matt 28.2% · guest 71.8%15:00 · Matt 6.9% · guest 93.1%15:00 · Matt 6.9% · guest 93.1%18:00 · Matt 8.2% · guest 91.8%18:00 · Matt 8.2% · guest 91.8%21:00 · Matt 15.7% · guest 84.3%21:00 · Matt 15.7% · guest 84.3%24:00 · Matt 47.9% · guest 52.1%24:00 · Matt 47.9% · guest 52.1%27:00 · Matt 17.1% · guest 82.9%27:00 · Matt 17.1% · guest 82.9%30:00 · Matt 43.4% · guest 56.6%30:00 · Matt 43.4% · guest 56.6%33:00 · Matt 42.4% · guest 57.6%33:00 · Matt 42.4% · guest 57.6%36:00 · Matt 23.5% · guest 76.5%36:00 · Matt 23.5% · guest 76.5%39:00 · Matt 13.3% · guest 86.7%39:00 · Matt 13.3% · guest 86.7%42:00 · Matt 34.2% · guest 65.8%42:00 · Matt 34.2% · guest 65.8%45:00 · Matt 30.4% · guest 69.6%45:00 · Matt 30.4% · guest 69.6%48:00 · Matt 16.7% · guest 83.3%48:00 · Matt 16.7% · guest 83.3%51:00 · Matt 30.1% · guest 69.9%51:00 · Matt 30.1% · guest 69.9%54:00 · Matt 46.5% · guest 53.5%54:00 · Matt 46.5% · guest 53.5%
Sharpest disagreement ▶ 35:28 Deflection on proprietary reasoning traces

The guest directly shuts down the host's probe into training on reasoning traces with a firm refusal to comment on specifics.

Hardest push from Matt ▶ 35:32 Host presses after guest deflection

Immediately after the guest refuses to comment on reasoning traces, the host cheekily pushes back, noting that his refusal confirms he asked the right question before reframing.

Biggest teaching moment ▶ 36:16 Correcting premise on finite data vs less data

The guest explicitly interrupts the host's premise to clarify that shifting to a finite data regime is not the same as learning with less data.

Matt holds his own ▶ 39:19 Connecting Retro paper to Gemini long context

The host demonstrates deep familiarity with the guest's academic work by contrasting his 2021 Retro paper on retrieval with Gemini 3's context expansion.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Deconstructing the 'Secret' Behind Gemini 3 3411 The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering.
Internal Utility and Realistic AI Trajectories 3411 The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering.
AI in Research Workflows and Paradigm Continuity Across Labs 3411 The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength.
Explorative Research and Post-Transformer Paradigms 4312 The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack.
Inside the Pre-Training Lead Role 2200 The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams.
Joining DeepMind and Shifting to Real-World Data 1300 The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets.
DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO 4500 The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism.
Defining 'Research Taste' and Managing Complexity 4501 The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability.
Balancing Exploration, Execution, and Product Pressure 3301 The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups.
DeepMind's Organizational Structure across Pre- and Post-Training 3400 The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size.
Multimodal Computational Costs and Optimization 4412 The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations.
Data Mixes, RL Scaling, and Data Shortages 4433 The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits.
Human vs. Machine Data Efficiency 3511 The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data.
Large Context Windows vs. Retrieval-Augmented Generation 5511 The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap.
AI Model Alignment and Harmful Pre-Training Data 4432 The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts.
Agentic Workflows, Screen Understanding, and Model Vibes 3411 The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling.
Finite Data Regimes and Inference Cost Optimization 4500 The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research.
Guidance for Startups Facing Rapidly Advancing Base Models 3401 The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers.

Statements from this episode (35)

Insight
Bourgeau: Gemini 3 progress came from small contributions, not one breakthrough
“In my experience, there's maybe one or two of those things that make a larger difference than other things, but it's really a combination of many, many changes and many, many things from a very large team that actually makes Gemini three so much better than th…”
Sebastien Bourgeau Dec 18, 2025 ▶ 1:45
Assertion Not checkable as stated
Bourgeau: AI progress from pre-training improvements is not slowing down
“It's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down.”
Sebastien Bourgeau Dec 18, 2025 ▶ 2:26
Insight
Bourgeau: Frontier AI development is about building systems, not just neural networks
“We're not really building a model anymore. I think we're really building a system at this point. People have sometimes this view that we're just training a neural network architecture and that's it. But it's really the entire system around the network as well …”
Sebastien Bourgeau Dec 18, 2025 ▶ 2:46
Assertion Not checkable as stated
Bourgeau: Gemini answers computer science benchmark questions taking humans significant time
“They are becoming increasingly difficult, and even for me, who has a background in computer science, some of the questions the model answers, it would take me a significant amount of time to answer.”
Sebastien Bourgeau Dec 18, 2025 ▶ 3:41
Disclosure
DeepMind's Bourgeau: AI progress is ahead of where I expected
“I think, if I'm being honest with myself, I think we're ahead of where I thought we could go.”
Sebastien Bourgeau Dec 18, 2025 ▶ 5:00
What-if
Bourgeau: Would not have bet heavily on scaling laws materializing
“I, I'm not sure if I would have bet a lot on, on that actually materializing and being where we are today.”
Sebastien Bourgeau Dec 18, 2025 ▶ 5:26
Prediction Not checkable as stated
DeepMind's Bourgeau predicts major AI-driven scientific breakthroughs within years
“I think we will be able to make some large scientific discoveries in the next few years.”
Sebastien Bourgeau Dec 18, 2025 ▶ 6:14
Prediction Not checkable as stated
Bourgeau: Agentic workflows will accelerate AI research tasks in the next year
“The first part, I think, especially in the next year with more agentic workflows being enabled more and more, that should be able to really accelerate our work there.”
Sebastien Bourgeau Dec 18, 2025 ▶ 7:35
Assertion Not checkable as stated
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Sebastien Bourgeau Dec 18, 2025 ▶ 10:38
Insight
Bourgeau: AI scale has blurred the line between research and engineering
“I think over time that boundary has blurred quite a lot because we're working on these very large systems now. Research really looks like engineering and vice versa.”
Sebastien Bourgeau Dec 18, 2025 ▶ 11:28
Disclosure
Bourgeau: DeepMind shifted focus from pure research to research engineering
“And I think that's a mindset that has really evolved over the last few years at DeepMind, especially where maybe there was a bit more of the traditional research mindset before, and now with Gemini, it's really more about research engineering.”
Sebastien Bourgeau Dec 18, 2025 ▶ 11:38
Disclosure
Bourgeau works with 150 to 200 people on Gemini pre-training
“So it's a fairly large team at this point. It's a bit hard to quantify exactly, but maybe a 152 hundred people I work on a day-to-day on the pre-training side between data, model, infrastructure, evals, and so coordinating the work of all of these people into …”
Sebastien Bourgeau Dec 18, 2025 ▶ 12:52
Assertion Not checkable as stated
Bourgeau: Early DeepMind research defaulted to synthetic data over real-world data
“And at the time we had to add this from real world data to the name of the project, because people would assume otherwise it would be synthetic environments or synthetic data. And that definitely has shifted completely since then.”
Sebastien Bourgeau Dec 18, 2025 ▶ 17:34
Insight
Bourgeau: Trading peak model performance for lower complexity enables faster progress
“Oftentimes we don't necessarily want to use the best performance version of a research idea, but we'd rather trade off some of the performance for a slightly lower complexity version because we think that will allow us to do more and more progress in the futur…”
Sebastien Bourgeau Dec 18, 2025 ▶ 21:35
Insight
Bourgeau: In deep learning, negative results often mean unoptimized techniques
“Especially in deep learning, a negative results doesn't mean something doesn't work. It means you haven't made it work yet often.”
Sebastien Bourgeau Dec 18, 2025 ▶ 22:48
Disclosure
Bourgeau: Google leadership's research background shields DeepMind from benchmark pressure
“There's actually very little of that. I think because all of the leadership has a research background that they're very much aware that yes, to some extent you can force and accelerate specific benchmarks and certain goals, but in the end, the progress and the…”
Sebastien Bourgeau Dec 18, 2025 ▶ 25:12
Disclosure
Bourgeau: Google DeepMind separates AI development into dedicated pre- and post-training teams
“At a super high level so we have a pre-training team, a post-training team. On the pre-training side, we have people working on the model, on the data, the infrastructure, evals as well, very important.”
Sebastien Bourgeau Dec 18, 2025 ▶ 25:50
Assertion Not checkable as stated
Bourgeau: Gemini 3's architecture hasn't changed much from Gemini 2.5
“At the high level, I don't think the architecture has changed that much compared to the previous one. It's more of what I was saying before, where a few different things come together to gather, give a large, large improvement.”
Sebastien Bourgeau Dec 18, 2025 ▶ 26:52
Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Sebastien Bourgeau Dec 18, 2025 ▶ 31:22
Insight
Bourgeau: Pre-training scaling lessons apply directly to RL scaling
“On the RL and RL scaling side, I think we're seeing a lot of the same things we're seeing in pre-training or we saw in pre-training. What's interesting here is because we have the experience of pre-training, a lot of the lessons apply, and we can reapply some …”
Sebastien Bourgeau Dec 18, 2025 ▶ 32:24
Assertion Not checkable as stated
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Sebastien Bourgeau Dec 18, 2025 ▶ 34:15
Insight
Bourgeau: AI research is shifting to a data-limited paradigm
“I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime where, where data would scale as much as you would like. And we're kind of shifting more to a data limited regime, which ac…”
Sebastien Bourgeau Dec 18, 2025 ▶ 34:26
Prediction Not checkable as stated
Bourgeau expects significant long-context AI innovations in the next year
“I think there's going to be a lot more innovation on that side in the next year or so to make long context more efficient, but also just to extend the context length of models themselves.”
Sebastien Bourgeau Dec 18, 2025 ▶ 37:45
Disclosure
DeepMind made recent attention discoveries that will shape near-term research
“For us, at least on the attention side, we've made some really interesting discoveries recently that I think will shape a lot of the research we do in the next few months, and I'm personally very excited about that.”
Sebastien Bourgeau Dec 18, 2025 ▶ 38:06
Prediction Not checkable as stated
Bourgeau: End-to-end differentiable retrieval and search in training will take years
“I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, Learn to retrieve as part of the training and learn how to do search as…”
Sebastien Bourgeau Dec 18, 2025 ▶ 40:09
Assertion Not checkable as stated
Bourgeau: External AI benchmarks quickly become contaminated via web data
“What we found is that external benchmarks, then you can use them for a little while, but very quickly they become contaminated. So they start to be replicated on different forms Different forms or different parts of the web, and then if we end up training on t…”
Sebastien Bourgeau Dec 18, 2025 ▶ 41:54
Insight
Bourgeau: Internal held-out evals are the only way to prevent benchmark self-deception
“The only way you really have to protect against cheating yourself and thinking you're doing better than you are is by actually creating held out evals and not really keeping them held out.”
Sebastien Bourgeau Dec 18, 2025 ▶ 42:18
Insight
Bourgeau: AI models must be trained on harmful data to avoid it
“So at a fundamental level, you did, you do need the model to know about those things. So you have to train a bit at least on those so that it knows what those things are and knows to stay away from those, right?”
Sebastien Bourgeau Dec 18, 2025 ▶ 43:14
Assertion Supported
Bourgeau: Thinking models compute across sequence length to test hypotheses before answering
“Rather than just doing compute in, in the depths or in, in the model side, you also do compute and allow the model to think more on, on the sequence length side of things. So the model actually starts to form hypotheses, test hypotheses, invoke some tools to v…”
Sebastien Bourgeau Dec 18, 2025 ▶ 44:07
Insight
Bourgeau: Pre-training perception and screen understanding is critical for agentic AI
“Bringing it back to the topics of pre-training I think that the perception and vision side is very important for this, because now you're asking models to interact with computer screens. So, so being able to do screen understanding really, really well, Is, is …”
Sebastien Bourgeau Dec 18, 2025 ▶ 45:05
Insight
Bourgeau: Vibe coding performance stems mostly from RL scaling and post-training
“I think this is, yeah, this is in general for vibe coding specifically, I think that's maybe more of an RL scaling and post training thing where, where you can actually get quite a lot of data and train them all to do that really well.”
Sebastien Bourgeau Dec 18, 2025 ▶ 46:15
Insight
Bourgeau: Recent continual learning progress has mostly occurred via post-training search tools
“First, I think a lot of progress has been made on this front since in the last few years. I think this is mostly around post-training, around search, use search tools and then make search calls, then they would have access to that new information.”
Sebastien Bourgeau Dec 18, 2025 ▶ 47:20
Insight
Bourgeau: Full-stack hardware and TPU understanding is a superpower for AI research
“So being able to understand how the stack works all the way down from TPUs to research is kind of a superpower, because then you're able to kind of find these gaps in between different layers that other people weren't necessarily able to see, but also to reaso…”
Sebastien Bourgeau Dec 18, 2025 ▶ 50:05
Prediction Not checkable as stated
Bourgeau: Retrieval-augmented pre-training could become viable in a few years
“I just think it's not unreasonable to think in the next few years, something like that might actually become viable for a leading model like general.”
Sebastien Bourgeau Dec 18, 2025 ▶ 50:54
Insight
Bourgeau: Focus shifts to AI model harnesses and error recovery mechanisms
“And then, so, so what that means is research in terms of how, how you use models and the harness, et cetera, is becoming increasingly important and also how you make models and these harnesses more robust to making errors and recover from such errors.”
Sebastien Bourgeau Dec 18, 2025 ▶ 52:21
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.