Dec 18, 2025 · 55m · mad
”We’re Ahead of Where I Thought We’d Be” — Gemini 3 & the Future of AI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, Google DeepMind Pre-Training Lead Sebastien Bourgeau discusses the technical, architectural, and organizational shifts driving Gemini 3 and frontier AI systems. He explores topics including data-limited regimes, native multimodality, test-time reasoning, vertical hardware integration, and the future of AI-driven research.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 28.5% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
The guest directly shuts down the host's probe into training on reasoning traces with a firm refusal to comment on specifics.
Hardest push from Matt ▶ 35:32 Host presses after guest deflectionImmediately after the guest refuses to comment on reasoning traces, the host cheekily pushes back, noting that his refusal confirms he asked the right question before reframing.
Biggest teaching moment ▶ 36:16 Correcting premise on finite data vs less dataThe guest explicitly interrupts the host's premise to clarify that shifting to a finite data regime is not the same as learning with less data.
Matt holds his own ▶ 39:19 Connecting Retro paper to Gemini long contextThe host demonstrates deep familiarity with the guest's academic work by contrasting his 2021 Retro paper on retrieval with Gemini 3's context expansion.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Deconstructing the 'Secret' Behind Gemini 3 | 3 | 4 | 1 | 1 | The host cites Oriel Vinyals' tweet about pre-training and post-training simplicity. The guest reframes this by explaining that Gemini 3's leap comes from thousands of combined micro-improvements across a massive team, transitioning the view from single neural nets to system engineering. | |
| Internal Utility and Realistic AI Trajectories | 3 | 4 | 1 | 1 | The host brings up benchmark overfitting fears. The guest counters with internal productivity metrics showing researchers increasingly rely on newer model generations to do their daily engineering. | |
| AI in Research Workflows and Paradigm Continuity Across Labs | 3 | 4 | 1 | 1 | The host references AI 2027 automation scenarios and rival releases like GPT-5.2. The guest breaks down research into infra execution versus high-level hypothesis formation, noting lab specialization branches like DeepMind's vision strength. | |
| Explorative Research and Post-Transformer Paradigms | 4 | 3 | 1 | 2 | The host probes for secret post-Transformer architecture research groups and quotes Demis Hassabis on full-stack integration. The guest acknowledges exploratory research exists but highlights how high failure rates require balancing research risk with Google's infra stack. | |
| Inside the Pre-Training Lead Role | 2 | 2 | 0 | 0 | The host asks about the pre-training lead responsibilities and guest background. The guest details managing 150-200 researchers across data, infra, and model teams. | |
| Joining DeepMind and Shifting to Real-World Data | 1 | 3 | 0 | 0 | The guest recounts joining DeepMind via a Cambridge referral and transitioning from synthetic Atari RL to real-world language datasets. | |
| DeepMind's LLM Milestones: Gopher, Chinchilla, and RETRO | 4 | 5 | 0 | 0 | The host specifically prompts the Retro paper. The guest details early scaling work on Gopher, Chinchilla's revision of OpenAI's compute-optimal scaling laws, and Retro's retrieval mechanism. | |
| Defining 'Research Taste' and Managing Complexity | 4 | 5 | 0 | 1 | The host catches the phrase 'research taste' and asks for a definition. The guest defines it as managing complexity budgets, trading maximum raw performance for system simplicity, and team interoperability. | |
| Balancing Exploration, Execution, and Product Pressure | 3 | 3 | 0 | 1 | The host asks about short-term vs long-term pressure and competing for benchmark wins like IMO. The guest explains how critical paths are de-risked before scale-ups. | |
| DeepMind's Organizational Structure across Pre- and Post-Training | 3 | 4 | 0 | 0 | The host asks about org structure and MoE architecture. The guest explains how Mixture of Experts decouples compute usage from total model parameter size. | |
| Multimodal Computational Costs and Optimization | 4 | 4 | 1 | 2 | The host asks if multimodality inflates token costs and whether pre-training scaling laws are dead. The guest dismisses death-of-scaling narratives as strange, explaining how scale compounds with architectural and data innovations. | |
| Data Mixes, RL Scaling, and Data Shortages | 4 | 4 | 3 | 3 | The host asks directly about training on reasoning traces. The guest explicitly declines to comment on proprietary techniques, leading the host to banter about hitting sensitive topics before shifting to data limits. | |
| Human vs. Machine Data Efficiency | 3 | 5 | 1 | 1 | The host asks if models can learn like children with less data. The guest explicitly corrects the host, clarifying that moving to a finite data regime is conceptually different from training with less data. | |
| Large Context Windows vs. Retrieval-Augmented Generation | 5 | 5 | 1 | 1 | The host connects the guest's earlier Retro paper on retrieval to Gemini 3's massive context windows. The guest explains the long-term vision for end-to-end differentiable retrieval and details the complex pre-training evaluation gap. | |
| AI Model Alignment and Harmful Pre-Training Data | 4 | 4 | 3 | 2 | The host asks if toxic web data should be filtered out during pre-training, then asks about DeepThink internals. The guest declines to give DeepThink details while educating on why models need exposure to bad data to recognize unsafe concepts. | |
| Agentic Workflows, Screen Understanding, and Model Vibes | 3 | 4 | 1 | 1 | The host asks about agentic workflows, Google Anti-Gravity, and 'vibe coding'. The guest highlights screen understanding in pre-training and attributes large model feel to pre-training and RL scaling. | |
| Finite Data Regimes and Inference Cost Optimization | 4 | 5 | 0 | 0 | The host brings up NeurIPS themes like continual learning and asks for career advice for students. The guest advises mastering the complete stack from TPU hardware to model research. | |
| Guidance for Startups Facing Rapidly Advancing Base Models | 3 | 4 | 0 | 1 | The host voices VC and startup founder concerns about rapidly expanding base models. The guest advises founders to extrapolate model capability trajectories rather than build narrow wrappers. |