Oct 23, 2025 · 1h 10m · mad

Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)

Julian Schrittwieser · 45m spoken Matt Turck · 16m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Anthropic researcher Julian Schrittwieser about the exponential growth trajectory of frontier AI, the evolution of reinforcement learning from AlphaGo to LLM agents, and the technical and societal implications of autonomous AI systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25.5% of the talking time here. How this is scored →

Matt as informed peer 3.9 Guest teaching 3.8 Guest disagreement 0.3 Matt pushing back 1.2
05100:0015:0030:0045:001:00:001:06–4:46 · Matt as informed peer 2/10 The Viral Blog Post: Failing to Understand the Exponential Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception.4:46–7:16 · Matt as informed peer 3/10 Unpacking Predictions for 2026–2027 AI Capabilities Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics.7:16–10:26 · Matt as informed peer 5/10 Evaluating AI Progress: GDPval vs. Real-World Friction Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity.10:26–13:55 · Matt as informed peer 3/10 Can AI Be Creative? Move 37 and LLM Novelty Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge.13:55–16:25 · Matt as informed peer 5/10 AI in Scientific Discovery and the Timeline to a Nobel Prize Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027.16:25–19:07 · Matt as informed peer 4/10 The Singularity Debate: Discontinuity vs. Scientific Scaling Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity.19:07–22:45 · Matt as informed peer 6/10 Is Pre-Training + RL Sufficient to Achieve AGI? Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training.22:45–26:23 · Matt as informed peer 1/10 Julian Schrittwieser's Personal Journey: From Austria to DeepMind Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak.26:23–32:15 · Matt as informed peer 5/10 The Evolution of Reinforcement Learning: AlphaGo to MuZero Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors.32:15–35:22 · Matt as informed peer 4/10 World Models and Translating Game RL to LLM Agents Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators.35:22–39:02 · Matt as informed peer 4/10 Implicit World Models in Language Models vs. MuZero Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space.39:02–41:43 · Matt as informed peer 4/10 Why Combining Pre-training and RL Took Time Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets.41:43–44:36 · Matt as informed peer 4/10 Scaling Laws and Compute Allocation in Reinforcement Learning Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI.44:36–48:02 · Matt as informed peer 4/10 Training Data for Modern RL and Model-Generated Data Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling.48:02–50:51 · Matt as informed peer 3/10 The Intersection of Reinforcement Learning and Agentic AI Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides.50:51–54:00 · Matt as informed peer 4/10 Building Agentic Applications: Frontier Models vs. Custom RL Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context.54:00–56:13 · Matt as informed peer 4/10 Goodhart's Law and the Challenges of Public AI Benchmarks Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics.56:13–58:15 · Matt as informed peer 3/10 Internal Evaluation Practices and Evaluating Complex Model Capabilities Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work.58:15–1:03:48 · Matt as informed peer 5/10 Mechanistic Interpretability and the Golden Gate Claude Feature Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment.1:03:48–1:06:33 · Matt as informed peer 4/10 Comparative Advantage and the Economic Impact of AI on Jobs Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement.1:06–4:46 · Guest teaching 3/10 The Viral Blog Post: Failing to Understand the Exponential Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception.4:46–7:16 · Guest teaching 2/10 Unpacking Predictions for 2026–2027 AI Capabilities Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics.7:16–10:26 · Guest teaching 3/10 Evaluating AI Progress: GDPval vs. Real-World Friction Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity.10:26–13:55 · Guest teaching 4/10 Can AI Be Creative? Move 37 and LLM Novelty Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge.13:55–16:25 · Guest teaching 3/10 AI in Scientific Discovery and the Timeline to a Nobel Prize Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027.16:25–19:07 · Guest teaching 5/10 The Singularity Debate: Discontinuity vs. Scientific Scaling Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity.19:07–22:45 · Guest teaching 4/10 Is Pre-Training + RL Sufficient to Achieve AGI? Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training.22:45–26:23 · Guest teaching 1/10 Julian Schrittwieser's Personal Journey: From Austria to DeepMind Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak.26:23–32:15 · Guest teaching 5/10 The Evolution of Reinforcement Learning: AlphaGo to MuZero Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors.32:15–35:22 · Guest teaching 4/10 World Models and Translating Game RL to LLM Agents Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators.35:22–39:02 · Guest teaching 4/10 Implicit World Models in Language Models vs. MuZero Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space.39:02–41:43 · Guest teaching 5/10 Why Combining Pre-training and RL Took Time Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets.41:43–44:36 · Guest teaching 4/10 Scaling Laws and Compute Allocation in Reinforcement Learning Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI.44:36–48:02 · Guest teaching 5/10 Training Data for Modern RL and Model-Generated Data Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling.48:02–50:51 · Guest teaching 5/10 The Intersection of Reinforcement Learning and Agentic AI Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides.50:51–54:00 · Guest teaching 4/10 Building Agentic Applications: Frontier Models vs. Custom RL Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context.54:00–56:13 · Guest teaching 3/10 Goodhart's Law and the Challenges of Public AI Benchmarks Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics.56:13–58:15 · Guest teaching 4/10 Internal Evaluation Practices and Evaluating Complex Model Capabilities Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work.58:15–1:03:48 · Guest teaching 4/10 Mechanistic Interpretability and the Golden Gate Claude Feature Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment.1:03:48–1:06:33 · Guest teaching 4/10 Comparative Advantage and the Economic Impact of AI on Jobs Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement.1:06–4:46 · Guest disagreement 1/10 The Viral Blog Post: Failing to Understand the Exponential Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception.4:46–7:16 · Guest disagreement 0/10 Unpacking Predictions for 2026–2027 AI Capabilities Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics.7:16–10:26 · Guest disagreement 1/10 Evaluating AI Progress: GDPval vs. Real-World Friction Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity.10:26–13:55 · Guest disagreement 1/10 Can AI Be Creative? Move 37 and LLM Novelty Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge.13:55–16:25 · Guest disagreement 0/10 AI in Scientific Discovery and the Timeline to a Nobel Prize Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027.16:25–19:07 · Guest disagreement 1/10 The Singularity Debate: Discontinuity vs. Scientific Scaling Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity.19:07–22:45 · Guest disagreement 2/10 Is Pre-Training + RL Sufficient to Achieve AGI? Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training.22:45–26:23 · Guest disagreement 0/10 Julian Schrittwieser's Personal Journey: From Austria to DeepMind Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak.26:23–32:15 · Guest disagreement 0/10 The Evolution of Reinforcement Learning: AlphaGo to MuZero Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors.32:15–35:22 · Guest disagreement 0/10 World Models and Translating Game RL to LLM Agents Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators.35:22–39:02 · Guest disagreement 0/10 Implicit World Models in Language Models vs. MuZero Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space.39:02–41:43 · Guest disagreement 0/10 Why Combining Pre-training and RL Took Time Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets.41:43–44:36 · Guest disagreement 0/10 Scaling Laws and Compute Allocation in Reinforcement Learning Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI.44:36–48:02 · Guest disagreement 0/10 Training Data for Modern RL and Model-Generated Data Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling.48:02–50:51 · Guest disagreement 0/10 The Intersection of Reinforcement Learning and Agentic AI Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides.50:51–54:00 · Guest disagreement 0/10 Building Agentic Applications: Frontier Models vs. Custom RL Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context.54:00–56:13 · Guest disagreement 0/10 Goodhart's Law and the Challenges of Public AI Benchmarks Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics.56:13–58:15 · Guest disagreement 0/10 Internal Evaluation Practices and Evaluating Complex Model Capabilities Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work.58:15–1:03:48 · Guest disagreement 0/10 Mechanistic Interpretability and the Golden Gate Claude Feature Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment.1:03:48–1:06:33 · Guest disagreement 0/10 Comparative Advantage and the Economic Impact of AI on Jobs Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement.1:06–4:46 · Matt pushing back 1/10 The Viral Blog Post: Failing to Understand the Exponential Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception.4:46–7:16 · Matt pushing back 0/10 Unpacking Predictions for 2026–2027 AI Capabilities Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics.7:16–10:26 · Matt pushing back 4/10 Evaluating AI Progress: GDPval vs. Real-World Friction Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity.10:26–13:55 · Matt pushing back 1/10 Can AI Be Creative? Move 37 and LLM Novelty Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge.13:55–16:25 · Matt pushing back 1/10 AI in Scientific Discovery and the Timeline to a Nobel Prize Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027.16:25–19:07 · Matt pushing back 2/10 The Singularity Debate: Discontinuity vs. Scientific Scaling Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity.19:07–22:45 · Matt pushing back 3/10 Is Pre-Training + RL Sufficient to Achieve AGI? Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training.22:45–26:23 · Matt pushing back 0/10 Julian Schrittwieser's Personal Journey: From Austria to DeepMind Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak.26:23–32:15 · Matt pushing back 1/10 The Evolution of Reinforcement Learning: AlphaGo to MuZero Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors.32:15–35:22 · Matt pushing back 1/10 World Models and Translating Game RL to LLM Agents Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators.35:22–39:02 · Matt pushing back 1/10 Implicit World Models in Language Models vs. MuZero Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space.39:02–41:43 · Matt pushing back 2/10 Why Combining Pre-training and RL Took Time Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets.41:43–44:36 · Matt pushing back 1/10 Scaling Laws and Compute Allocation in Reinforcement Learning Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI.44:36–48:02 · Matt pushing back 1/10 Training Data for Modern RL and Model-Generated Data Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling.48:02–50:51 · Matt pushing back 0/10 The Intersection of Reinforcement Learning and Agentic AI Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides.50:51–54:00 · Matt pushing back 1/10 Building Agentic Applications: Frontier Models vs. Custom RL Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context.54:00–56:13 · Matt pushing back 1/10 Goodhart's Law and the Challenges of Public AI Benchmarks Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics.56:13–58:15 · Matt pushing back 1/10 Internal Evaluation Practices and Evaluating Complex Model Capabilities Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work.58:15–1:03:48 · Matt pushing back 1/10 Mechanistic Interpretability and the Golden Gate Claude Feature Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment.1:03:48–1:06:33 · Matt pushing back 1/10 Comparative Advantage and the Economic Impact of AI on Jobs Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 28.3% · guest 71.7%0:00 · Matt 28.3% · guest 71.7%3:00 · Matt 5.3% · guest 94.7%3:00 · Matt 5.3% · guest 94.7%6:00 · Matt 28.8% · guest 71.2%6:00 · Matt 28.8% · guest 71.2%9:00 · Matt 32% · guest 68%9:00 · Matt 32% · guest 68%12:00 · Matt 22.3% · guest 77.7%12:00 · Matt 22.3% · guest 77.7%15:00 · Matt 28.2% · guest 71.8%15:00 · Matt 28.2% · guest 71.8%18:00 · Matt 19% · guest 81%18:00 · Matt 19% · guest 81%21:00 · Matt 40.9% · guest 59.1%21:00 · Matt 40.9% · guest 59.1%24:00 · Matt 17.1% · guest 82.9%24:00 · Matt 17.1% · guest 82.9%27:00 · Matt 18.7% · guest 81.3%27:00 · Matt 18.7% · guest 81.3%30:00 · Matt 24.2% · guest 75.8%30:00 · Matt 24.2% · guest 75.8%33:00 · Matt 32.9% · guest 67.1%33:00 · Matt 32.9% · guest 67.1%36:00 · Matt 23.5% · guest 76.5%36:00 · Matt 23.5% · guest 76.5%39:00 · Matt 24.9% · guest 75.1%39:00 · Matt 24.9% · guest 75.1%42:00 · Matt 41% · guest 59%42:00 · Matt 41% · guest 59%45:00 · Matt 5% · guest 95%45:00 · Matt 5% · guest 95%48:00 · Matt 25.5% · guest 74.5%48:00 · Matt 25.5% · guest 74.5%51:00 · Matt 32.7% · guest 67.3%51:00 · Matt 32.7% · guest 67.3%54:00 · Matt 23.4% · guest 76.6%54:00 · Matt 23.4% · guest 76.6%57:00 · Matt 22.4% · guest 77.6%57:00 · Matt 22.4% · guest 77.6%1:00:00 · Matt 54.8% · guest 45.2%1:00:00 · Matt 54.8% · guest 45.2%1:03:00 · Matt 17.4% · guest 82.6%1:03:00 · Matt 17.4% · guest 82.6%1:06:00 · Matt 12.1% · guest 87.9%1:06:00 · Matt 12.1% · guest 87.9%1:09:00 · Matt 49.1% · guest 50.9%1:09:00 · Matt 49.1% · guest 50.9%
Sharpest disagreement ▶ 21:26 Pushing back on Richard Sutton's RL-from-scratch stance

Julian explicitly rejects the premise put forward by AI pioneer Richard Sutton, arguing that abandoning pre-training in favor of pure from-scratch RL is practically inefficient and bad for safety alignment.

Hardest push from Matt ▶ 7:44 Challenging benchmark performance vs real-world messiness

Matt directly challenges the optimism around capability benchmarks by asking how well GDPval or Meter predict actual economic utility once messy data, compliance, liability, and friction are factored in.

Biggest teaching moment ▶ 48:39 Explaining why pre-trained text cannot create effective agents

Julian clearly articulates why raw base models make poor agents, explaining that static pre-training corpora contain static knowledge rather than interactive action loops or self-correction patterns.

Matt holds his own ▶ 20:55 Citing Richard Sutton's recent podcast argument on pure RL

Matt displays deep context and industry familiarity by bringing up Richard Sutton's recent argument on Dwarkesh Patel's podcast regarding whether future models should replace pre-training entirely with RL.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Viral Blog Post: Failing to Understand the Exponential 2311 Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception.
Unpacking Predictions for 2026–2027 AI Capabilities 3200 Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics.
Evaluating AI Progress: GDPval vs. Real-World Friction 5314 Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity.
Can AI Be Creative? Move 37 and LLM Novelty 3411 Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge.
AI in Scientific Discovery and the Timeline to a Nobel Prize 5301 Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027.
The Singularity Debate: Discontinuity vs. Scientific Scaling 4512 Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity.
Is Pre-Training + RL Sufficient to Achieve AGI? 6423 Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training.
Julian Schrittwieser's Personal Journey: From Austria to DeepMind 1100 Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak.
The Evolution of Reinforcement Learning: AlphaGo to MuZero 5501 Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors.
World Models and Translating Game RL to LLM Agents 4401 Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators.
Implicit World Models in Language Models vs. MuZero 4401 Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space.
Why Combining Pre-training and RL Took Time 4502 Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets.
Scaling Laws and Compute Allocation in Reinforcement Learning 4401 Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI.
Training Data for Modern RL and Model-Generated Data 4501 Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling.
The Intersection of Reinforcement Learning and Agentic AI 3500 Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides.
Building Agentic Applications: Frontier Models vs. Custom RL 4401 Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context.
Goodhart's Law and the Challenges of Public AI Benchmarks 4301 Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics.
Internal Evaluation Practices and Evaluating Complex Model Capabilities 3401 Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work.
Mechanistic Interpretability and the Golden Gate Claude Feature 5401 Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment.
Comparative Advantage and the Economic Impact of AI on Jobs 4401 Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement.

Statements from this episode (27)

Assertion Not checkable as stated
AI autonomous task duration doubles every three to four months
“We are seeing this very consistent improvement over many, many years where every say like, you know, three, four months is able to like do a task that is twice as long as before completely on its own.”
Julian Schrittwieser Oct 23, 2025 ▶ 2:41
Prediction Held up
Top AI models will work autonomously for full days within two years
“In a year from now, maybe two years from now, it's the top models are going to be able to work completely on their own for like a whole day or more”
Julian Schrittwieser Oct 23, 2025 ▶ 3:01
Opinion
Valuations for OpenAI, Anthropic, and Google are fairly conservative
“If you look at OpenAI, if you look at Anthropic, if you look at Google, those evaluations, those revenue numbers are actually fairly conservative.”
Julian Schrittwieser Oct 23, 2025 ▶ 3:32
Opinion
Schrittwieser: Wider AI ecosystem may face bubble while frontier labs thrive
“There may simultaneously be like some sort of bubble in, you know, the wider ecosystem, while at the same time, the frontier labs on a very solid trajectory, having a lot of revenue, making a lot of money.”
Julian Schrittwieser Oct 23, 2025 ▶ 4:13
Insight
Schrittwieser: Task duration dictates how much work can be delegated to AI
“The reason I think why task length specifically is interesting is because that's What allows you to delegate more and more work to language models, to agents. Now, even if you have a very clever model, but if it needs feedback or the interaction with you very …”
Julian Schrittwieser Oct 23, 2025 ▶ 5:58
Opinion
Schrittwieser: OpenAI's GDPval is a strong benchmark for economic impact
“I think that GDP is like a super cool evaluation from OpenAI where they collected a lot of, like, you know, real world tasks from real domain experts to make sure that it is actually representative of what you might do in the economy.”
Julian Schrittwieser Oct 23, 2025 ▶ 7:16
Assertion Not checkable as stated
Schrittwieser: AI language models are clearly capable of novel output
“For me, as somebody who has been doing research a long time, I think it's pretty clear that these models can do novel things.”
Julian Schrittwieser Oct 23, 2025 ▶ 12:26
Prediction Not checkable as stated
AI will make Nobel Prize-level scientific discoveries by 2027 or 2028
“I think my guess for that level of capability might be maybe 2027. I think we're probably not going to find out for quite some time afterwards because of the delay in getting prices. But I think by 20, 27, 20, 28, I think extremely likely that the models will …”
Julian Schrittwieser Oct 23, 2025 ▶ 15:35
Prediction Not checkable as stated
A sudden AI singularity or intelligence explosion is extremely unlikely
“Yeah, I think a true discontinuity is extremely unlikely from, you know, obviously AI researchers are already using AI to accelerate themselves. And so what's, what's already happening and like what is likely to continue to happening is that we see like a smoo…”
Julian Schrittwieser Oct 23, 2025 ▶ 17:11
Prediction Not checkable as stated
Schrittwieser: Current AI paradigm likely to achieve human-level performance in productivity tasks
“I think if you're thinking of, oh, we want some kind of system that can perform at roughly human level in basically all tasks that we care about. Productivity wise. Then I think, yeah, it's extremely likely that the current approach, pre-training RL, you know,…”
Julian Schrittwieser Oct 23, 2025 ▶ 19:43
Prediction Not checkable as stated
Future AI models will continue to rely on pre-training data
“Personally, I think that's unlikely. Not, not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch, as we've been able to do in other domains, but more because pre-training on this vast data sets th…”
Julian Schrittwieser Oct 23, 2025 ▶ 21:27
Insight
Pre-training aids AI alignment by implicitly instilling human values
“I definitely think we would keep using pre-training data, not just from an efficiency point of view as well, but also I think there is interesting safety angles, because by pre-training and, you know, all this human knowledge, we're implicitly creating an agen…”
Julian Schrittwieser Oct 23, 2025 ▶ 22:07
What-if
AlphaGo would have probably lost to Lee Sedol if played earlier
“And I think if we had done it a few months earlier, we would have probably lost.”
Julian Schrittwieser Oct 23, 2025 ▶ 29:57
Assertion Not checkable as stated
Schrittwieser: Language models possess an implicit world model
“So I think, yes, I would say that language models have an, not an explicit world model, but they do have an implicit model of the world.”
Julian Schrittwieser Oct 23, 2025 ▶ 35:23
Insight
Schrittwieser: AI pre-training risks over-restricting an agent's exploration search space
“I think the main, you know, the main challenge or the main thing you need to watch out for is that you don't over encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that …”
Julian Schrittwieser Oct 23, 2025 ▶ 38:41
Assertion Partly supported
Reinforcement learning scaling yields returns on compute similar to pre-training
“If you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL, where we can invest exponentially more compute in RL and keep getting benefits.”
Julian Schrittwieser Oct 23, 2025 ▶ 42:01
Prediction Not checkable as stated
Schrittwieser: Reliable reward sources will be key to scaling reinforcement learning
“Figuring out what are the best reward sources, and how do we scale it up, and how do we, you know, get more rewards, more reliable rewards. That will be one of the key ingredients in scaling up RL further.”
Julian Schrittwieser Oct 23, 2025 ▶ 44:21
Assertion Not checkable as stated
AI research still lacks scaling laws for training data quality
“I think we don't have any good scaling laws yet. That tell us the trade off, especially I think because it's very hard to measure what is the quality of a data point, right? Like how good is this example compared to this other example without being able to mea…”
Julian Schrittwieser Oct 23, 2025 ▶ 46:26
Insight
Schrittwieser: Adding model reasoning improves RL training stability and scaling
“One direction of scaling RL and making it more stable is by improving this by, for example, putting more reasoning into your language model to generate much more high quality training data. That can then give us training that is much more stable, and then we c…”
Julian Schrittwieser Oct 23, 2025 ▶ 47:40
Insight
Schrittwieser: Raw pre-trained AI models make poor agents without RL
“Our pre-training data is not very agent-like. If you think of the pre-training data, right, there is like websites and books and, you know, all kinds of recent text that has a lot of information, but it doesn't have a lot of actions. It doesn't really capture …”
Julian Schrittwieser Oct 23, 2025 ▶ 49:17
Insight
AI app developers do not need custom fine-tuning for top models
“I think nowadays, with the capabilities of, like, you know, top, probably cloud models, top OpenAI, GPT models, You don't need to do any fine tuning. You can take the model as is, ride your own tools, your own harness, and benefit from that agentic training. B…”
Julian Schrittwieser Oct 23, 2025 ▶ 51:30
Prediction Not checkable as stated
Schrittwieser: AI progress will remain smooth and incremental without a single bottleneck
“There's probably, yeah, not one individual blocker. And that's why we will continue to see sort of smooth incremental progress over model releases”
Julian Schrittwieser Oct 23, 2025 ▶ 53:07
Insight
Internal benchmarks are the most accurate way to select AI models
“Just make your own internal benchmark that really represents what you care about, and then measure on that. And I think that's likely to be the most objective, most accurate way of measuring.”
Julian Schrittwieser Oct 23, 2025 ▶ 56:03
Insight
Using chain-of-thought as an RL reward destroys model interpretability
“If you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts…”
Julian Schrittwieser Oct 23, 2025 ▶ 58:25
Insight
Schrittwieser: AI safety must span the entire stack, not just RL
“Yeah, I wouldn't view it alignment adjust like an RL problem. I think it sort of, it goes throughout the whole stack. You might, you know, for example, filter the pre-training data in some way. You might, after training, you might have classifiers that, you kn…”
Julian Schrittwieser Oct 23, 2025 ▶ 1:03:10
Insight
Schrittwieser: Distributing AI productivity gains is a political problem, not technological
“I think it's much more like a political, social problem of, like, figuring out how do we actually benefit from all these improvements, and, like, you know, bring the increases in wealth and productivity to everybody, and it's much less a technological problem.…”
Julian Schrittwieser Oct 23, 2025 ▶ 1:06:10
Assertion Not checkable as stated
Scientific advances are currently bottlenecked on applied intelligence
“All of those are basically bottlenecked on how much intelligence we have access to, and how can we apply it?”
Julian Schrittwieser Oct 23, 2025 ▶ 1:09:08
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.