Oct 23, 2025 · 1h 10m · mad
Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Anthropic researcher Julian Schrittwieser about the exponential growth trajectory of frontier AI, the evolution of reinforcement learning from AlphaGo to LLM agents, and the technical and societal implications of autonomous AI systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25.5% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Julian explicitly rejects the premise put forward by AI pioneer Richard Sutton, arguing that abandoning pre-training in favor of pure from-scratch RL is practically inefficient and bad for safety alignment.
Hardest push from Matt ▶ 7:44 Challenging benchmark performance vs real-world messinessMatt directly challenges the optimism around capability benchmarks by asking how well GDPval or Meter predict actual economic utility once messy data, compliance, liability, and friction are factored in.
Biggest teaching moment ▶ 48:39 Explaining why pre-trained text cannot create effective agentsJulian clearly articulates why raw base models make poor agents, explaining that static pre-training corpora contain static knowledge rather than interactive action loops or self-correction patterns.
Matt holds his own ▶ 20:55 Citing Richard Sutton's recent podcast argument on pure RLMatt displays deep context and industry familiarity by bringing up Richard Sutton's recent argument on Dwarkesh Patel's podcast regarding whether future models should replace pre-training entirely with RL.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| The Viral Blog Post: Failing to Understand the Exponential | 2 | 3 | 1 | 1 | Matt opens by referencing Julian's viral blog post on exponential growth in AI. Julian elaborates on why people struggle to extrapolate exponential progress, drawing parallels to early COVID-19 perception. | |
| Unpacking Predictions for 2026–2027 AI Capabilities | 3 | 2 | 0 | 0 | Julian breaks down baseline extrapolations for AI autonomous task length into 2026 and 2027. Matt accurately summarizes task length and time-to-complete as core evaluation metrics. | |
| Evaluating AI Progress: GDPval vs. Real-World Friction | 5 | 3 | 1 | 4 | Matt offers a sharp pushback question, asking how benchmark metrics account for real-world friction like compliance and messy data. Julian explains how duration benchmarks naturally force models to deal with complexity. | |
| Can AI Be Creative? Move 37 and LLM Novelty | 3 | 4 | 1 | 1 | Matt asks about AlphaGo's Move 37 and whether LLMs can demonstrate true novelty. Julian details Move 37 context and explains that LLM novelty is easy, but making novel outputs useful is the real challenge. | |
| AI in Scientific Discovery and the Timeline to a Nobel Prize | 5 | 3 | 0 | 1 | Matt cites specific DeepMind and biomedical research breakthroughs (AlphaCode, AlphaTensor). Julian outlines a timeline leading to potential Nobel-prize-tier AI scientific discoveries by 2027. | |
| The Singularity Debate: Discontinuity vs. Scientific Scaling | 4 | 5 | 1 | 2 | Matt questions whether self-improving AI leads to a technological singularity or faces diminishing returns. Julian reframes the singularity concept using pharmacology and scientific scaling laws, arguing against a sudden discontinuity. | |
| Is Pre-Training + RL Sufficient to Achieve AGI? | 6 | 4 | 2 | 3 | Matt demonstrates high domain familiarity by citing Richard Sutton's podcast appearance arguing for pure RL from scratch. Julian explicitly rejects Sutton's premise, defending the practical and alignment benefits of pre-training. | |
| Julian Schrittwieser's Personal Journey: From Austria to DeepMind | 1 | 1 | 0 | 0 | Matt invites Julian to share his personal career trajectory. Julian discusses growing up in Austria, building game engines, interning at Google, and deciding to join DeepMind after seeing Demis Hassabis speak. | |
| The Evolution of Reinforcement Learning: AlphaGo to MuZero | 5 | 5 | 0 | 1 | Matt guides the chronology of AlphaGo, AlphaGo Zero, AlphaZero, and MuZero, while clarifying technical concepts like tree search versus database search. Julian explains the step-by-step removal of human priors. | |
| World Models and Translating Game RL to LLM Agents | 4 | 4 | 0 | 1 | Matt asks how game RL translated into LLM agents and modern world models. Julian explains why MuZero learned an environment predictor to handle tasks without explicit simulators. | |
| Implicit World Models in Language Models vs. MuZero | 4 | 4 | 0 | 1 | Matt probes implicit versus explicit world models. Julian explains that language models naturally build implicit world representations to predict future text tokens without reconstructing full high-dimensional state space. | |
| Why Combining Pre-training and RL Took Time | 4 | 5 | 0 | 2 | Matt asks why combining pre-training and RL took so long industry-wide. Julian educates on the engineering difficulty of RL feedback loops compared to the stability of supervised pre-training targets. | |
| Scaling Laws and Compute Allocation in Reinforcement Learning | 4 | 4 | 0 | 1 | Matt asks about compute allocation and scaling laws in RL. Julian notes that RL scales predictably with compute and discusses flexible reward signals like Constitutional AI. | |
| Training Data for Modern RL and Model-Generated Data | 4 | 5 | 0 | 1 | Matt asks whether quality, quantity, or recency matters most for RL data. Julian contrasts the high stability of AlphaZero planning search data with the stability challenges of direct LLM RL sampling. | |
| The Intersection of Reinforcement Learning and Agentic AI | 3 | 5 | 0 | 0 | Matt asks Julian to explain the connection between RL and agentic AI for a general tech audience. Julian highlights how pre-training text lacks action trajectories and error recovery, which RL provides. | |
| Building Agentic Applications: Frontier Models vs. Custom RL | 4 | 4 | 0 | 1 | Matt asks whether app developers should train custom RL models or rely on frontier APIs. Julian advises relying on base frontier models while investing effort into tool harnesses and execution context. | |
| Goodhart's Law and the Challenges of Public AI Benchmarks | 4 | 3 | 0 | 1 | Matt introduces Goodhart's Law and leaderboard gaming. Julian confirms the benchmark optimization phenomenon and recommends held-out internal benchmarks to avoid gaming metrics. | |
| Internal Evaluation Practices and Evaluating Complex Model Capabilities | 3 | 4 | 0 | 1 | Matt inquires about internal evaluation practices at DeepMind and Anthropic. Julian points out the difficulty of building evals that are simultaneously cheap, reliable, and representative of complex work. | |
| Mechanistic Interpretability and the Golden Gate Claude Feature | 5 | 4 | 0 | 1 | Matt asks about mechanistic interpretability and Anthropic's alignment commitments. Julian highlights the Golden Gate Claude feature discovery and explains how interpretability validates model alignment. | |
| Comparative Advantage and the Economic Impact of AI on Jobs | 4 | 4 | 0 | 1 | Matt asks about economic impact on employment and inequality. Julian discusses comparative advantage, noting AI elevates productive capacity rather than executing flat human replacement. |