Oct 2, 2025 · 1h 10m · mad

Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)

Sholto Douglas · 52m spoken Matt Turck · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Anthropic AI researcher Sholto Douglas about the release of Claude Sonnet 4.5, frontier model scaling paradigms, and why claims of an AI plateau are premature. Douglas shares insights on reinforcement learning, 30-hour autonomous coding agents, and his personal journey from competitive fencing to leading infrastructure research.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 16.8% of the talking time here. How this is scored →

Matt as informed peer 3.9 Guest teaching 4.0 Guest disagreement 0.6 Matt pushing back 0.9
05100:0015:0030:0045:001:00:001:08–4:14 · Matt as informed peer 3/10 Rapid Model Releases and Hardware Lead Times Matt asks about Anthropic's rapid release pacing and model hierarchy. Sholto explains the two-paradigm regime of pre-training plus RL scaling and compute supply lead times.4:14–6:38 · Matt as informed peer 1/10 Growing Up in Australia and Athletic Mentorship Matt asks about Sholto's background in Australia. Sholto describes his mother's medical mentorship and his elite international fencing trajectory.6:38–9:27 · Matt as informed peer 4/10 Fencing as Reinforcement Learning and Discovering AI Matt perceptively framing YouTube-based athletic analysis as an early form of reinforcement learning. Sholto agrees and outlines reading Gwern's scaling essay and starting independent robotics research.9:27–12:39 · Matt as informed peer 4/10 Academic Obstacles versus Independent High-Signal Work Matt probes the difference between academic gatekeeping and AI research effectiveness. Sholto explains how independent artifacts like Simon Boehm's CUDA guide provide higher signal than traditional PhD credentials.12:39–15:20 · Matt as informed peer 2/10 Navigating Gemini's Formation and Building Inference Stacks Sholto details joining Google before ChatGPT, building Gemini's inference stack from scratch to save hundreds of millions, and navigating multi-team sociopolitical dynamics.15:20–18:40 · Matt as informed peer 3/10 Transition to Anthropic and Shared Alignment Focus Matt asks how top AI labs differ given similar compute resources. Sholto contrasts DeepMind's scientific focus with Anthropic's laser focus on alignment and near-term economic impact through coding.18:40–21:25 · Matt as informed peer 6/10 Mechanistic Understanding and Sutton's Bitter Lesson Matt demonstrates explicit AI research knowledge by interjecting to define Richard Sutton's Bitter Lesson essay. Sholto explains how researcher taste acts as a simplicity regularizer in single training runs.21:25–24:45 · Matt as informed peer 2/10 Convolutional Networks, Vision Transformers, and Language Structure Sholto provides concrete architectural examples of Sutton's Bitter Lesson, explaining why CNN priors and explicit grammar structures get washed away by general vision transformers and language scale.24:45–27:23 · Matt as informed peer 4/10 Balancing Short-Term Wins with Long-Term Techniques Matt asks about research failure rates and managing expensive compute budgets. Sholto quotes Noam Shazeer's 10 percent idea hit rate and discusses balancing short-term tweaks with fundamental long-term research.27:23–31:46 · Matt as informed peer 4/10 Anthropic's Targeted AGI Bet versus DeepMind's Broad Scope Matt asks whether Anthropic researchers investigate non-transformer or non-RL paradigms. Sholto explains Anthropic's focused AGI bet versus DeepMind's broader exploration, highlighting coding's unique verification tractability.31:46–34:16 · Matt as informed peer 4/10 Benchmark Metrics, SWE-Bench Saturation, and Real-World Impact Matt requests hard facts and metrics on Sonnet 4.5's SWE-bench score. Sholto cites the jump from 72 to 78 percent, notes benchmark saturation, and emphasizes real-world partner adoption like Cognition's Devin.34:16–36:40 · Matt as informed peer 5/10 From Short Prompts to 30-Hour Autonomous Execution Matt articulates a sharp startup thesis about betting on six-month model capabilities. Sholto agrees enthusiastically and highlights Sonnet 4.5's ability to run autonomously for 30 hours to build full applications.36:40–41:02 · Matt as informed peer 4/10 Terminal Loops, Memory Management, and Self-Correction Matt asks what an agent physically does during 30 hours of execution. Sholto breaks down terminal execution loops, tool use, markdown memory files, and emergent self-correction capabilities.41:02–43:09 · Matt as informed peer 5/10 End-to-End Application Generation and Replicating Claude.ai Matt brings up Anthropic's viral showcase demo where Claude autonomously replicated Claude.ai with artifacts. Sholto notes this represents the first steps toward fully functional application synthesis.43:09–46:28 · Matt as informed peer 4/10 Overcoming Context Limits and Teaching Model Taste Matt pushes on the underlying technical differences between 7-hour and 30-hour agent limits. Sholto discusses Tesla-style intervention rates, global context drift, and multi-agent systems for teaching coding taste.46:28–50:48 · Matt as informed peer 3/10 Cumulative AI Progress and Moore's Law Analogies Matt asks about the core technical breakthroughs behind Sonnet 4.5. Sholto provides an intuitive analogy explaining pre-training as skim reading textbooks and RL as working out problems with graded feedback.50:48–55:55 · Matt as informed peer 5/10 Overlap of Test-Time Compute and Reinforcement Learning Matt raises historical RL achievements like AlphaGo and Sutton's decades of research to ask why LLM RL converged in 2025. Sholto explains that simple reward feedback on strong base models beat complex search trees.55:55–59:38 · Matt as informed peer 4/10 Defining AGI and Evaluating Future AI Capabilities Matt playfully pushes back on Sholto calling a harder AGI criteria stronger. Sholto defines AGI as outperforming most humans on computer tasks within 2-3 years and discusses parallel execution leverage.59:38–1:02:05 · Matt as informed peer 6/10 Addressing Counter-Theses and Architectural Debates in AI Matt directly confronts Sholto with Yann LeCun's counter-thesis that transformers and current paradigms are fundamentally flawed. Sholto rejects the premise, emphasizing consistent benchmark progress.1:02:05–1:05:39 · Matt as informed peer 5/10 Debunking the AI Plateau Myth and Economic Evaluations Sholto forcefully debunks the AI plateau narrative, comparing LLM pipelines to duct tape versus centuries-old sailboat design. Matt adds details about economic GDP evaluations.1:05:39–1:09:30 · Matt as informed peer 4/10 Preparing for Individual Leverage in an AI-Driven World Matt asks how individuals should prepare and raises physical hand manipulation limits in robotics. Sholto debunks Moravec's paradox as a data artifact and explains generator-verifier gaps in robotics.1:08–4:14 · Guest teaching 3/10 Rapid Model Releases and Hardware Lead Times Matt asks about Anthropic's rapid release pacing and model hierarchy. Sholto explains the two-paradigm regime of pre-training plus RL scaling and compute supply lead times.4:14–6:38 · Guest teaching 1/10 Growing Up in Australia and Athletic Mentorship Matt asks about Sholto's background in Australia. Sholto describes his mother's medical mentorship and his elite international fencing trajectory.6:38–9:27 · Guest teaching 2/10 Fencing as Reinforcement Learning and Discovering AI Matt perceptively framing YouTube-based athletic analysis as an early form of reinforcement learning. Sholto agrees and outlines reading Gwern's scaling essay and starting independent robotics research.9:27–12:39 · Guest teaching 5/10 Academic Obstacles versus Independent High-Signal Work Matt probes the difference between academic gatekeeping and AI research effectiveness. Sholto explains how independent artifacts like Simon Boehm's CUDA guide provide higher signal than traditional PhD credentials.12:39–15:20 · Guest teaching 4/10 Navigating Gemini's Formation and Building Inference Stacks Sholto details joining Google before ChatGPT, building Gemini's inference stack from scratch to save hundreds of millions, and navigating multi-team sociopolitical dynamics.15:20–18:40 · Guest teaching 4/10 Transition to Anthropic and Shared Alignment Focus Matt asks how top AI labs differ given similar compute resources. Sholto contrasts DeepMind's scientific focus with Anthropic's laser focus on alignment and near-term economic impact through coding.18:40–21:25 · Guest teaching 4/10 Mechanistic Understanding and Sutton's Bitter Lesson Matt demonstrates explicit AI research knowledge by interjecting to define Richard Sutton's Bitter Lesson essay. Sholto explains how researcher taste acts as a simplicity regularizer in single training runs.21:25–24:45 · Guest teaching 5/10 Convolutional Networks, Vision Transformers, and Language Structure Sholto provides concrete architectural examples of Sutton's Bitter Lesson, explaining why CNN priors and explicit grammar structures get washed away by general vision transformers and language scale.24:45–27:23 · Guest teaching 4/10 Balancing Short-Term Wins with Long-Term Techniques Matt asks about research failure rates and managing expensive compute budgets. Sholto quotes Noam Shazeer's 10 percent idea hit rate and discusses balancing short-term tweaks with fundamental long-term research.27:23–31:46 · Guest teaching 5/10 Anthropic's Targeted AGI Bet versus DeepMind's Broad Scope Matt asks whether Anthropic researchers investigate non-transformer or non-RL paradigms. Sholto explains Anthropic's focused AGI bet versus DeepMind's broader exploration, highlighting coding's unique verification tractability.31:46–34:16 · Guest teaching 4/10 Benchmark Metrics, SWE-Bench Saturation, and Real-World Impact Matt requests hard facts and metrics on Sonnet 4.5's SWE-bench score. Sholto cites the jump from 72 to 78 percent, notes benchmark saturation, and emphasizes real-world partner adoption like Cognition's Devin.34:16–36:40 · Guest teaching 3/10 From Short Prompts to 30-Hour Autonomous Execution Matt articulates a sharp startup thesis about betting on six-month model capabilities. Sholto agrees enthusiastically and highlights Sonnet 4.5's ability to run autonomously for 30 hours to build full applications.36:40–41:02 · Guest teaching 5/10 Terminal Loops, Memory Management, and Self-Correction Matt asks what an agent physically does during 30 hours of execution. Sholto breaks down terminal execution loops, tool use, markdown memory files, and emergent self-correction capabilities.41:02–43:09 · Guest teaching 3/10 End-to-End Application Generation and Replicating Claude.ai Matt brings up Anthropic's viral showcase demo where Claude autonomously replicated Claude.ai with artifacts. Sholto notes this represents the first steps toward fully functional application synthesis.43:09–46:28 · Guest teaching 5/10 Overcoming Context Limits and Teaching Model Taste Matt pushes on the underlying technical differences between 7-hour and 30-hour agent limits. Sholto discusses Tesla-style intervention rates, global context drift, and multi-agent systems for teaching coding taste.46:28–50:48 · Guest teaching 5/10 Cumulative AI Progress and Moore's Law Analogies Matt asks about the core technical breakthroughs behind Sonnet 4.5. Sholto provides an intuitive analogy explaining pre-training as skim reading textbooks and RL as working out problems with graded feedback.50:48–55:55 · Guest teaching 5/10 Overlap of Test-Time Compute and Reinforcement Learning Matt raises historical RL achievements like AlphaGo and Sutton's decades of research to ask why LLM RL converged in 2025. Sholto explains that simple reward feedback on strong base models beat complex search trees.55:55–59:38 · Guest teaching 5/10 Defining AGI and Evaluating Future AI Capabilities Matt playfully pushes back on Sholto calling a harder AGI criteria stronger. Sholto defines AGI as outperforming most humans on computer tasks within 2-3 years and discusses parallel execution leverage.59:38–1:02:05 · Guest teaching 4/10 Addressing Counter-Theses and Architectural Debates in AI Matt directly confronts Sholto with Yann LeCun's counter-thesis that transformers and current paradigms are fundamentally flawed. Sholto rejects the premise, emphasizing consistent benchmark progress.1:02:05–1:05:39 · Guest teaching 4/10 Debunking the AI Plateau Myth and Economic Evaluations Sholto forcefully debunks the AI plateau narrative, comparing LLM pipelines to duct tape versus centuries-old sailboat design. Matt adds details about economic GDP evaluations.1:05:39–1:09:30 · Guest teaching 5/10 Preparing for Individual Leverage in an AI-Driven World Matt asks how individuals should prepare and raises physical hand manipulation limits in robotics. Sholto debunks Moravec's paradox as a data artifact and explains generator-verifier gaps in robotics.1:08–4:14 · Guest disagreement 1/10 Rapid Model Releases and Hardware Lead Times Matt asks about Anthropic's rapid release pacing and model hierarchy. Sholto explains the two-paradigm regime of pre-training plus RL scaling and compute supply lead times.4:14–6:38 · Guest disagreement 0/10 Growing Up in Australia and Athletic Mentorship Matt asks about Sholto's background in Australia. Sholto describes his mother's medical mentorship and his elite international fencing trajectory.6:38–9:27 · Guest disagreement 0/10 Fencing as Reinforcement Learning and Discovering AI Matt perceptively framing YouTube-based athletic analysis as an early form of reinforcement learning. Sholto agrees and outlines reading Gwern's scaling essay and starting independent robotics research.9:27–12:39 · Guest disagreement 1/10 Academic Obstacles versus Independent High-Signal Work Matt probes the difference between academic gatekeeping and AI research effectiveness. Sholto explains how independent artifacts like Simon Boehm's CUDA guide provide higher signal than traditional PhD credentials.12:39–15:20 · Guest disagreement 0/10 Navigating Gemini's Formation and Building Inference Stacks Sholto details joining Google before ChatGPT, building Gemini's inference stack from scratch to save hundreds of millions, and navigating multi-team sociopolitical dynamics.15:20–18:40 · Guest disagreement 1/10 Transition to Anthropic and Shared Alignment Focus Matt asks how top AI labs differ given similar compute resources. Sholto contrasts DeepMind's scientific focus with Anthropic's laser focus on alignment and near-term economic impact through coding.18:40–21:25 · Guest disagreement 0/10 Mechanistic Understanding and Sutton's Bitter Lesson Matt demonstrates explicit AI research knowledge by interjecting to define Richard Sutton's Bitter Lesson essay. Sholto explains how researcher taste acts as a simplicity regularizer in single training runs.21:25–24:45 · Guest disagreement 0/10 Convolutional Networks, Vision Transformers, and Language Structure Sholto provides concrete architectural examples of Sutton's Bitter Lesson, explaining why CNN priors and explicit grammar structures get washed away by general vision transformers and language scale.24:45–27:23 · Guest disagreement 0/10 Balancing Short-Term Wins with Long-Term Techniques Matt asks about research failure rates and managing expensive compute budgets. Sholto quotes Noam Shazeer's 10 percent idea hit rate and discusses balancing short-term tweaks with fundamental long-term research.27:23–31:46 · Guest disagreement 1/10 Anthropic's Targeted AGI Bet versus DeepMind's Broad Scope Matt asks whether Anthropic researchers investigate non-transformer or non-RL paradigms. Sholto explains Anthropic's focused AGI bet versus DeepMind's broader exploration, highlighting coding's unique verification tractability.31:46–34:16 · Guest disagreement 1/10 Benchmark Metrics, SWE-Bench Saturation, and Real-World Impact Matt requests hard facts and metrics on Sonnet 4.5's SWE-bench score. Sholto cites the jump from 72 to 78 percent, notes benchmark saturation, and emphasizes real-world partner adoption like Cognition's Devin.34:16–36:40 · Guest disagreement 0/10 From Short Prompts to 30-Hour Autonomous Execution Matt articulates a sharp startup thesis about betting on six-month model capabilities. Sholto agrees enthusiastically and highlights Sonnet 4.5's ability to run autonomously for 30 hours to build full applications.36:40–41:02 · Guest disagreement 0/10 Terminal Loops, Memory Management, and Self-Correction Matt asks what an agent physically does during 30 hours of execution. Sholto breaks down terminal execution loops, tool use, markdown memory files, and emergent self-correction capabilities.41:02–43:09 · Guest disagreement 0/10 End-to-End Application Generation and Replicating Claude.ai Matt brings up Anthropic's viral showcase demo where Claude autonomously replicated Claude.ai with artifacts. Sholto notes this represents the first steps toward fully functional application synthesis.43:09–46:28 · Guest disagreement 0/10 Overcoming Context Limits and Teaching Model Taste Matt pushes on the underlying technical differences between 7-hour and 30-hour agent limits. Sholto discusses Tesla-style intervention rates, global context drift, and multi-agent systems for teaching coding taste.46:28–50:48 · Guest disagreement 0/10 Cumulative AI Progress and Moore's Law Analogies Matt asks about the core technical breakthroughs behind Sonnet 4.5. Sholto provides an intuitive analogy explaining pre-training as skim reading textbooks and RL as working out problems with graded feedback.50:48–55:55 · Guest disagreement 0/10 Overlap of Test-Time Compute and Reinforcement Learning Matt raises historical RL achievements like AlphaGo and Sutton's decades of research to ask why LLM RL converged in 2025. Sholto explains that simple reward feedback on strong base models beat complex search trees.55:55–59:38 · Guest disagreement 1/10 Defining AGI and Evaluating Future AI Capabilities Matt playfully pushes back on Sholto calling a harder AGI criteria stronger. Sholto defines AGI as outperforming most humans on computer tasks within 2-3 years and discusses parallel execution leverage.59:38–1:02:05 · Guest disagreement 2/10 Addressing Counter-Theses and Architectural Debates in AI Matt directly confronts Sholto with Yann LeCun's counter-thesis that transformers and current paradigms are fundamentally flawed. Sholto rejects the premise, emphasizing consistent benchmark progress.1:02:05–1:05:39 · Guest disagreement 3/10 Debunking the AI Plateau Myth and Economic Evaluations Sholto forcefully debunks the AI plateau narrative, comparing LLM pipelines to duct tape versus centuries-old sailboat design. Matt adds details about economic GDP evaluations.1:05:39–1:09:30 · Guest disagreement 1/10 Preparing for Individual Leverage in an AI-Driven World Matt asks how individuals should prepare and raises physical hand manipulation limits in robotics. Sholto debunks Moravec's paradox as a data artifact and explains generator-verifier gaps in robotics.1:08–4:14 · Matt pushing back 1/10 Rapid Model Releases and Hardware Lead Times Matt asks about Anthropic's rapid release pacing and model hierarchy. Sholto explains the two-paradigm regime of pre-training plus RL scaling and compute supply lead times.4:14–6:38 · Matt pushing back 0/10 Growing Up in Australia and Athletic Mentorship Matt asks about Sholto's background in Australia. Sholto describes his mother's medical mentorship and his elite international fencing trajectory.6:38–9:27 · Matt pushing back 0/10 Fencing as Reinforcement Learning and Discovering AI Matt perceptively framing YouTube-based athletic analysis as an early form of reinforcement learning. Sholto agrees and outlines reading Gwern's scaling essay and starting independent robotics research.9:27–12:39 · Matt pushing back 1/10 Academic Obstacles versus Independent High-Signal Work Matt probes the difference between academic gatekeeping and AI research effectiveness. Sholto explains how independent artifacts like Simon Boehm's CUDA guide provide higher signal than traditional PhD credentials.12:39–15:20 · Matt pushing back 0/10 Navigating Gemini's Formation and Building Inference Stacks Sholto details joining Google before ChatGPT, building Gemini's inference stack from scratch to save hundreds of millions, and navigating multi-team sociopolitical dynamics.15:20–18:40 · Matt pushing back 1/10 Transition to Anthropic and Shared Alignment Focus Matt asks how top AI labs differ given similar compute resources. Sholto contrasts DeepMind's scientific focus with Anthropic's laser focus on alignment and near-term economic impact through coding.18:40–21:25 · Matt pushing back 1/10 Mechanistic Understanding and Sutton's Bitter Lesson Matt demonstrates explicit AI research knowledge by interjecting to define Richard Sutton's Bitter Lesson essay. Sholto explains how researcher taste acts as a simplicity regularizer in single training runs.21:25–24:45 · Matt pushing back 0/10 Convolutional Networks, Vision Transformers, and Language Structure Sholto provides concrete architectural examples of Sutton's Bitter Lesson, explaining why CNN priors and explicit grammar structures get washed away by general vision transformers and language scale.24:45–27:23 · Matt pushing back 1/10 Balancing Short-Term Wins with Long-Term Techniques Matt asks about research failure rates and managing expensive compute budgets. Sholto quotes Noam Shazeer's 10 percent idea hit rate and discusses balancing short-term tweaks with fundamental long-term research.27:23–31:46 · Matt pushing back 1/10 Anthropic's Targeted AGI Bet versus DeepMind's Broad Scope Matt asks whether Anthropic researchers investigate non-transformer or non-RL paradigms. Sholto explains Anthropic's focused AGI bet versus DeepMind's broader exploration, highlighting coding's unique verification tractability.31:46–34:16 · Matt pushing back 1/10 Benchmark Metrics, SWE-Bench Saturation, and Real-World Impact Matt requests hard facts and metrics on Sonnet 4.5's SWE-bench score. Sholto cites the jump from 72 to 78 percent, notes benchmark saturation, and emphasizes real-world partner adoption like Cognition's Devin.34:16–36:40 · Matt pushing back 1/10 From Short Prompts to 30-Hour Autonomous Execution Matt articulates a sharp startup thesis about betting on six-month model capabilities. Sholto agrees enthusiastically and highlights Sonnet 4.5's ability to run autonomously for 30 hours to build full applications.36:40–41:02 · Matt pushing back 1/10 Terminal Loops, Memory Management, and Self-Correction Matt asks what an agent physically does during 30 hours of execution. Sholto breaks down terminal execution loops, tool use, markdown memory files, and emergent self-correction capabilities.41:02–43:09 · Matt pushing back 0/10 End-to-End Application Generation and Replicating Claude.ai Matt brings up Anthropic's viral showcase demo where Claude autonomously replicated Claude.ai with artifacts. Sholto notes this represents the first steps toward fully functional application synthesis.43:09–46:28 · Matt pushing back 1/10 Overcoming Context Limits and Teaching Model Taste Matt pushes on the underlying technical differences between 7-hour and 30-hour agent limits. Sholto discusses Tesla-style intervention rates, global context drift, and multi-agent systems for teaching coding taste.46:28–50:48 · Matt pushing back 1/10 Cumulative AI Progress and Moore's Law Analogies Matt asks about the core technical breakthroughs behind Sonnet 4.5. Sholto provides an intuitive analogy explaining pre-training as skim reading textbooks and RL as working out problems with graded feedback.50:48–55:55 · Matt pushing back 1/10 Overlap of Test-Time Compute and Reinforcement Learning Matt raises historical RL achievements like AlphaGo and Sutton's decades of research to ask why LLM RL converged in 2025. Sholto explains that simple reward feedback on strong base models beat complex search trees.55:55–59:38 · Matt pushing back 2/10 Defining AGI and Evaluating Future AI Capabilities Matt playfully pushes back on Sholto calling a harder AGI criteria stronger. Sholto defines AGI as outperforming most humans on computer tasks within 2-3 years and discusses parallel execution leverage.59:38–1:02:05 · Matt pushing back 3/10 Addressing Counter-Theses and Architectural Debates in AI Matt directly confronts Sholto with Yann LeCun's counter-thesis that transformers and current paradigms are fundamentally flawed. Sholto rejects the premise, emphasizing consistent benchmark progress.1:02:05–1:05:39 · Matt pushing back 1/10 Debunking the AI Plateau Myth and Economic Evaluations Sholto forcefully debunks the AI plateau narrative, comparing LLM pipelines to duct tape versus centuries-old sailboat design. Matt adds details about economic GDP evaluations.1:05:39–1:09:30 · Matt pushing back 1/10 Preparing for Individual Leverage in an AI-Driven World Matt asks how individuals should prepare and raises physical hand manipulation limits in robotics. Sholto debunks Moravec's paradox as a data artifact and explains generator-verifier gaps in robotics.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 47.9% · guest 52.1%0:00 · Matt 47.9% · guest 52.1%3:00 · Matt 9.5% · guest 90.5%3:00 · Matt 9.5% · guest 90.5%6:00 · Matt 11.6% · guest 88.4%6:00 · Matt 11.6% · guest 88.4%9:00 · Matt 16% · guest 84%9:00 · Matt 16% · guest 84%12:00 · Matt 10.1% · guest 89.9%12:00 · Matt 10.1% · guest 89.9%15:00 · Matt 21.7% · guest 78.3%15:00 · Matt 21.7% · guest 78.3%18:00 · Matt 7.6% · guest 92.4%18:00 · Matt 7.6% · guest 92.4%21:00 · Matt 20.6% · guest 79.4%21:00 · Matt 20.6% · guest 79.4%24:00 · Matt 24.6% · guest 75.4%24:00 · Matt 24.6% · guest 75.4%27:00 · Matt 11.2% · guest 88.8%27:00 · Matt 11.2% · guest 88.8%30:00 · Matt 12.9% · guest 87.1%30:00 · Matt 12.9% · guest 87.1%33:00 · Matt 7.2% · guest 92.8%33:00 · Matt 7.2% · guest 92.8%36:00 · Matt 24.3% · guest 75.7%36:00 · Matt 24.3% · guest 75.7%39:00 · Matt 6.7% · guest 93.3%39:00 · Matt 6.7% · guest 93.3%42:00 · Matt 30.1% · guest 69.9%42:00 · Matt 30.1% · guest 69.9%45:00 · Matt 19.2% · guest 80.8%45:00 · Matt 19.2% · guest 80.8%48:00 · Matt 9.3% · guest 90.7%48:00 · Matt 9.3% · guest 90.7%51:00 · Matt 13.3% · guest 86.7%51:00 · Matt 13.3% · guest 86.7%54:00 · Matt 14% · guest 86%54:00 · Matt 14% · guest 86%57:00 · Matt 8% · guest 92%57:00 · Matt 8% · guest 92%1:00:00 · Matt 13% · guest 87%1:00:00 · Matt 13% · guest 87%1:03:00 · Matt 27.7% · guest 72.3%1:03:00 · Matt 27.7% · guest 72.3%1:06:00 · Matt 9.2% · guest 90.8%1:06:00 · Matt 9.2% · guest 90.8%1:09:00 · Matt 48.3% · guest 51.7%1:09:00 · Matt 48.3% · guest 51.7%
Sharpest disagreement ▶ 1:02:26 Debunking the AI plateau myth with the sailboat analogy

Sholto aggressively dismisses the popular AI plateau narrative, arguing that current LLM pipelines are primitive duct-tape systems with massive room for expansion.

Hardest push from Matt ▶ 59:38 Challenging the lab consensus with Yann LeCun's counter-thesis

Matt directly pushes back against lab optimism by citing prominent skeptics like Yann LeCun who claim current architectures cannot reach AGI.

Biggest teaching moment ▶ 48:08 Textbook skimming versus workbook problem solving

Sholto offers a crystal-clear masterclass analogy explaining why reinforcement learning unlocks capabilities that pre-training alone can never produce.

Matt holds his own ▶ 21:07 Defining Richard Sutton's Bitter Lesson

Matt demonstrates deep familiarity with foundational AI literature by interjecting to cite and explain Richard Sutton's Bitter Lesson essay.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Rapid Model Releases and Hardware Lead Times 3311 Matt asks about Anthropic's rapid release pacing and model hierarchy. Sholto explains the two-paradigm regime of pre-training plus RL scaling and compute supply lead times.
Growing Up in Australia and Athletic Mentorship 1100 Matt asks about Sholto's background in Australia. Sholto describes his mother's medical mentorship and his elite international fencing trajectory.
Fencing as Reinforcement Learning and Discovering AI 4200 Matt perceptively framing YouTube-based athletic analysis as an early form of reinforcement learning. Sholto agrees and outlines reading Gwern's scaling essay and starting independent robotics research.
Academic Obstacles versus Independent High-Signal Work 4511 Matt probes the difference between academic gatekeeping and AI research effectiveness. Sholto explains how independent artifacts like Simon Boehm's CUDA guide provide higher signal than traditional PhD credentials.
Navigating Gemini's Formation and Building Inference Stacks 2400 Sholto details joining Google before ChatGPT, building Gemini's inference stack from scratch to save hundreds of millions, and navigating multi-team sociopolitical dynamics.
Transition to Anthropic and Shared Alignment Focus 3411 Matt asks how top AI labs differ given similar compute resources. Sholto contrasts DeepMind's scientific focus with Anthropic's laser focus on alignment and near-term economic impact through coding.
Mechanistic Understanding and Sutton's Bitter Lesson 6401 Matt demonstrates explicit AI research knowledge by interjecting to define Richard Sutton's Bitter Lesson essay. Sholto explains how researcher taste acts as a simplicity regularizer in single training runs.
Convolutional Networks, Vision Transformers, and Language Structure 2500 Sholto provides concrete architectural examples of Sutton's Bitter Lesson, explaining why CNN priors and explicit grammar structures get washed away by general vision transformers and language scale.
Balancing Short-Term Wins with Long-Term Techniques 4401 Matt asks about research failure rates and managing expensive compute budgets. Sholto quotes Noam Shazeer's 10 percent idea hit rate and discusses balancing short-term tweaks with fundamental long-term research.
Anthropic's Targeted AGI Bet versus DeepMind's Broad Scope 4511 Matt asks whether Anthropic researchers investigate non-transformer or non-RL paradigms. Sholto explains Anthropic's focused AGI bet versus DeepMind's broader exploration, highlighting coding's unique verification tractability.
Benchmark Metrics, SWE-Bench Saturation, and Real-World Impact 4411 Matt requests hard facts and metrics on Sonnet 4.5's SWE-bench score. Sholto cites the jump from 72 to 78 percent, notes benchmark saturation, and emphasizes real-world partner adoption like Cognition's Devin.
From Short Prompts to 30-Hour Autonomous Execution 5301 Matt articulates a sharp startup thesis about betting on six-month model capabilities. Sholto agrees enthusiastically and highlights Sonnet 4.5's ability to run autonomously for 30 hours to build full applications.
Terminal Loops, Memory Management, and Self-Correction 4501 Matt asks what an agent physically does during 30 hours of execution. Sholto breaks down terminal execution loops, tool use, markdown memory files, and emergent self-correction capabilities.
End-to-End Application Generation and Replicating Claude.ai 5300 Matt brings up Anthropic's viral showcase demo where Claude autonomously replicated Claude.ai with artifacts. Sholto notes this represents the first steps toward fully functional application synthesis.
Overcoming Context Limits and Teaching Model Taste 4501 Matt pushes on the underlying technical differences between 7-hour and 30-hour agent limits. Sholto discusses Tesla-style intervention rates, global context drift, and multi-agent systems for teaching coding taste.
Cumulative AI Progress and Moore's Law Analogies 3501 Matt asks about the core technical breakthroughs behind Sonnet 4.5. Sholto provides an intuitive analogy explaining pre-training as skim reading textbooks and RL as working out problems with graded feedback.
Overlap of Test-Time Compute and Reinforcement Learning 5501 Matt raises historical RL achievements like AlphaGo and Sutton's decades of research to ask why LLM RL converged in 2025. Sholto explains that simple reward feedback on strong base models beat complex search trees.
Defining AGI and Evaluating Future AI Capabilities 4512 Matt playfully pushes back on Sholto calling a harder AGI criteria stronger. Sholto defines AGI as outperforming most humans on computer tasks within 2-3 years and discusses parallel execution leverage.
Addressing Counter-Theses and Architectural Debates in AI 6423 Matt directly confronts Sholto with Yann LeCun's counter-thesis that transformers and current paradigms are fundamentally flawed. Sholto rejects the premise, emphasizing consistent benchmark progress.
Debunking the AI Plateau Myth and Economic Evaluations 5431 Sholto forcefully debunks the AI plateau narrative, comparing LLM pipelines to duct tape versus centuries-old sailboat design. Matt adds details about economic GDP evaluations.
Preparing for Individual Leverage in an AI-Driven World 4511 Matt asks how individuals should prepare and raises physical hand manipulation limits in robotics. Sholto debunks Moravec's paradox as a data artifact and explains generator-verifier gaps in robotics.

Statements from this episode (50)

Insight
Douglas: AI releases accelerating due to dual-paradigm scaling
“There's now this two paradigm regime where previously you did free training scaling and reinforcement learning scaling, and now we're in a mix of the two basically. And so I think that gives you more opportunities to update models because it means that you can…”
Sholto Douglas Oct 2, 2025 ▶ 1:45
What-if
Douglas: TSMC bottlenecked desired AI chip acquisitions in 2024
“Even if you, as much as you wanted chips last year, it would have been impossible to get them because TSMC was, you know, booked out and so forth.”
Sholto Douglas Oct 2, 2025 ▶ 2:34
Assertion Not checkable as stated
Douglas: The post-ChatGPT AI compute supercycle begins properly in 2025
“So finally, this year is where the compute, like, super cycle is, like, beginning properly in effect.”
Sholto Douglas Oct 2, 2025 ▶ 2:40
Assertion Supported
Douglas: Anthropic's mid-tier Sonnet is smarter than its flagship Opus
“One of the interesting things about this most recent release is actually Sonnet is smarter than Opus.”
Sholto Douglas Oct 2, 2025 ▶ 3:12
Insight
Douglas: Reinforcement learning elevates mid-tier models to match older flagships
“So that allows you to take a mid-tier model and make it as good as a larger-tier model of six months ago or three months ago.”
Sholto Douglas Oct 2, 2025 ▶ 4:08
Insight
Douglas: Independent technical blogs are the highest AI hiring signals
“The fastest route, or like, the most immediate one is whenever we see a really good blog post where people have, like, done incredible amount of work in an independent fashion, it's one of the highest signal things there is.”
Sholto Douglas Oct 2, 2025 ▶ 10:54
Assertion Not checkable as stated
Douglas: Google LLM inference stack saved hundreds of millions in months
“This ended up saving several hundred million dollars, I think, like, even over the first six months”
Sholto Douglas Oct 2, 2025 ▶ 14:18
Prediction Not checkable as stated
Douglas predicts DeepMind will lead the world in AI science discoveries
“DeepMind, if you wanted to solve science, is the best place in the world. Like, I think that DeepMind will directly contribute to more scientific discoveries from AI than anything else, right?”
Sholto Douglas Oct 2, 2025 ▶ 16:54
Disclosure
Douglas: Anthropic focuses on alignment and near-term economic coding utility
“Anthropic has been laser focused on on two things. One is, like, long-term AI alignment, and two is near-term economic impact. So Anthropik has been laser-focused on coding and computer use, and things that we think will make a direct impact to the economy, li…”
Sholto Douglas Oct 2, 2025 ▶ 17:25
Disclosure
Douglas: Anthropic intentionally deprioritized math reasoning unlike OpenAI and DeepMind
“You know, one thing that Anthropik, like, You noticeably hasn't focused on compared to DeepMind and to OpenAI is is mathematical reasoning, right? DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and for sci…”
Sholto Douglas Oct 2, 2025 ▶ 17:47
Insight
Douglas: Compute scale consistently wipes out clever human priors in AI
“Generations of people have developed clever methods of encoding priors about how they think an artificial intelligence should reason. And encoding it into the model, and all of this gets wiped out by scale and, you know, planning, like, basically, like, search…”
Sholto Douglas Oct 2, 2025 ▶ 20:40
Insight
Douglas: Vision transformers outscale CNNs by leveraging global image context
“Convolutional neural networks will be better than a more general, like, a vision transformer for the vast majority of images up to a certain, and up to a certain amount of scale. But past a point, actually, you need to be able to flexibly integrate information…”
Sholto Douglas Oct 2, 2025 ▶ 22:23
Assertion Not checkable as stated
Douglas: Even AI pioneer Noam Shazeer only sees 10% of ideas work
“I once asked this question of Noam Chazier. And he was like, yeah, maybe like 10% of my ideas work, and that's not, right? You know, one of the, you know, an absolute genius, one of the best in the field. So if only 10% of his ideas work, then I think that, yo…”
Sholto Douglas Oct 2, 2025 ▶ 24:47
Disclosure
Douglas: Top AI labs give researchers months to explore unproven ideas
“It's one of the things that at both Anthropic and at DeepMind we really tried to build, which is like a culture of safe experimentation where People were trusted to explore ideas for a long time out in the wild. Cause you often need months of independent resea…”
Sholto Douglas Oct 2, 2025 ▶ 25:44
Insight
Douglas: Examining training data and simple tweaks yields massive model gains
“A really high ROI use on, on your time would probably be to just go and look at the data and think hard about what, like the model is learning or doing and like make some tweaks that makes it like, there's so many, like you could even just The simplest things …”
Sholto Douglas Oct 2, 2025 ▶ 26:33
Prediction Not checkable as stated
Douglas: Anthropic believes AGI is reachable in a couple of years
“We think that, you know, AGI is within reach in the next couple of years.”
Sholto Douglas Oct 2, 2025 ▶ 27:33
Disclosure
Douglas: Anthropic's ethos is that scaling current techniques achieves AGI
“Like really for the last five or six years, Anthropics ethos has been scaling compute with broadly the current set of techniques is like AGI is tractable within those bounds.”
Sholto Douglas Oct 2, 2025 ▶ 27:52
Assertion Not checkable as stated
Douglas: DeepMind has 1,000 on Gemini and 10,000 on foundational research
“But if you look, Gemini is like, you know, a thousand people, there's a, there's still like 10,000 plus people doing all kinds of really like longterm foundational research at DMI.”
Sholto Douglas Oct 2, 2025 ▶ 28:37
Insight
Douglas: AI takeoff speed depends on AI assisting AI research
“We think that one of the most important signals of whether or not we are basically the speed of takeoff, the speed of progress is driven by how much AI is able to assist AI research.”
Sholto Douglas Oct 2, 2025 ▶ 29:18
Insight
Douglas: Coding is uniquely tractable for AI because execution is verifiable
“Coding is a uniquely tractable problem in some respects for the techniques that we have in terms, the data exists in many ways. You can containerize and run things in parallel. You can run unit tests, and so you can verify.”
Sholto Douglas Oct 2, 2025 ▶ 30:29
Assertion Supported
Douglas: Sonnet 4.5 pushed SWE-bench scores from roughly 72% to 78%
“We moved recently from roughly 72 to roughly 78 in Sweepbench”
Sholto Douglas Oct 2, 2025 ▶ 32:58
Assertion Partly supported
Douglas: The entire AI industry scored under 20% on SWE-bench last year
“As recently as a year ago, I think we were under 20% or something like that as a field.”
Sholto Douglas Oct 2, 2025 ▶ 33:05
Assertion Supported
Douglas: Cognition rebuilt Devin's architecture around Anthropic's highly useful Claude Sonnet
“The Cognition folks in Devon found the model so useful, they had to, like, rebuild their architecture around it.”
Sholto Douglas Oct 2, 2025 ▶ 33:56
Opinion
Douglas: Claude 3.5 Sonnet drove product-market fit for code editor Cursor
“In many ways, this model is what caused PMF for Cursor. Cursor took off like a rocket, right, with because they were in the right place, and they were able to capitalize on that model as offering a coding experience that didn't previously exist.”
Sholto Douglas Oct 2, 2025 ▶ 34:37
Insight
Turck: AI startups in 2025 must build for capabilities six months away
“One of the key lessons for anybody in the startup world. In 2025, which is bet on what the models will be able to do in six months from now.”
Matt Turck Oct 2, 2025 ▶ 35:19
Prediction Not checkable as stated
Douglas: AI coding agent supervision will drop to 20-minute intervals within months
“Over time you know, over the next couple of months, you're probably gonna end up in a situation where you only need to supervise the models every, you know, 10 minutes, 20 minutes or so.”
Sholto Douglas Oct 2, 2025 ▶ 35:49
Assertion Not checkable as stated
Douglas: An Anthropic AI agent operated autonomously for 30 hours building apps
“We asked it to build something that looks roughly like a chat app, you know, something like Slack or, you know. And it was, it, the model just worked for 30 hours. Like, it was just spinning there on a computer for 30 hours, and came out with a really good wor…”
Sholto Douglas Oct 2, 2025 ▶ 36:03
Disclosure
Douglas: Anthropic built persistent memory files into its AI agentic harness
“One of the things that we're pretty happy with about the recent launches, we've finally taught the models to use what's called memory. And so, and we've built that into the agentic harness. So it's able to create a markdown file of to-dos and things that it th…”
Sholto Douglas Oct 2, 2025 ▶ 37:49
Assertion Not checkable as stated
Douglas: The current generation of AI agents are astonishingly good at self-correcting
“I think one of the things that made me remarkable about the current generation of agents is that they can self-correct. In fact, they're astonishingly good at self-correcting.”
Sholto Douglas Oct 2, 2025 ▶ 38:23
Assertion Supported
Douglas: AI autonomous task execution time horizons double every six months
“And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something, something crazy. Maybe, maybe every six months the time horizon doubles”
Sholto Douglas Oct 2, 2025 ▶ 40:13
Assertion Contradicted
Douglas: Anthropic models autonomously replicated the Claude.ai website in hours
“And in this case, the model replicated Claude.ai with artifacts, with everything else I can't quite remember how long that one took. Maybe a couple hours to do.”
Sholto Douglas Oct 2, 2025 ▶ 42:36
Prediction Not checkable as stated
Douglas: AI application development will see another massive leap next year
“Over the next six months, over the next year, expect dramatic progress here. And like look at where we are now versus where we were a year ago. And the difference is I expect the same jump basically.”
Sholto Douglas Oct 2, 2025 ▶ 42:57
Assertion Not checkable as stated
Douglas: AI coding interventions stem from taste, not raw programming capability
“Right now you need to intervene quite frequently, but it's usually on questions of taste rather than it is questions of, like, raw programming ability.”
Sholto Douglas Oct 2, 2025 ▶ 43:54
Assertion Partly supported
Douglas: AI progress on METR evaluations plots as a straight line
“Like on the meter eval, if you look at progress over the last two years, you can plot it with a straight line, right?”
Sholto Douglas Oct 2, 2025 ▶ 47:16
Insight
Douglas: AI progress relies on accumulated work and compute, not single breakthroughs
“It's not any one critical breakthrough. It's more the accumulation of a huge amount of work in an environment where there's a sort of exogenous force of compute pushing progress forward.”
Sholto Douglas Oct 2, 2025 ▶ 47:31
Insight
Douglas: Solving AI hallucinations intrinsically requires reinforcement learning
“Saying, I don't know, or solving, you know, hallucinations is, ah, intrinsically requires reinforcement learning in many ways.”
Sholto Douglas Oct 2, 2025 ▶ 49:35
Assertion Not checkable as stated
Douglas: Reinforcement learning on language models finally started working in late 2024
“I think also an important change in, you know, in this sort of like era of reasoning models and RL on language models is, at the end of last year, RL on language models finally started to work.”
Sholto Douglas Oct 2, 2025 ▶ 49:50
Assertion Not checkable as stated
Douglas: OpenAI's o1 established test-time compute and RL as a scaling axis
“And I think OpenAI deserves a lot of credit for you know, releasing the first, like, serious RL plus LLMs release with O-one. And I think this really kicked off a pretty, you know, substantial change because it opened up a new axis of scaling, right? There was…”
Sholto Douglas Oct 2, 2025 ▶ 50:01
Insight
Douglas: Test-time compute executes reasoning while RL provides feedback on correctness
“One way of thinking about this is test time compute is doing a lot of reasoning, and then RL is the feedback signal on whether or not that reasoning was right or wrong.”
Sholto Douglas Oct 2, 2025 ▶ 50:58
Insight
Douglas: Test-time compute solves harder tasks before RL distills them into models
“Test time compute lets you do harder problems than you can currently do, than you can like currently do off the cuff, and RL Then allows you to sort of distill that back into the model.”
Sholto Douglas Oct 2, 2025 ▶ 51:52
Insight
Douglas: Simple RL methods work better on language models than complex strategies
“One of the craziest things about RL on language models in the, in, like, the RL from verified rewards regime, is it's almost the simplest possible thing. It's, like, almost too simple to work. And this is, again, comes back to that question of taste, where Rea…”
Sholto Douglas Oct 2, 2025 ▶ 52:58
Insight
Douglas: AI reasoning strategies emerge naturally with enough compute and RL feedback
“Give it math questions, tell it whether it got them right or wrong, and the model will learn. This is, it comes down to a bit of lesson in scale and search, is just allow the model to search, have enough compute to run the experiments, and the model actually e…”
Sholto Douglas Oct 2, 2025 ▶ 55:19
Prediction Open · timeframe Oct 2028
Douglas: AI industry will reach human-level computer capabilities in 2-3 years
“Which is that in the next two or three years, given the right feedback loops, given the right compute, given the right, you know, elbow grease and this kind of stuff, we think that we as the AI industry are all on track to create something that is at least as …”
Sholto Douglas Oct 2, 2025 ▶ 59:06
Assertion Not checkable as stated
Douglas: Transformers successfully model any domain given sufficient data and compute
“I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute.”
Sholto Douglas Oct 2, 2025 ▶ 1:00:25
Opinion
Douglas: LLM pipelines are two and a half years of desperate effort
“When I look at an LLM training pipeline, it is two and a half years of best effort, last minute, desperate effort. And there's just so much room to go on every part of it.”
Sholto Douglas Oct 2, 2025 ▶ 1:03:26
Assertion Supported
Douglas: Anthropic's Claude Opus 4.1 led OpenAI's GDP eval benchmark
“4.1 Opus was the leading model there.”
Sholto Douglas Oct 2, 2025 ▶ 1:04:08
Prediction Not checkable as stated
Douglas: AI beating GDP benchmarks won't immediately alter the broader economy
“We'll probably reach like better than human on the GDP eval, and it won't change anything economically because It'll be all the connective tissue, and all the, like, you know, the context, and actually, like, the task won't be representative.”
Sholto Douglas Oct 2, 2025 ▶ 1:04:53
Prediction Not checkable as stated
Douglas: Individuals will manage 24/7 AI agent teams within two years
“If coding agents progress in the way I've been saying, in a year or two, you'll be able to manage a team, basically, that works 24 seven for you doing work.”
Sholto Douglas Oct 2, 2025 ▶ 1:06:11
Opinion
Douglas: Moravec's paradox is fake and primarily a data availability problem
“I actually think Moravec's paradox is a little bit fake, and I think this is mostly a question of data availability and like you know, RL signal and stuff.”
Sholto Douglas Oct 2, 2025 ▶ 1:07:31
Assertion Not checkable as stated
Douglas: Robotic locomotion is essentially solved using basic reinforcement learning
“Locomotion's kind of solved, to be honest, with basic RL.”
Sholto Douglas Oct 2, 2025 ▶ 1:08:09
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.