Jun 4, 2026 · 49m · mad

OpenAI's Dan Roberts: Why AI Can Now Make Discoveries

Dan Roberts · 37m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Dan Roberts, theoretical physicist and Foundations of Reinforcement Learning lead at OpenAI, to explore how reinforcement learning and test-time compute are empowering AI models to make scientific discoveries. Dan shares insights on mathematical problem-solving, scaling laws, physics-inspired theoretical frameworks, and the evolving role of AI as an autonomous research collaborator.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 18.4% of the talking time here. How this is scored →

Matt as informed peer 3.8 Guest teaching 5.2 Guest disagreement 1.8 Matt pushing back 1.4
05100:0015:0030:0045:001:18–7:05 · Matt as informed peer 1/10 Understanding the Foundations of Reinforcement Learning Team Matt asks open-ended introductory questions about Dan's role at OpenAI and his academic trajectory. Dan details his work on RL foundations, scaling laws, and his theoretical physics background studying quantum chaos and black hole information processing.7:05–12:12 · Matt as informed peer 4/10 The Evolution of AI in Autonomous Scientific Discovery Matt demonstrates domain knowledge by referencing recent developments on Erdos problems by OpenAI, DeepMind, and Anthropic. Dan educates Matt on how OpenAI refuted a lower-bound conjecture via algebraic number theory and contrasts DeepMind's Lean auto-formalization with OpenAI's natural language proof generation.12:12–17:04 · Matt as informed peer 2/10 Defining Reinforcement Learning: Analogies and Core Principles Matt asks fundamental definition questions about RL, why it works, and how it fails. Dan delivers an educational breakdown using a Super Mario Bros analogy to contrast supervised learning from demonstrations with interactive reinforcement learning.17:04–24:50 · Matt as informed peer 4/10 Applying RL to Large Language Models and Proxy Reward Models Matt prompts Dan on RLHF history, reward proxies, AlphaGo's Move 37, and exploration vs exploitation. Dan shares a story about Noam Brown's poker bot strategy to explain equilibrium play versus game-theory exploitation.24:50–28:47 · Matt as informed peer 6/10 Reframing the Paradigm: Why RL is the Main 'Cake' Matt actively pushes back by citing Yann LeCun's 'cake' meme and Andrej Karpathy's quote about RL 'sucking supervision through a straw.' Dan defends the paradigm by emphasizing that test-time compute and reasoning breakthroughs justify the compute investment.28:47–32:41 · Matt as informed peer 5/10 Language as the Core Medium of Intelligence: Debating Rich Sutton Matt references Rich Sutton's argument on the Dwarkesh podcast asserting that LLMs are not true intelligence. Dan strongly disagrees with Sutton's premise, arguing language is the essential medium of intelligence and challenging the simple 'Bitter Lesson' view of raw scaling.32:41–35:40 · Matt as informed peer 3/10 Demystifying Test-Time Compute and Chain of Thought Matt asks about the internal mechanics of test-time compute and chain of thought. Dan demystifies the process, explaining that generated tokens function as a scratchpad enabling larger total compute per response.35:40–41:57 · Matt as informed peer 5/10 Defining Verifiable Rewards in AI Matt asks how RL can apply to non-verifiable domains and how theoretical physics applies to AI. Dan refutes the narrative of abrupt emergent behavior or grokking, arguing scaling sequences can be made smooth using physics-style simplified modeling.41:57–45:51 · Matt as informed peer 5/10 The Search for a Thermodynamics of AI Matt poses a high-level question about creating a 'thermodynamics of AI' and brings up Dan's past joke predicting 9 years to Einstein-level AI. Dan explains scaling laws as macroscopic effective theories akin to thermodynamics.45:51–48:45 · Matt as informed peer 3/10 Automating AI Research and Unlocking Fundamental Science Matt asks about timelines for self-automating AI research. Dan provides a balanced perspective, expressing optimistic excitement about using AI models to solve fundamental open questions in physics and mathematics.1:18–7:05 · Guest teaching 4/10 Understanding the Foundations of Reinforcement Learning Team Matt asks open-ended introductory questions about Dan's role at OpenAI and his academic trajectory. Dan details his work on RL foundations, scaling laws, and his theoretical physics background studying quantum chaos and black hole information processing.7:05–12:12 · Guest teaching 5/10 The Evolution of AI in Autonomous Scientific Discovery Matt demonstrates domain knowledge by referencing recent developments on Erdos problems by OpenAI, DeepMind, and Anthropic. Dan educates Matt on how OpenAI refuted a lower-bound conjecture via algebraic number theory and contrasts DeepMind's Lean auto-formalization with OpenAI's natural language proof generation.12:12–17:04 · Guest teaching 6/10 Defining Reinforcement Learning: Analogies and Core Principles Matt asks fundamental definition questions about RL, why it works, and how it fails. Dan delivers an educational breakdown using a Super Mario Bros analogy to contrast supervised learning from demonstrations with interactive reinforcement learning.17:04–24:50 · Guest teaching 5/10 Applying RL to Large Language Models and Proxy Reward Models Matt prompts Dan on RLHF history, reward proxies, AlphaGo's Move 37, and exploration vs exploitation. Dan shares a story about Noam Brown's poker bot strategy to explain equilibrium play versus game-theory exploitation.24:50–28:47 · Guest teaching 5/10 Reframing the Paradigm: Why RL is the Main 'Cake' Matt actively pushes back by citing Yann LeCun's 'cake' meme and Andrej Karpathy's quote about RL 'sucking supervision through a straw.' Dan defends the paradigm by emphasizing that test-time compute and reasoning breakthroughs justify the compute investment.28:47–32:41 · Guest teaching 6/10 Language as the Core Medium of Intelligence: Debating Rich Sutton Matt references Rich Sutton's argument on the Dwarkesh podcast asserting that LLMs are not true intelligence. Dan strongly disagrees with Sutton's premise, arguing language is the essential medium of intelligence and challenging the simple 'Bitter Lesson' view of raw scaling.32:41–35:40 · Guest teaching 5/10 Demystifying Test-Time Compute and Chain of Thought Matt asks about the internal mechanics of test-time compute and chain of thought. Dan demystifies the process, explaining that generated tokens function as a scratchpad enabling larger total compute per response.35:40–41:57 · Guest teaching 6/10 Defining Verifiable Rewards in AI Matt asks how RL can apply to non-verifiable domains and how theoretical physics applies to AI. Dan refutes the narrative of abrupt emergent behavior or grokking, arguing scaling sequences can be made smooth using physics-style simplified modeling.41:57–45:51 · Guest teaching 6/10 The Search for a Thermodynamics of AI Matt poses a high-level question about creating a 'thermodynamics of AI' and brings up Dan's past joke predicting 9 years to Einstein-level AI. Dan explains scaling laws as macroscopic effective theories akin to thermodynamics.45:51–48:45 · Guest teaching 4/10 Automating AI Research and Unlocking Fundamental Science Matt asks about timelines for self-automating AI research. Dan provides a balanced perspective, expressing optimistic excitement about using AI models to solve fundamental open questions in physics and mathematics.1:18–7:05 · Guest disagreement 1/10 Understanding the Foundations of Reinforcement Learning Team Matt asks open-ended introductory questions about Dan's role at OpenAI and his academic trajectory. Dan details his work on RL foundations, scaling laws, and his theoretical physics background studying quantum chaos and black hole information processing.7:05–12:12 · Guest disagreement 1/10 The Evolution of AI in Autonomous Scientific Discovery Matt demonstrates domain knowledge by referencing recent developments on Erdos problems by OpenAI, DeepMind, and Anthropic. Dan educates Matt on how OpenAI refuted a lower-bound conjecture via algebraic number theory and contrasts DeepMind's Lean auto-formalization with OpenAI's natural language proof generation.12:12–17:04 · Guest disagreement 1/10 Defining Reinforcement Learning: Analogies and Core Principles Matt asks fundamental definition questions about RL, why it works, and how it fails. Dan delivers an educational breakdown using a Super Mario Bros analogy to contrast supervised learning from demonstrations with interactive reinforcement learning.17:04–24:50 · Guest disagreement 1/10 Applying RL to Large Language Models and Proxy Reward Models Matt prompts Dan on RLHF history, reward proxies, AlphaGo's Move 37, and exploration vs exploitation. Dan shares a story about Noam Brown's poker bot strategy to explain equilibrium play versus game-theory exploitation.24:50–28:47 · Guest disagreement 3/10 Reframing the Paradigm: Why RL is the Main 'Cake' Matt actively pushes back by citing Yann LeCun's 'cake' meme and Andrej Karpathy's quote about RL 'sucking supervision through a straw.' Dan defends the paradigm by emphasizing that test-time compute and reasoning breakthroughs justify the compute investment.28:47–32:41 · Guest disagreement 4/10 Language as the Core Medium of Intelligence: Debating Rich Sutton Matt references Rich Sutton's argument on the Dwarkesh podcast asserting that LLMs are not true intelligence. Dan strongly disagrees with Sutton's premise, arguing language is the essential medium of intelligence and challenging the simple 'Bitter Lesson' view of raw scaling.32:41–35:40 · Guest disagreement 1/10 Demystifying Test-Time Compute and Chain of Thought Matt asks about the internal mechanics of test-time compute and chain of thought. Dan demystifies the process, explaining that generated tokens function as a scratchpad enabling larger total compute per response.35:40–41:57 · Guest disagreement 3/10 Defining Verifiable Rewards in AI Matt asks how RL can apply to non-verifiable domains and how theoretical physics applies to AI. Dan refutes the narrative of abrupt emergent behavior or grokking, arguing scaling sequences can be made smooth using physics-style simplified modeling.41:57–45:51 · Guest disagreement 2/10 The Search for a Thermodynamics of AI Matt poses a high-level question about creating a 'thermodynamics of AI' and brings up Dan's past joke predicting 9 years to Einstein-level AI. Dan explains scaling laws as macroscopic effective theories akin to thermodynamics.45:51–48:45 · Guest disagreement 1/10 Automating AI Research and Unlocking Fundamental Science Matt asks about timelines for self-automating AI research. Dan provides a balanced perspective, expressing optimistic excitement about using AI models to solve fundamental open questions in physics and mathematics.1:18–7:05 · Matt pushing back 0/10 Understanding the Foundations of Reinforcement Learning Team Matt asks open-ended introductory questions about Dan's role at OpenAI and his academic trajectory. Dan details his work on RL foundations, scaling laws, and his theoretical physics background studying quantum chaos and black hole information processing.7:05–12:12 · Matt pushing back 1/10 The Evolution of AI in Autonomous Scientific Discovery Matt demonstrates domain knowledge by referencing recent developments on Erdos problems by OpenAI, DeepMind, and Anthropic. Dan educates Matt on how OpenAI refuted a lower-bound conjecture via algebraic number theory and contrasts DeepMind's Lean auto-formalization with OpenAI's natural language proof generation.12:12–17:04 · Matt pushing back 0/10 Defining Reinforcement Learning: Analogies and Core Principles Matt asks fundamental definition questions about RL, why it works, and how it fails. Dan delivers an educational breakdown using a Super Mario Bros analogy to contrast supervised learning from demonstrations with interactive reinforcement learning.17:04–24:50 · Matt pushing back 1/10 Applying RL to Large Language Models and Proxy Reward Models Matt prompts Dan on RLHF history, reward proxies, AlphaGo's Move 37, and exploration vs exploitation. Dan shares a story about Noam Brown's poker bot strategy to explain equilibrium play versus game-theory exploitation.24:50–28:47 · Matt pushing back 4/10 Reframing the Paradigm: Why RL is the Main 'Cake' Matt actively pushes back by citing Yann LeCun's 'cake' meme and Andrej Karpathy's quote about RL 'sucking supervision through a straw.' Dan defends the paradigm by emphasizing that test-time compute and reasoning breakthroughs justify the compute investment.28:47–32:41 · Matt pushing back 2/10 Language as the Core Medium of Intelligence: Debating Rich Sutton Matt references Rich Sutton's argument on the Dwarkesh podcast asserting that LLMs are not true intelligence. Dan strongly disagrees with Sutton's premise, arguing language is the essential medium of intelligence and challenging the simple 'Bitter Lesson' view of raw scaling.32:41–35:40 · Matt pushing back 1/10 Demystifying Test-Time Compute and Chain of Thought Matt asks about the internal mechanics of test-time compute and chain of thought. Dan demystifies the process, explaining that generated tokens function as a scratchpad enabling larger total compute per response.35:40–41:57 · Matt pushing back 2/10 Defining Verifiable Rewards in AI Matt asks how RL can apply to non-verifiable domains and how theoretical physics applies to AI. Dan refutes the narrative of abrupt emergent behavior or grokking, arguing scaling sequences can be made smooth using physics-style simplified modeling.41:57–45:51 · Matt pushing back 2/10 The Search for a Thermodynamics of AI Matt poses a high-level question about creating a 'thermodynamics of AI' and brings up Dan's past joke predicting 9 years to Einstein-level AI. Dan explains scaling laws as macroscopic effective theories akin to thermodynamics.45:51–48:45 · Matt pushing back 1/10 Automating AI Research and Unlocking Fundamental Science Matt asks about timelines for self-automating AI research. Dan provides a balanced perspective, expressing optimistic excitement about using AI models to solve fundamental open questions in physics and mathematics.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 36.8% · guest 63.2%0:00 · Matt 36.8% · guest 63.2%3:00 · Matt 3.8% · guest 96.2%3:00 · Matt 3.8% · guest 96.2%6:00 · Matt 33.7% · guest 66.3%6:00 · Matt 33.7% · guest 66.3%9:00 · Matt 8.7% · guest 91.3%9:00 · Matt 8.7% · guest 91.3%12:00 · Matt 12.1% · guest 87.9%12:00 · Matt 12.1% · guest 87.9%15:00 · Matt 15.4% · guest 84.6%15:00 · Matt 15.4% · guest 84.6%18:00 · Matt 10.3% · guest 89.7%18:00 · Matt 10.3% · guest 89.7%21:00 · Matt 10.4% · guest 89.6%21:00 · Matt 10.4% · guest 89.6%24:00 · Matt 20.4% · guest 79.6%24:00 · Matt 20.4% · guest 79.6%27:00 · Matt 37.2% · guest 62.8%27:00 · Matt 37.2% · guest 62.8%30:00 · Matt 13.1% · guest 86.9%30:00 · Matt 13.1% · guest 86.9%33:00 · Matt 23.6% · guest 76.4%33:00 · Matt 23.6% · guest 76.4%36:00 · Matt 35.6% · guest 64.4%36:00 · Matt 35.6% · guest 64.4%39:00 · Matt 1.1% · guest 98.9%39:00 · Matt 1.1% · guest 98.9%42:00 · Matt 14.4% · guest 85.6%42:00 · Matt 14.4% · guest 85.6%45:00 · Matt 9.1% · guest 90.9%45:00 · Matt 9.1% · guest 90.9%48:00 · Matt 42.2% · guest 57.8%48:00 · Matt 42.2% · guest 57.8%
Sharpest disagreement ▶ 38:32 Dan rejects grokking and emergent scale narratives

Dan forcefully rejects the common industry view that AI capabilities abruptly 'grok' or emerge out of nowhere at scale, asserting that such claims reflect a failure to properly understand the underlying scaling sequence.

Hardest push from Matt ▶ 27:30 Matt challenges RL token efficiency with Karpathy's quote

Matt directly challenges Dan on the core efficiency of RL training by citing Andrej Karpathy's critique that RL sucks supervision through a straw with less than one bit of information per 10,000 tokens.

Biggest teaching moment ▶ 12:35 Dan breaks down RL vs Supervised Learning with Mario analogy

Dan provides a clear, insightful conceptual framework comparing passive observation of a video game to active trial-and-error play to illustrate the fundamental mechanics of RL.

Matt holds his own ▶ 28:47 Matt frames debate citing Rich Sutton on the Dwarkesh podcast

Matt demonstrates deep industry knowledge by directly citing Rich Sutton's recent interview on Dwarkesh Patel's podcast regarding pure RL versus hybrid LLM+RL models.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Understanding the Foundations of Reinforcement Learning Team 1410 Matt asks open-ended introductory questions about Dan's role at OpenAI and his academic trajectory. Dan details his work on RL foundations, scaling laws, and his theoretical physics background studying quantum chaos and black hole information processing.
The Evolution of AI in Autonomous Scientific Discovery 4511 Matt demonstrates domain knowledge by referencing recent developments on Erdos problems by OpenAI, DeepMind, and Anthropic. Dan educates Matt on how OpenAI refuted a lower-bound conjecture via algebraic number theory and contrasts DeepMind's Lean auto-formalization with OpenAI's natural language proof generation.
Defining Reinforcement Learning: Analogies and Core Principles 2610 Matt asks fundamental definition questions about RL, why it works, and how it fails. Dan delivers an educational breakdown using a Super Mario Bros analogy to contrast supervised learning from demonstrations with interactive reinforcement learning.
Applying RL to Large Language Models and Proxy Reward Models 4511 Matt prompts Dan on RLHF history, reward proxies, AlphaGo's Move 37, and exploration vs exploitation. Dan shares a story about Noam Brown's poker bot strategy to explain equilibrium play versus game-theory exploitation.
Reframing the Paradigm: Why RL is the Main 'Cake' 6534 Matt actively pushes back by citing Yann LeCun's 'cake' meme and Andrej Karpathy's quote about RL 'sucking supervision through a straw.' Dan defends the paradigm by emphasizing that test-time compute and reasoning breakthroughs justify the compute investment.
Language as the Core Medium of Intelligence: Debating Rich Sutton 5642 Matt references Rich Sutton's argument on the Dwarkesh podcast asserting that LLMs are not true intelligence. Dan strongly disagrees with Sutton's premise, arguing language is the essential medium of intelligence and challenging the simple 'Bitter Lesson' view of raw scaling.
Demystifying Test-Time Compute and Chain of Thought 3511 Matt asks about the internal mechanics of test-time compute and chain of thought. Dan demystifies the process, explaining that generated tokens function as a scratchpad enabling larger total compute per response.
Defining Verifiable Rewards in AI 5632 Matt asks how RL can apply to non-verifiable domains and how theoretical physics applies to AI. Dan refutes the narrative of abrupt emergent behavior or grokking, arguing scaling sequences can be made smooth using physics-style simplified modeling.
The Search for a Thermodynamics of AI 5622 Matt poses a high-level question about creating a 'thermodynamics of AI' and brings up Dan's past joke predicting 9 years to Einstein-level AI. Dan explains scaling laws as macroscopic effective theories akin to thermodynamics.
Automating AI Research and Unlocking Fundamental Science 3411 Matt asks about timelines for self-automating AI research. Dan provides a balanced perspective, expressing optimistic excitement about using AI models to solve fundamental open questions in physics and mathematics.

Statements from this episode (17)

Disclosure
OpenAI studied reinforcement learning reasoning models internally long before releasing o1
“So before we released a one and thinking reasoning models, we were studying this internally”
Dan Roberts Jun 4, 2026 ▶ 1:54
Disclosure
OpenAI's RL team researches models two generations ahead of current releases
“And to do that, we need to make thinking models and some, somewhere along the way, we interact with that process. Usually at the earlier stage for models, you know, not the next model, but things that are like the next model or the next, next model.”
Dan Roberts Jun 4, 2026 ▶ 2:54
Prediction Not checkable as stated
AI systems evolving into fully fledged scientists will be a gradual transition
“The, there's no sharp point, or I don't think there will be a sharp point where we'll say that systems didn't, weren't able to be useful for scientific, the scientific process to their fully fledged scientists. There'll be sort of a gradual shift.”
Dan Roberts Jun 4, 2026 ▶ 7:32
Assertion Supported
ChatGPT disproved an Erdős conjecture using cross-disciplinary mathematical reasoning
“The big result was that this conjecture of this lower bound for the number of pairs that you can make is, is false. Not only is it false, it was false due to a really interesting connection to another field of mathematics.”
Dan Roberts Jun 4, 2026 ▶ 9:57
Disclosure
OpenAI publishes most math research results using informal language settings
“Most of our results that we publicize, as far as I can think, are all in the informal setting.”
Dan Roberts Jun 4, 2026 ▶ 11:57
Disclosure
OpenAI plans to increasingly rely on reinforcement learning to scale intelligence
“When you have a lot of compute, you want to turn that compute into intelligence in a way that's useful, and RL is one way of doing it, and we just started doing it then, and we're going to do a lot more of it now.”
Dan Roberts Jun 4, 2026 ▶ 25:35
Insight
Roberts: Powerful pre-trained models are necessary for effective RL and reasoning
“If you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to like think at use test time compute to for instance, solve, solve math problems that it wouldn't otherwise be able to do.”
Dan Roberts Jun 4, 2026 ▶ 27:15
Insight
Pre-training models on language before reinforcement learning is the correct architecture
“Having the model have a prior of language and being able to like, think in language and then train on top of that, that seems like clearly the right. The right thing to do.”
Dan Roberts Jun 4, 2026 ▶ 31:16
Insight
Dan Roberts: AI scaling requires novel algorithms, not just pure compute
“It's not that scale is all you need. You need to also have good ideas to guide the scaling.”
Dan Roberts Jun 4, 2026 ▶ 31:56
Assertion Not checkable as stated
Combining reinforcement learning with pre-training outperforms scaling pre-training alone
“If you were just trying to scale pre-training, you wouldn't get anywhere near as far as also trying to scale RL on top of pre-training, which is what we do now.”
Dan Roberts Jun 4, 2026 ▶ 32:05
Insight
Roberts: AI models improve performance by generating running thought tokens in language
“The natural way it thinks is in language. It's a language model, and so that's sort of this key insight that, that you can cause it to do better just by producing a thought process in, in, in token space, in, in language.”
Dan Roberts Jun 4, 2026 ▶ 34:01
Prediction Not checkable as stated
OpenAI will release reinforcement learning products for consulting, banking, and legal
“I definitely think OpenAI will have amazing products that will be relevant in those domains, and some amount of RL will play a role in there.”
Dan Roberts Jun 4, 2026 ▶ 37:01
Insight
Dan Roberts: AI scaling laws should be analyzed big-to-small
“The way to think about scaling and say, scaling laws is not small to big, but big to small.”
Dan Roberts Jun 4, 2026 ▶ 38:38
Insight
Dan Roberts: AI models do not experience discontinuous emergence or grokking
“You have these crazy, huge systems that have all sorts of interesting phenomena, and, you know, if you think about it the right way, they don't grok. There's just this nice continuity.”
Dan Roberts Jun 4, 2026 ▶ 41:47
Prediction Not checkable as stated
Continuous AI improvements render multi-year autonomous agent runs highly inefficient
“In, in general, we're not just going to like set up a system and let it think autonomously for eight years, if anything, because like the system's eight years. After will be so much more powerful that it probably doesn't make sense to let a system think for a …”
Dan Roberts Jun 4, 2026 ▶ 44:03
Assertion Not checkable as stated
Current AI models lack research taste and problem formulation ability
“There's part of the scientific process. I think that the models haven't been imbued with yet. And I'm sure people are thinking about how to do that. You know, like what, trying to get to what is the right question as opposed to here's a well-defined thing and …”
Dan Roberts Jun 4, 2026 ▶ 44:53
Prediction Not checkable as stated
Dan Roberts expects more AI-driven math and science breakthroughs within six months
“For the next six months. Like, I think we'll see more of these sorts of math and science breakthroughs.”
Dan Roberts Jun 4, 2026 ▶ 47:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.