Dan Roberts, lead of the Foundations of Reinforcement Learning team at OpenAI, describes his team's focus on long-term R&D rather than immediate model deployments.
“And to do that, we need to make thinking models and some, somewhere along the way, we interact with that process. Usually at the earlier stage for models, you know, not the next model, but things that are like the next model or the next, next model.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dan Roberts
Insight
Dan Roberts: AI scaling requires novel algorithms, not just pure compute
“It's not that scale is all you need. You need to also have good ideas to guide the scaling.”
Dan RobertsJun 4, 2026▶ 31:56OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Insight
Dan Roberts: AI models do not experience discontinuous emergence or grokking
“You have these crazy, huge systems that have all sorts of interesting phenomena, and, you know, if you think about it the right way, they don't grok. There's just this nice continuity.”
Dan RobertsJun 4, 2026▶ 41:47OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Insight
Pre-training models on language before reinforcement learning is the correct architecture
“Having the model have a prior of language and being able to like, think in language and then train on top of that, that seems like clearly the right. The right thing to do.”
Dan RobertsJun 4, 2026▶ 31:16OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
AssertionNot checkable as stated
Combining reinforcement learning with pre-training outperforms scaling pre-training alone
“If you were just trying to scale pre-training, you wouldn't get anywhere near as far as also trying to scale RL on top of pre-training, which is what we do now.”
Dan RobertsJun 4, 2026▶ 32:05OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
AssertionSupported
ChatGPT disproved an Erdős conjecture using cross-disciplinary mathematical reasoning
“The big result was that this conjecture of this lower bound for the number of pairs that you can make is, is false. Not only is it false, it was false due to a really interesting connection to another field of mathematics.”
Dan RobertsJun 4, 2026▶ 9:57OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Disclosure
OpenAI plans to increasingly rely on reinforcement learning to scale intelligence
“When you have a lot of compute, you want to turn that compute into intelligence in a way that's useful, and RL is one way of doing it, and we just started doing it then, and we're going to do a lot more of it now.”
Dan RobertsJun 4, 2026▶ 25:35OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.