reinforcement learning
22 statements across 18 episodes · 16 bullish · 1 bearish · 18 people on the record · first statement Jan 2, 2019 by Murray Shanahan · across every show →
Everything said about reinforcement learning, oldest first
Jan 2, 2019 positive
Shanahan: Reinforcement learning progress does not require massive datasets
“Actually, DeepMind are another example of the same thing, because if you want to apply reinforcement learning to games, and that's enabled them to make some quite fundamental sort of progress, you don't need vast amounts of data either.”
Jan 2, 2019 positive
Schuler: Rich Sutton's textbook is computer science's most influential publication
“According to the Alan Soufre AI Semantics Scholar, he's the most highly cited researcher in reinforcement learning, 11th most influential researcher in all of competing science. And his textbook on reinforcement learning was ranked as the single most influenti…”
Jan 2, 2019 bullish
Chen: AlphaGo Zero proves reinforcement learning succeeds without labeled data
“Right, where there's so much momentum right now that says, basically, we're one labeled data set away from glory. Right, and this result basically shows you, wow, there's a lot of mileage that you can get out of reinforcement learning where there's no data, no…”
Feb 28, 2025 negative
Ayrey: Filtering API keys from training risks degrading AI data science skills
“So if we do our reinforcement learning and we skew it towards code snippets that's generating that don't have API keys, inadvertently, we may be training this thing to behave less like a data scientist. And then we lose the entire discipline of data science in…”
Mar 5, 2025 bullish
Mascorro: DeepSeek-R1 proved reinforcement learning improves models without human feedback
“And I think the big thing in, in R-one, or generally with these reasoning models is, We were doing before there was a human in the loop always, right? Like when we have this SFT training and these other techniques that we're doing after like RLHF and having R …”
May 29, 2025 positive
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Jul 23, 2025 positive
Jul 23, 2025 positive
Caldwell: Humans are poorly suited for multivariable plant optimization compared to RL
“As you like build more of that large scale infrastructure and kind of like see the problems that humans have to solve on a daily basis, it becomes like pretty obvious that these are problems that humans don't, aren't best positioned to solve. Like these are la…”
Jul 28, 2025
Martin Casado: Domain-specific reinforcement learning requires a plurality of models
“But as soon as you're doing RL, where you're training it in a specific domain with a specific verifier, you're likely losing, you know, other areas. So you make it really, really good at playing chess. It's going to be less good at, you know, something else li…”
Aug 8, 2025
Aug 8, 2025 positive
Fulford: RL breakthroughs in math and coding unlocked functional AI agents
“When we saw the reinforcement learning algorithm working really well on math and physics problems and coding problems, It became pretty clear, like, just from reading through the chain of thought, like, okay, this thing's actually, like, thinking and reasoning…”
Sep 25, 2025 positive
Pachocki: AI reward modeling will evolve toward simpler, human-like learning
“I expect this will evolve quite rapidly. I expect it will become simpler, right? Like I think, you know, maybe like two years ago, we would have been talking about like, what is the right way to craft my fine tuning data set? And I don't think we are like at t…”
Sep 30, 2025 positive
Fedus: High-compute reinforcement learning is essential for AI tool use
“High compute reinforcement learning is really effective. This is how you should think about the strategies it's using. This is how you create effective tool using towards those problems, and this is how you optimize it effectively.”
Oct 14, 2025 bullish
Oct 23, 2025 positive
Oct 23, 2025 bullish
Masad: Reinforcement learning paired with TreeSearch still has substantial runway
“I think the breakthroughs in RL are incredibly exciting, but we also knew about them now for like over 10 years where you marry generative systems with TreeSearch and things like that. But there's a lot more to go there”
Oct 23, 2025 positive
Andreessen: AI labs hire mathematicians and coders for reinforcement learning data
“Foundation model companies are, in some cases, they are hiring, they're actually hiring human experts. To generate new training data. So they're actually hiring mathematicians and physicists and coders to basically sit, and, you know, they're hiring human prog…”
Nov 5, 2025
Kuyda: OpenAI temporarily abandoned language models for video game agents
“And very quickly they stopped working on language models, and we were very upset because we really wanted to continue going there, but they didn't want to talk about any language models because no one was really working on them and that may have made us feel v…”
Nov 17, 2025
Shear: In AI agents, care equates to loss and reward correlation
“Care is basically like reward. Like how much does this state correlate with survival? How much does this state correlate with your inclusive, your full inclusive reproductive fitness? For a somewhat thing that learns evolutionarily, or for a reinforcement lear…”
May 13, 2026 positive
Aug 26, 2026 bullish
Acharya: Domain-specialized open models outperform general models through RL and reasoning traces
“If you actually have a problem that you can specialize the model around with your reasoning traces, You can start to create this compounding advantage in your domain for your customer base, where you're able to kind of shape the intelligence to be better than …”
Sep 1, 2026
Litt: Reinforcement learning struggles to reward intermediate mathematical theory building
“I think what is definitely true is that, like, the skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one. So it might be harder, you know I guess you can try, you can tell it, you kno…”