reinforcement learning

also referred to as: rl

15 statements across 12 episodes · 11 bullish · 1 bearish · 11 people on the record · first statement Mar 19, 2025 by Scott Wu · across every show →

Everything said about reinforcement learning, oldest first

Mar 19, 2025 bullish
Opinion
Wu: AI progress is primarily driven by reinforcement learning
“You know, I think the at a high level, you know, I think all the continued progress that we're seeing everywhere in AI actually is, is primarily due to Essentially like reinforcement learning, RL, that's really working.”
Scott Wu Mar 19, 2025 ▶ 11:44 Scott Wu on Cognition, Devin, and the Evolution of AI Agents.
Mar 24, 2025 bullish
Prediction Not checkable as stated
Foody: Human data market will grow dramatically as RL requires ubiquitous evals
“My broader take on the human data market is that It's going to grow dramatically because, so we're getting to the point where RL is so effective that you can create almost any eval and it will be able to solve that eval. And so the barrier to applying AI throu…”
Brendan Foody Mar 24, 2025 ▶ 6:38 Brendan Foody on using AI to predict job performance and scaling from $1M to $100M in 11 months
Apr 25, 2025 positive
Assertion Supported
Kilpatrick: Gemini 2.5 Pro Relied on Pre-Training Innovations, Not Just RL
“But I think if you look at like a, an example of this in practice, like 2.5 pro is actually an example where it wasn't just like RL scaling that made that model better. Yes, RL was part of the story, but like there was also a bunch of pre-training innovation a…”
Logan Kilpatrick Apr 25, 2025 ▶ 12:21 Google's AI Comeback in Their Own Words - Logan Kilpatrick
Apr 25, 2025 bullish
Opinion
Kilpatrick: Pre-Training Isn't Dead; Gains Multiply Through Post-Training and RL
“And this is why, like, I don't subscribe to the, like pre-training is, you know, dead and all that stuff, because the more work that you can do at the pre-training level, those capabilities, as you do post-training and as you give the models RL capability, it'…”
Logan Kilpatrick Apr 25, 2025 ▶ 12:50 Google's AI Comeback in Their Own Words - Logan Kilpatrick
Apr 25, 2025 bullish
Prediction Not checkable as stated
Srinivas: Most AI foundation model investment will go toward reinforcement learning
“I think RL is the place where most investments are gonna go to especially with models like O-three that are able to do tool calls pretty natively rather than being prompt engineered to do that”
Aravind Srinivas Apr 25, 2025 ▶ 27:35 Perplexity Founder Explains What Comes Next - Aravind Srinivas on TBPN April 23rd
Apr 26, 2025 bullish
Insight
Will Brown: OpenAI o3 succeeds because reinforcement learning enables tool use
“The reason O-three is good is because it's trained to use tools. The way you train a model to use the right tool for the job is reinforcement learning. And they've said as much, like, deep research, reinforcement learning.”
Will Brown Apr 26, 2025 ▶ 7:33 Why Humor Is the True Test of AI Intelligence | Will Brown on TBPN
Jun 7, 2025 positive
Assertion Not checkable as stated
Douglas: 10x RL compute increases still deliver linear performance gains
“We're still seeing these huge gains where you go, you know, 10 X compute increase in RL. We still getting like very distinct linear gains for that.”
Sholto Douglas Jun 7, 2025 ▶ 1:30:12 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Jun 7, 2025 bullish
Insight
Mark Chen: Compute Can Scale Heavily into RL Given Right Levers
“I think, like, if you find the right levers, you can really pump a lot of compute into RL as well as pre-training.”
Mark Chen Jun 7, 2025 ▶ 1:28:16 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Jun 7, 2025 bullish
Prediction Not checkable as stated
Douglas: AI will use long-horizon reward signals within two years
“We've got very little rewards right now, but pretty quickly over the next year or two, you're going to start to see much more meaningful and long horizon rewards.”
Sholto Douglas Jun 7, 2025 ▶ 1:35:46 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Jun 14, 2025 neutral
Assertion Not checkable as stated
Much AI reinforcement learning research focuses on difficult math problems
“A lot of the reinforcement learning is now done on like extremely hard math problems. And so you might need to go to PhD programs to recruit.”
John Coogan Jun 14, 2025 ▶ 55:27 Weekly Recap: Apple WWDC, Andrew Huberman, YC Demo Day, Intern Phone Assembly, $14B Meta x Scale
Jul 5, 2025 bullish
Insight
Depue: RL task specification scales with compute far better than pre-training data
“Because like tokens, you kind of hit diminishing returns pretty quick in terms of just like more internet texts or more like human written math solution examples. And it doesn't scale super well, but the thing, the nice thing about the task specification is yo…”
Will Depue Jul 5, 2025 ▶ 3:05:03 Weekly Recap: The Soham Parekh Drama, Meta Attacks OpenAI, Trump's Mega Bill
Jul 8, 2025 negative
Insight
Patel: AI automation is bottlenecked by task depth, not task breadth
“I think the bigger problem is not just the width or the width of the pool, how many different tasks you have to RL on, but it's a depth in the sense that a job doesn't involve doing a thousand different five minute tasks individually. It's the fact that you're…”
Dwarkesh Patel Jul 8, 2025 ▶ 10:01 DWARKESH PATEL on the Biggest Problem of AI
Jul 26, 2025 neutral
Insight
Coogan: Task-specific RL creates models that fail to generalize
“There's a ton of situations where you can go and RL on a specific task and it gets really, really good at it, but then you try and get it to do anything else and it's not that great.”
John Coogan Jul 26, 2025 ▶ 14:35 Weekly Recap: Casey Neistat, OpenAI Cracks Math, The Future of ChatGPT, Apple x F1, Intel Layoffs
Oct 28, 2025 positive
Insight
Nadella: Pre-training Remains More Efficient Than RL Due to Amortization
“Pre-training is a more efficient form of training. Because you can advertise it.”
Satya Nadella Oct 28, 2025 ▶ 11:33 Microsoft CEO Satya Nadella Live on TBPN
Feb 6, 2026 neutral
Insight
Coogan: Git commit history makes reinforcement learning uniquely effective for coding
“Long context reinforcement learning has been very, very successful in the coding world because Git has a complete history of every line of code that's been written, every comment, why it happened.”
John Coogan Feb 6, 2026 ▶ 18:17 Anthropic’s Trust Nuke, OpenAI’s new releases, Google Claims AI Crown | Diet TBPN
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.