reinforcement learning
also referred to as: rl
15 statements across 12 episodes · 11 bullish · 1 bearish · 11 people on the record · first statement Mar 19, 2025 by Scott Wu · across every show →
Everything said about reinforcement learning, oldest first
Mar 19, 2025 bullish
Mar 24, 2025 bullish
Foody: Human data market will grow dramatically as RL requires ubiquitous evals
“My broader take on the human data market is that It's going to grow dramatically because, so we're getting to the point where RL is so effective that you can create almost any eval and it will be able to solve that eval. And so the barrier to applying AI throu…”
Apr 25, 2025 positive
Kilpatrick: Gemini 2.5 Pro Relied on Pre-Training Innovations, Not Just RL
“But I think if you look at like a, an example of this in practice, like 2.5 pro is actually an example where it wasn't just like RL scaling that made that model better. Yes, RL was part of the story, but like there was also a bunch of pre-training innovation a…”
Apr 25, 2025 bullish
Kilpatrick: Pre-Training Isn't Dead; Gains Multiply Through Post-Training and RL
“And this is why, like, I don't subscribe to the, like pre-training is, you know, dead and all that stuff, because the more work that you can do at the pre-training level, those capabilities, as you do post-training and as you give the models RL capability, it'…”
Apr 25, 2025 bullish
Apr 26, 2025 bullish
Will Brown: OpenAI o3 succeeds because reinforcement learning enables tool use
“The reason O-three is good is because it's trained to use tools. The way you train a model to use the right tool for the job is reinforcement learning. And they've said as much, like, deep research, reinforcement learning.”
Jun 7, 2025 positive
Jun 7, 2025 bullish
Jun 7, 2025 bullish
Jun 14, 2025 neutral
Jul 5, 2025 bullish
Depue: RL task specification scales with compute far better than pre-training data
“Because like tokens, you kind of hit diminishing returns pretty quick in terms of just like more internet texts or more like human written math solution examples. And it doesn't scale super well, but the thing, the nice thing about the task specification is yo…”
Jul 8, 2025 negative
Patel: AI automation is bottlenecked by task depth, not task breadth
“I think the bigger problem is not just the width or the width of the pool, how many different tasks you have to RL on, but it's a depth in the sense that a job doesn't involve doing a thousand different five minute tasks individually. It's the fact that you're…”
Jul 26, 2025 neutral
Oct 28, 2025 positive
Feb 6, 2026 neutral