reinforcement learning

also referred to as: rl

9 statements across 5 episodes · 5 bullish · 2 bearish · 5 people on the record · first statement Aug 21, 2017 by Daniel Gross · across every show →

Everything said about reinforcement learning, oldest first

Aug 21, 2017 bullish
Prediction Not checkable as stated
Gross: Breakthroughs in reinforcement learning will transition to market quickly
“Because when something really big happens with, say, reinforcement learning, which is an area mostly attributed to research today, it's going to happen really quickly.”
Daniel Gross Aug 21, 2017 ▶ 12:25 20VC: YC's Daniel Gross on How YC Can Democratise AI & Reduce Incumbency Advantages, Why ML Enabled Software Will Eat The Software That Ate The World & Whether AI Will Produce Independent Companies or Be Technology within Incumbents
Feb 20, 2025 bullish
Prediction Not checkable as stated
Hiremath: Shift to reinforcement learning will create domain-specific reasoning AI
“The whole market is shifting to reinforcement learning, right? You're already seeing this with O-one, O-three, the deep seek models. And as a result, I think we're going to see really, really powerful models in specific domains that can reason extremely well.”
Adarsh Hiremath Feb 20, 2025 ▶ 16:43 Adarsh Hiremath @ Mercor: The Fastest Growing Startup in Silicon Valley | E1261 · 20VC with Harry Stebbings
Jul 18, 2025 positive
Opinion
Wu: Reinforcement learning is AI's biggest breakthrough of the last 18 months
“You know, it's RL, for example, is I would say the biggest breakthrough of the last year and a half call it.”
Scott Wu Jul 18, 2025 ▶ 15:07 Cognition CEO Scott Wu on Acquiring Windsurf: The Process, The Deal, The Rationale · 20VC with Harry Stebbings
Nov 3, 2025 negative
Insight
Pineau: Using reinforcement learning to teach AI social behavior remains unsolved
“RL, to shape the behavior of models, to get them to be social creatures, that we have no idea how to do. I mean, I don't know if you have children, but like shaping their behaviors, you know, the number of times you can repeat the same thing, and still they do…”
Joelle Pineau Nov 3, 2025 ▶ 6:21 Cohere's Chief AI Officer, Joelle Pineau: Why Scaling Laws Will Continue & Future of Synthetic Data · 20VC with Harry Stebbings
Nov 3, 2025 bearish
Opinion
Pineau: Out-of-the-box reinforcement learning will not deliver AGI
“Now, you know, where we're maybe getting a little bit ahead is thinking that just RL out of the box is gonna give us AGI. That part, a lot less so. You know, if you look at the curve of progress, RL is terribly inefficient, and so the amount of signal you need…”
Joelle Pineau Nov 3, 2025 ▶ 3:24 Cohere's Chief AI Officer, Joelle Pineau: Why Scaling Laws Will Continue & Future of Synthetic Data · 20VC with Harry Stebbings
Nov 3, 2025 bullish
Opinion
Pineau: Reinforcement learning is fundamental to AI and will not disappear
“Oh, I'm still super bullish on RL in that, like, the concept itself is so fundamental. You know, this idea of training through a system of rewards, of indicating what's valuable and what's not valuable through numerical values, like, that is so fundamental. It…”
Joelle Pineau Nov 3, 2025 ▶ 3:05 Cohere's Chief AI Officer, Joelle Pineau: Why Scaling Laws Will Continue & Future of Synthetic Data · 20VC with Harry Stebbings
Nov 3, 2025 positive
Assertion Supported
Pineau: RL costs are falling where clear reward functions exist
“It's coming down, especially in domains where we have good reward functions.”
Joelle Pineau Nov 3, 2025 ▶ 5:30 Cohere's Chief AI Officer, Joelle Pineau: Why Scaling Laws Will Continue & Future of Synthetic Data · 20VC with Harry Stebbings
Dec 1, 2025
Assertion Not checkable as stated
Siddharth: Reinforcement learning is the dominant paradigm for training AI agents
“Today, the dominant paradigm is reinforcement learning.”
Jonathan Siddharth Dec 1, 2025 ▶ 5:14 Turing CEO Jonathan Siddharth: Who Wins in Data Labelling & Why 99% of Knowledge Work Will Disappear · 20VC with Harry Stebbings
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.