reinforcement learning

also referred to as: rl

12 statements across 9 episodes · 6 bullish · 2 bearish · 9 people on the record · first statement May 17, 2017 by Wojciech Zaremba · across every show →

Everything said about reinforcement learning, oldest first

May 17, 2017 negative
Insight
Zaremba: Reinforcement learning struggles in reality due to reward and reset assumptions
“The assumption underlying reinforcement learning Is that then there is some environment, and environment, you are an agent, and you are acting in environment by executing actions and getting rewards from the environment. And the rewards might be taught as, let…”
Wojciech Zaremba May 17, 2017 ▶ 7:12 An AI Primer with Wojciech Zaremba · Y Combinator
May 17, 2017 neutral
Assertion Partly supported
Zaremba: RL models need three years of real-time play to learn games
“Well, for instance, in terms of real-time execution, it takes something around three, three, three years of play to learn to play simple games . I mean it, it can be hugely parallelized, therefore it takes a few days to train it on current computers”
Wojciech Zaremba May 17, 2017 ▶ 6:24 An AI Primer with Wojciech Zaremba · Y Combinator
Jul 21, 2017 neutral
Insight
Eck: RL is slower to train than GANs but offers more flexibility
“So another way to do this is to use reinforcement learning. Yeah. And it, it's slower to train, because all you have is a single number, scalar reward, instead of this whole gradient flowing. Than GANs, but it also is more flexible.”
Doug Eck Jul 21, 2017 ▶ 32:54 Making Music and Art Through Machine Learning - Doug Eck of Magenta · Y Combinator
Nov 8, 2017 positive
Insight
Brockman: AI can discover non-obvious strategies that transfer to humans
“The bot had, like, taught him the strategy that he could use against a human and I think that was, like, very interesting and a good example of the kinds of things that you can get out of these systems, that they can discover these very, sort of, you know, Non…”
Greg Brockman Nov 8, 2017 ▶ 48:21 Building Dota Bots That Beat Pros - OpenAI's Greg Brockman, Szymon Sidor, and Sam Altman · Y Combinator
Nov 8, 2024 neutral
Disclosure
Altman: OpenAI's original goals included never exceeding 120 people
“I believe there were three goals for the For the effort at the time. It was like, figure out how to do unsupervised learning, solve RL, and never get more than a 120 people. Missed on the third one.”
Sam Altman Nov 8, 2024 ▶ 16:51 How To Build The Future: Sam Altman · Y Combinator
Jun 25, 2025 bullish
Prediction Not checkable as stated
Nadella: AI labs will shift toward integrated reasoning models within a year
“What is the pre-training to RL the end-to-end training loop that's the next, you know, big sample? That I think is also what I think will happen in the next year. So I would say, if that is another scaling law breakthrough, because we will be, like, if you sor…”
Satya Nadella Jun 25, 2025 ▶ 16:48 Satya Nadella: Microsoft's AI Bets, Hyperscaling, Quantum Computing Breakthroughs · Y Combinator
Jul 22, 2025
Insight
Chelsea Finn: Synthetic data's robotic analog is RL, not simulation
“I think that the analog of synthetic data in language models is actually not necessarily simulation in robotics, but closer to something like reinforcement learning.”
Chelsea Finn Jul 22, 2025 ▶ 41:47 Chelsea Finn: Building Robots That Can Do Anything · Y Combinator
Jul 22, 2025 positive
Insight
Finn: Reinforcement Learning Outperforms Pure Imitation Learning in Robotics
“I think that reinforcement learning can play a very large role in it actually, in post-training. I think that online data from the robots which reinforcement learning allows you to use, Can allow robots to have a much higher success rate and also be faster tha…”
Chelsea Finn Jul 22, 2025 ▶ 32:16 Chelsea Finn: Building Robots That Can Do Anything · Y Combinator
Jul 29, 2025 positive
Insight
Kaplan: Compute scaling drives AI progress more than researcher cleverness
“Basically you can Scale up the compute in both pre-training and RL and get better and better performance. And I think that's sort of the fundamental thing that is driving AI progress. It's not that AI researchers are really smart or they suddenly got smart. It…”
Jared Kaplan Jul 29, 2025 ▶ 7:54 Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan · Y Combinator
Jul 29, 2025 positive
Assertion Supported
Kaplan: Scaling laws apply to reinforcement learning in AI training
“You can see scaling laws in the reinforcement learning phase of AI training.”
Jared Kaplan Jul 29, 2025 ▶ 6:15 Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan · Y Combinator
Sep 30, 2025 bullish
Insight
Joseph: Reinforcement learning exhibits scaling laws where compute yields better models
“You can get pretty big wins from RL. You sort of have another set of scaling laws. It's like you put more and more compute into RL, you can get better and better models out of that.”
Nick Joseph Sep 30, 2025 ▶ 29:15 Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI · Y Combinator
Jul 29, 2026 negative
Assertion Not checkable as stated
Chaubard: CPU-based environment simulation bottlenecks on-policy reinforcement learning rollouts
“The, it's amazing how much of simulators, when you call environment.step, is still run on the CPU, and so that's usually the bottleneck for a lot of your on-policy rollouts”
Francois Chaubard Jul 29, 2026 ▶ 2:50 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.