reinforcement learning
also referred to as: rl
12 statements across 9 episodes · 6 bullish · 2 bearish · 9 people on the record · first statement May 17, 2017 by Wojciech Zaremba · across every show →
Everything said about reinforcement learning, oldest first
May 17, 2017 negative
Zaremba: Reinforcement learning struggles in reality due to reward and reset assumptions
“The assumption underlying reinforcement learning Is that then there is some environment, and environment, you are an agent, and you are acting in environment by executing actions and getting rewards from the environment. And the rewards might be taught as, let…”
May 17, 2017 neutral
Zaremba: RL models need three years of real-time play to learn games
“Well, for instance, in terms of real-time execution, it takes something around three, three, three years of play to learn to play simple games . I mean it, it can be hugely parallelized, therefore it takes a few days to train it on current computers”
Jul 21, 2017 neutral
Nov 8, 2017 positive
Brockman: AI can discover non-obvious strategies that transfer to humans
“The bot had, like, taught him the strategy that he could use against a human and I think that was, like, very interesting and a good example of the kinds of things that you can get out of these systems, that they can discover these very, sort of, you know, Non…”
Nov 8, 2024 neutral
Jun 25, 2025 bullish
Nadella: AI labs will shift toward integrated reasoning models within a year
“What is the pre-training to RL the end-to-end training loop that's the next, you know, big sample? That I think is also what I think will happen in the next year. So I would say, if that is another scaling law breakthrough, because we will be, like, if you sor…”
Jul 22, 2025
Jul 22, 2025 positive
Finn: Reinforcement Learning Outperforms Pure Imitation Learning in Robotics
“I think that reinforcement learning can play a very large role in it actually, in post-training. I think that online data from the robots which reinforcement learning allows you to use, Can allow robots to have a much higher success rate and also be faster tha…”
Jul 29, 2025 positive
Kaplan: Compute scaling drives AI progress more than researcher cleverness
“Basically you can Scale up the compute in both pre-training and RL and get better and better performance. And I think that's sort of the fundamental thing that is driving AI progress. It's not that AI researchers are really smart or they suddenly got smart. It…”
Jul 29, 2025 positive
Sep 30, 2025 bullish