reinforcement learning
20 statements across 16 episodes · 7 bullish · 6 bearish · 14 people on the record · first statement Dec 10, 2021 by Yann LeCun · across every show →
Everything said about reinforcement learning, oldest first
Dec 10, 2021 bearish
LeCun: Reinforcement learning cannot safely or efficiently train self-driving cars
“If we were to use, let's say, reinforcement learning to train a self-driving car to drive itself,
It would have to drive itself for millions of hours and cause, you know, until thousands of accidents and destroy itself multiple times before it learns to drive …”
Dec 10, 2021 bearish
LeCun: Supervised and reinforcement learning do not reflect biological learning
“The type of learning that we are currently able to reproduce in machine, which is supervised learning and reinforcement learning do not seem to Reflect what we observe in humans and animals. There is another type of learning, another paradigm of learning that …”
Feb 23, 2023 neutral
Lemoine: Reinforcement learning transforms language models into goal-oriented systems
“The subsequent fine-tuning, and especially once you add reinforcement learning, it's no longer just trying to predict the next token in a stream of text. Specifically in the reinforcement learning paradigm, it's trying to accomplish a goal.”
Feb 23, 2023 negative
Aug 3, 2023 bullish
Wood: AI agents and output vetting will become as important as models
“So there's a lot of focus on models today, but those models are going to remain important. But over time, there's going to be additional capabilities like agents, like reinforcement learning, like the ability to be able to understand and vet The responses that…”
Sep 14, 2023 bullish
Murdock: Merging LLMs and reinforcement learning will drive major AI breakthroughs
“I don't know exactly what the breakthrough will be, but I'm really excited about the union of these LLMs plus reinforcement learning. And I'm excited about that because I think there's a lot more to come from reinforcement learning. And I know at Google D mine…”
May 9, 2024 bullish
Clark: Scaling RL compute will unlock major new AI capabilities
“Everyone is trying to figure out how they can spend more and more of their compute on reinforcement learning, because I think everyone has this intuition that the more RL you add, the more sophisticated you're going to be able to make these things, and a lot o…”
May 15, 2024
Long-horizon reinforcement learning for AI agents is constrained by sparse rewards
“People have been talking about long horizon RL, which is the training method. You need to get something like this, where you go tell it to do something and then you reward it at the end for having achieved that outcome. But the difficulty with those kinds of a…”
Mar 28, 2025 bullish
Hendrycks: RL-based reasoning models are improving faster than pre-training did
“That is separate from the new reasoning paradigm that has emerged in the past year which is where you train models to on math and coding types of questions with reinforcement learning, and that has a very steep slope, and I don't see any signs of that slowing …”
Apr 23, 2025
Patel: Human labeling is unscalable, forcing reliance on synthetic AI data
“Using humans to train models is just so expensive, right? So then there's the magic of sort of reinforcement learning and other synthetic data technologies, right? Where the model is helping teach the model, right? So you have many models in, in, in a sort of,…”
Jun 18, 2025 bearish
Patel: Delayed Reward Signals Will Slow AI Progress on Long-Horizon Tasks
“Now if we're getting to the world where you got to like do a project for seven hours and then at the end of those seven hours, then we tell you, Hey, did you did you get this right? Then like the progress just goes on a bunch. Cause you've gone from like getti…”
Jun 18, 2025 bearish
Jul 30, 2025 bullish
Jan 16, 2026 neutral
Mensch: AI customization techniques will be abstracted away for enterprises
“I do expect the part of the software in those deployment to increase. So the amount of the way customization occurs today with fine tuning, reinforcement learning, this kind of things, this is going to be abstracted away from the enterprise buyer because it's …”
Mar 9, 2026 neutral
Kantrowitz: Scale AI shifted majority of training to reinforcement learning
“I just did the story for with about scale AI saying that the majority of their training has moved to reinforcement learning where they train models to act in specific environments like filling out forms, and then they baked those capabilities back Into the mod…”
Apr 27, 2026 neutral
Kantrowitz: Scale AI now does most of its training via reinforcement learning
“Scale AI, Alexander Wang's company, they told me recently that most of the training that they're doing is reinforcement learning, where you build environments for the bots and they go and they try to figure out what to do.”
Jul 7, 2026 positive
Kutylowski: Task-focused reinforcement learning outperforms broad multi-task training in specific domains
“And if you run this reinforcement learning step on too many different tasks, the model will be able to do all of that. But once again, it's going to be very, very, very broad. And if you focus on making sure that the model understands and knows that it needs t…”
Jul 8, 2026 bullish
Sep 7, 2026 neutral
Sep 7, 2026 bearish
Kantrowitz: Reinforcement learning layers add ruthlessness to AI models
“What we have now is that the reinforcement learning type of AI technology has been put on top of the self supervised learning to get these AI models working better, which has added a level of ruthlessness to them. Because one of the things we know about RL is …”