reinforcement learning

20 statements across 16 episodes · 7 bullish · 6 bearish · 14 people on the record · first statement Dec 10, 2021 by Yann LeCun · across every show →

Everything said about reinforcement learning, oldest first

Dec 10, 2021 bearish
What-if
LeCun: Reinforcement learning cannot safely or efficiently train self-driving cars
“If we were to use, let's say, reinforcement learning to train a self-driving car to drive itself, It would have to drive itself for millions of hours and cause, you know, until thousands of accidents and destroy itself multiple times before it learns to drive …”
Yann LeCun Dec 10, 2021 ▶ 11:58 Daniel Kahneman and Yann LeCun: How To Get AI To Think Like Humans (Full Episode)
Dec 10, 2021 bearish
Insight
LeCun: Supervised and reinforcement learning do not reflect biological learning
“The type of learning that we are currently able to reproduce in machine, which is supervised learning and reinforcement learning do not seem to Reflect what we observe in humans and animals. There is another type of learning, another paradigm of learning that …”
Yann LeCun Dec 10, 2021 ▶ 5:04 Daniel Kahneman and Yann LeCun: How To Get AI To Think Like Humans (Full Episode)
Feb 23, 2023 neutral
Insight
Lemoine: Reinforcement learning transforms language models into goal-oriented systems
“The subsequent fine-tuning, and especially once you add reinforcement learning, it's no longer just trying to predict the next token in a stream of text. Specifically in the reinforcement learning paradigm, it's trying to accomplish a goal.”
Blake Lemoine Feb 23, 2023 ▶ 3:22 Blake Lemoine and Gary Marcus Debate AI Chatbots
Feb 23, 2023 negative
Opinion
Marcus: We Have No Way to Truly Debug AI Systems
“I think that's actually the deepest problem here is we have no way to debug these systems, really. We have the band-aids of, like, reinforcement learning and things like that that are so indirect.”
Gary Marcus Feb 23, 2023 ▶ 19:33 Blake Lemoine and Gary Marcus Debate AI Chatbots
Aug 3, 2023 bullish
Prediction Not checkable as stated
Wood: AI agents and output vetting will become as important as models
“So there's a lot of focus on models today, but those models are going to remain important. But over time, there's going to be additional capabilities like agents, like reinforcement learning, like the ability to be able to understand and vet The responses that…”
Matt Wood Aug 3, 2023 ▶ 9:46 Amazon Reveals Its AI Master Plan — With Matt Wood
Sep 14, 2023 bullish
Prediction Not checkable as stated
Murdock: Merging LLMs and reinforcement learning will drive major AI breakthroughs
“I don't know exactly what the breakthrough will be, but I'm really excited about the union of these LLMs plus reinforcement learning. And I'm excited about that because I think there's a lot more to come from reinforcement learning. And I know at Google D mine…”
Colin Murdock Sep 14, 2023 ▶ 50:27 Google's DeepMind Wants To Make Human-Level Artificial Intelligence, Says Its Chief Business Officer
May 9, 2024 bullish
Prediction Not checkable as stated
Clark: Scaling RL compute will unlock major new AI capabilities
“Everyone is trying to figure out how they can spend more and more of their compute on reinforcement learning, because I think everyone has this intuition that the more RL you add, the more sophisticated you're going to be able to make these things, and a lot o…”
Jack Clark May 9, 2024 ▶ 22:39 Anthropic's Co-Founder on AI Agents, General Intelligence, and Sentience — With Jack Clark
May 15, 2024
Insight
Long-horizon reinforcement learning for AI agents is constrained by sparse rewards
“People have been talking about long horizon RL, which is the training method. You need to get something like this, where you go tell it to do something and then you reward it at the end for having achieved that outcome. But the difficulty with those kinds of a…”
Dwarkesh Patel May 15, 2024 ▶ 36:58 AI Scaling, Alignment, and the Path to Superintelligence — With Dwarkesh Patel
Mar 28, 2025 bullish
Assertion Not checkable as stated
Hendrycks: RL-based reasoning models are improving faster than pre-training did
“That is separate from the new reasoning paradigm that has emerged in the past year which is where you train models to on math and coding types of questions with reinforcement learning, and that has a very steep slope, and I don't see any signs of that slowing …”
Dan Hendrycks Mar 28, 2025 ▶ 24:15 AI's Rising Risks: Hacking, Virology, Loss of Control — With Dan Hendrycks
Apr 23, 2025
Insight
Patel: Human labeling is unscalable, forcing reliance on synthetic AI data
“Using humans to train models is just so expensive, right? So then there's the magic of sort of reinforcement learning and other synthetic data technologies, right? Where the model is helping teach the model, right? So you have many models in, in, in a sort of,…”
Dylan Patel Apr 23, 2025 ▶ 18:14 Generative AI 101: Tokens, Pre-training, Fine-tuning, Reasoning — With SemiAnalysis CEO Dylan Patel
Jun 18, 2025 bearish
Insight
Patel: Delayed Reward Signals Will Slow AI Progress on Long-Horizon Tasks
“Now if we're getting to the world where you got to like do a project for seven hours and then at the end of those seven hours, then we tell you, Hey, did you did you get this right? Then like the progress just goes on a bunch. Cause you've gone from like getti…”
Dwarkesh Patel Jun 18, 2025 ▶ 20:41 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Jun 18, 2025 bearish
Prediction Not checkable as stated
Patel: Reinforcement learning may not generalize beyond verifiable domains
“I still think I I'm like, I'm not confident that this will generalize to domains that are not so verifiable or text-based.”
Dwarkesh Patel Jun 18, 2025 ▶ 7:34 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Jul 30, 2025 bullish
Prediction Not checkable as stated
Amodei: AI lag on subjective tasks is a temporary obstacle
“We've seen more progress on say math and code where, where the models are, you know, getting pretty close to like a high professional Level and less on more subjective tasks. But I think that is very much a temporary obstacle.”
Dario Amodei Jul 30, 2025 ▶ 7:27 Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism
Jan 16, 2026 neutral
Prediction Not checkable as stated
Mensch: AI customization techniques will be abstracted away for enterprises
“I do expect the part of the software in those deployment to increase. So the amount of the way customization occurs today with fine tuning, reinforcement learning, this kind of things, this is going to be abstracted away from the enterprise buyer because it's …”
Arthur Mensch Jan 16, 2026 ▶ 21:11 Who Wins if AI Models Commoditize? — With Mistral CEO Arthur Mensch
Mar 9, 2026 neutral
Assertion Supported
Kantrowitz: Scale AI shifted majority of training to reinforcement learning
“I just did the story for with about scale AI saying that the majority of their training has moved to reinforcement learning where they train models to act in specific environments like filling out forms, and then they baked those capabilities back Into the mod…”
Alex Kantrowitz Mar 9, 2026 ▶ 28:43 AI Revenue Explodes, Dario’s Memo, McDonalds’ CEO’s Baby Burger Bite
Apr 27, 2026 neutral
Assertion Not checkable as stated
Kantrowitz: Scale AI now does most of its training via reinforcement learning
“Scale AI, Alexander Wang's company, they told me recently that most of the training that they're doing is reinforcement learning, where you build environments for the bots and they go and they try to figure out what to do.”
Alex Kantrowitz Apr 27, 2026 ▶ 42:52 Apple After Tim Cook, OpenAI’s New Mojo, Meta’s Internal Tracking Escapade
Jul 7, 2026 positive
Insight
Kutylowski: Task-focused reinforcement learning outperforms broad multi-task training in specific domains
“And if you run this reinforcement learning step on too many different tasks, the model will be able to do all of that. But once again, it's going to be very, very, very broad. And if you focus on making sure that the model understands and knows that it needs t…”
Jarek Kutylowski Jul 7, 2026 ▶ 7:15 Why Specialized AI Models Are Challenging the Frontier Labs — With DeepL CEO Jarek Kutylowski
Jul 8, 2026 bullish
Assertion Not checkable as stated
Bosworth: Reinforcement learning plays a far bigger role in AI than predicted
“Reinforcement learning is playing a huge, a much bigger role in today's kind of AI than people had maybe predicted two or three years ago that it would.”
Andrew Bosworth Jul 8, 2026 ▶ 38:01 Meta CTO Andrew Bosworth: Our Path To Frontier AI, Renting Models, Consumer AI's Struggles
Sep 7, 2026 neutral
Assertion Supported
Kantrowitz: OpenAI paused some reinforcement learning on new model training
“We are starting to see some of the labs do things like opening. I, for instance, paused some reinforcement learning for a bit on the training of its new models.”
Alex Kantrowitz Sep 7, 2026 ▶ 43:08 GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
Sep 7, 2026 bearish
Insight
Kantrowitz: Reinforcement learning layers add ruthlessness to AI models
“What we have now is that the reinforcement learning type of AI technology has been put on top of the self supervised learning to get these AI models working better, which has added a level of ruthlessness to them. Because one of the things we know about RL is …”
Alex Kantrowitz Sep 7, 2026 ▶ 41:22 GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.