reinforcement learning

22 statements across 18 episodes · 16 bullish · 1 bearish · 18 people on the record · first statement Jan 2, 2019 by Murray Shanahan · across every show →

Everything said about reinforcement learning, oldest first

Jan 2, 2019 positive
Insight
Shanahan: Reinforcement learning progress does not require massive datasets
“Actually, DeepMind are another example of the same thing, because if you want to apply reinforcement learning to games, and that's enabled them to make some quite fundamental sort of progress, you don't need vast amounts of data either.”
Murray Shanahan Jan 2, 2019 ▶ 32:49 a16z Podcast | Artificial Intelligence and the 'Space of Possible Minds'
Jan 2, 2019 positive
Assertion Supported
Schuler: Rich Sutton's textbook is computer science's most influential publication
“According to the Alan Soufre AI Semantics Scholar, he's the most highly cited researcher in reinforcement learning, 11th most influential researcher in all of competing science. And his textbook on reinforcement learning was ranked as the single most influenti…”
Cameron Schuler Jan 2, 2019 ▶ 16:26 a16z Podcast | Machine Intelligence, from University to Industry
Jan 2, 2019 bullish
Insight
Chen: AlphaGo Zero proves reinforcement learning succeeds without labeled data
“Right, where there's so much momentum right now that says, basically, we're one labeled data set away from glory. Right, and this result basically shows you, wow, there's a lot of mileage that you can get out of reinforcement learning where there's no data, no…”
Frank Chen Jan 2, 2019 ▶ 34:50 a16z Podcast | Revenge of the Algorithms (Over Data)... Go! No?
Feb 28, 2025 negative
Insight
Ayrey: Filtering API keys from training risks degrading AI data science skills
“So if we do our reinforcement learning and we skew it towards code snippets that's generating that don't have API keys, inadvertently, we may be training this thing to behave less like a data scientist. And then we lose the entire discipline of data science in…”
AI Security Researcher Feb 28, 2025 ▶ 9:48 Avoiding vulnerabilities in AI code
Mar 5, 2025 bullish
Insight
Mascorro: DeepSeek-R1 proved reinforcement learning improves models without human feedback
“And I think the big thing in, in R-one, or generally with these reasoning models is, We were doing before there was a human in the loop always, right? Like when we have this SFT training and these other techniques that we're doing after like RLHF and having R …”
Marco Mascorro Mar 5, 2025 ▶ 6:42 DeepSeek, Reasoning Models, and the Future of LLMs
May 29, 2025 positive
Insight
Angelopoulos: Reinforcement learning allows AI models to surpass human teachers
“And supervised learning, you can only do as well as the best human that you have. Because what's happening is that you're learning from the teacher. In reinforcement learning, you're learning from the world. You're able to learn things better than the best hum…”
Anastasios Angelopoulos May 29, 2025 ▶ 1:00:48 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Jul 23, 2025 positive
Assertion Not checkable as stated
Caldwell: Reinforcement learning enables fully automated control of industrial plants
“We're at this point where kind of like the compute and machine learning, reinforcement learning Do really enable, enable you to go no humans in the loop and how a lot of these plants are controlled.”
Turner Caldwell Jul 23, 2025 ▶ 16:44 The U.S. Can’t Build AI Without These Materials
Jul 23, 2025 positive
Insight
Caldwell: Humans are poorly suited for multivariable plant optimization compared to RL
“As you like build more of that large scale infrastructure and kind of like see the problems that humans have to solve on a daily basis, it becomes like pretty obvious that these are problems that humans don't, aren't best positioned to solve. Like these are la…”
Turner Caldwell Jul 23, 2025 ▶ 16:55 The U.S. Can’t Build AI Without These Materials
Jul 28, 2025
Insight
Martin Casado: Domain-specific reinforcement learning requires a plurality of models
“But as soon as you're doing RL, where you're training it in a specific domain with a specific verifier, you're likely losing, you know, other areas. So you make it really, really good at playing chess. It's going to be less good at, you know, something else li…”
Martin Casado Jul 28, 2025 ▶ 47:34 Balaji Srinivasan: How AI Will Change Politics, War, and Money
Aug 8, 2025
Insight
Isa Fulford: Reinforcement learning for specific model capabilities is data-efficient
“Training a model to be good at a specific capability is very data efficient. You don't need that many examples to teach it something new.”
Isa Fulford Aug 8, 2025 ▶ 7:16 GPT-5 and Agents Breakdown – w/ OpenAI Researchers Isa Fulford & Christina Kim
Aug 8, 2025 positive
Insight
Fulford: RL breakthroughs in math and coding unlocked functional AI agents
“When we saw the reinforcement learning algorithm working really well on math and physics problems and coding problems, It became pretty clear, like, just from reading through the chain of thought, like, okay, this thing's actually, like, thinking and reasoning…”
Isa Fulford Aug 8, 2025 ▶ 13:37 GPT-5 and Agents Breakdown – w/ OpenAI Researchers Isa Fulford & Christina Kim
Sep 25, 2025 positive
Prediction Not checkable as stated
Pachocki: AI reward modeling will evolve toward simpler, human-like learning
“I expect this will evolve quite rapidly. I expect it will become simpler, right? Like I think, you know, maybe like two years ago, we would have been talking about like, what is the right way to craft my fine tuning data set? And I don't think we are like at t…”
Jakub Pachocki Sep 25, 2025 ▶ 15:26 From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
Sep 30, 2025 positive
Insight
Fedus: High-compute reinforcement learning is essential for AI tool use
“High compute reinforcement learning is really effective. This is how you should think about the strategies it's using. This is how you create effective tool using towards those problems, and this is how you optimize it effectively.”
Liam Fedus Sep 30, 2025 ▶ 40:35 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Oct 14, 2025 bullish
Prediction Not checkable as stated
Labenz: LLM fine-tuning and reinforcement learning will successfully power humanoid robotics
“All these techniques that have been developed over the last few years, Seems to me they're absolutely gonna apply to a problem like a humanoid robot as well.”
Nathan Labenz Oct 14, 2025 ▶ 58:00 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Oct 23, 2025 positive
Insight
Masad: Reinforcement learning enabled long-horizon reasoning in AI models
“I think it's RL. I think it's reinforcement learning.”
Amjad Masad Oct 23, 2025 ▶ 15:21 Marc Andreessen & Amjad Masad on “Good Enough” AI, AGI, and the End of Coding
Oct 23, 2025 bullish
Insight
Masad: Reinforcement learning paired with TreeSearch still has substantial runway
“I think the breakthroughs in RL are incredibly exciting, but we also knew about them now for like over 10 years where you marry generative systems with TreeSearch and things like that. But there's a lot more to go there”
Amjad Masad Oct 23, 2025 ▶ 52:34 Marc Andreessen & Amjad Masad on “Good Enough” AI, AGI, and the End of Coding
Oct 23, 2025 positive
Assertion Supported
Andreessen: AI labs hire mathematicians and coders for reinforcement learning data
“Foundation model companies are, in some cases, they are hiring, they're actually hiring human experts. To generate new training data. So they're actually hiring mathematicians and physicists and coders to basically sit, and, you know, they're hiring human prog…”
Marc Andreessen Oct 23, 2025 ▶ 29:16 Marc Andreessen & Amjad Masad on “Good Enough” AI, AGI, and the End of Coding
Nov 5, 2025
Assertion Partly supported
Kuyda: OpenAI temporarily abandoned language models for video game agents
“And very quickly they stopped working on language models, and we were very upset because we really wanted to continue going there, but they didn't want to talk about any language models because no one was really working on them and that may have made us feel v…”
Eugenia Kuyda Nov 5, 2025 ▶ 40:29 Seeing The Future from AI Companions to Personal Software
Nov 17, 2025
Insight
Shear: In AI agents, care equates to loss and reward correlation
“Care is basically like reward. Like how much does this state correlate with survival? How much does this state correlate with your inclusive, your full inclusive reproductive fitness? For a somewhat thing that learns evolutionarily, or for a reinforcement lear…”
Emmett Shear Nov 17, 2025 ▶ 23:57 Emmett Shear on Building AI That Actually Cares: Beyond Control and Steering
May 13, 2026 positive
Disclosure
Caldwell: Mariana Minerals uses reinforcement learning to automate mineral refineries
“We're making a big bet on autonomy and refineries, where we use reinforcement learning to actually remove humans from the loop in determining how refineries operate.”
Turner Caldwell May 13, 2026 ▶ 0:28 The Founders Who Left Tesla to Rebuild America | a16z
Aug 26, 2026 bullish
Insight
Acharya: Domain-specialized open models outperform general models through RL and reasoning traces
“If you actually have a problem that you can specialize the model around with your reasoning traces, You can start to create this compounding advantage in your domain for your customer base, where you're able to kind of shape the intelligence to be better than …”
Anish Acharya Aug 26, 2026 ▶ 9:19 The State of AI: Models, Moats, and the Consumer Renaissance
Sep 1, 2026
Insight
Litt: Reinforcement learning struggles to reward intermediate mathematical theory building
“I think what is definitely true is that, like, the skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one. So it might be harder, you know I guess you can try, you can tell it, you kno…”
Daniel Litt Sep 1, 2026 ▶ 27:08 Can AI Learn Mathematical Intuition?
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.