reinforcement learning

also referred to as: rl

69 statements across 36 episodes · 30 bullish · 11 bearish · 37 people on the record · first statement Oct 21, 2023 by Kanjun Qiu · across every show →

Everything said about reinforcement learning, oldest first

Oct 21, 2023 bearish
Insight
Qiu: Pure reinforcement learning cannot deliver planning and reasoning
“The second thing we learned is that reinforcement learning is not a good vehicle. Like, pure reinforcement learning is not a good vehicle for planning and reasoning.”
Kanjun Qiu Oct 21, 2023 ▶ 25:16 Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
Jan 11, 2024 negative
Opinion
Lambert: Reinforcement learning in language models is contrived and not real RL
“And the view of RL in language models is pretty contrived already. So it's not like we're doing real RL.”
Nathan Lambert Jan 11, 2024 ▶ 7:59 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 negative
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 neutral
Insight
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Nathan Lambert Jan 11, 2024 ▶ 16:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Prediction Not checkable as stated
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”
Nathan Lambert Jan 11, 2024 ▶ 7:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Insight
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Prediction Not checkable as stated
Lambert: AI community will clarify if chain-of-thought maps to RL within a year
“I think in the next year that'll probably get kind of made more concrete by the community on, like, if you can easily draw out, like, if chain of thought reasoning is more like RL.”
Nathan Lambert Jan 11, 2024 ▶ 18:00 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Mar 27, 2024 positive
Insight
Luan: LLMs shortcut evolutionary RL by behaviorally cloning all human knowledge
“Like de novo RL is like a pretty terrible way to get there quickly. Why are we rediscovering all the knowledge about the world? Like years ago, I had a debate with a Berkeley professor as to like what will it actually take to build HCI? And his view is basical…”
David Luan Mar 27, 2024 ▶ 5:38 Why Google failed to make GPT-3 -- with David Luan of Adept
Sep 27, 2024 positive
Insight
Yao: Reflexion replaces scalar RL rewards with verbal gradient descent
“I think one way to think of reflection is that the traditional idea of reinforcement learning is you have a scalar reward, and then you somehow back propagate the signal of the scalar reward. To the rest of your neural network through whatever algorithm, like …”
Shunyu Yao Sep 27, 2024 ▶ 15:35 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Nov 11, 2024 neutral
Opinion
Polu: DeepMind IMO breakthrough relied on scaling RL and autoformalization
“I think the DeepMind team just did a good job of scaling. I think there's nothing too magical in their approach, even if it hasn't been published as a Dan Silver talk from seven days ago, where it goes a little bit into more details. It feels like there's noth…”
Stanislas Polu Nov 11, 2024 ▶ 9:00 Agents @ Work: Dust.tt — with Stanislas Polu
Mar 7, 2025 neutral
Insight
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”
Misha Laskin Mar 7, 2025 ▶ 11:38 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 7, 2025 positive
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Misha Laskin Mar 7, 2025 ▶ 4:34 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 23, 2025 positive
Insight
Agarwal: Optimal Post-Training Pipeline Combines Heavy Distillation Followed by RL
“So, so I would think maybe an optimal pipeline would look like you do distillation heavily, but then you still do some RL afterwards, because maybe there's still something you can get out of your reward functions or whatever your post-training stack is.”
Rishabh Agarwal Mar 23, 2025 ▶ 8:05 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Assertion Supported
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Rishabh Agarwal Mar 23, 2025 ▶ 38:45 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Apr 29, 2025 bullish
Insight
Jin: RL enables models to surpass expert labelers and develop self-direction
“The model outperforming expert labelers is, is possible. The model learning, like, self-direction is, like, expected. And yeah, we've seen, like, kind of cool emergent behaviors with, like, you know, like, O-one, O-three, R-one, kind of, like, these, like, thi…”
Roger Jin Apr 29, 2025 ▶ 6:50 What is an RL environment? w/ Nous Research's Roger Jin
May 21, 2025 neutral
Prediction Not checkable as stated
Alberti: AI App Companies Must Encode Product Needs Into RL Feedback
“I think that's almost how I view like the future of application layer companies. Cause I mean, yeah, you see like the different, the labs are also now creating these like RL platforms and you can soon like customize models with RL on your personal, like on you…”
Silas Alberti May 21, 2025 ▶ 31:16 DeepWiki: The GitHub Encyclopedia
Jun 6, 2025
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39 The Utility of Interpretability — Emmanuel Amiesen
Jul 18, 2025 negative
Insight
Kamradt: Synthetic RL Transfers Developer Intelligence Rather Than Creating True Intelligence
“Often what happens is the human or developer intelligence is often injected into that environment itself, and so the model isn't actually Intelligent. You're just almost like taking the intelligence from the developer, injecting it into the environment, and th…”
Greg Kamradt Jul 18, 2025 ▶ 6:24 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Jul 18, 2025 neutral
Assertion Supported
Kamradt: xAI Increased RL Compute on Grok 4 Tenfold
“They tend X the RL that they put on top of Grok for that.”
Greg Kamradt Jul 18, 2025 ▶ 36:42 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Jul 24, 2025 bullish
Prediction Not checkable as stated
Scaling math AI becomes purely compute and data once auto-evaluation works
“And then my guess is, I believe in IL, so if for each category, we can figure out the A way to auto-evaluate the results, then after that, it will just be compute and data.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 21:19 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 24, 2025 bullish
Prediction Not checkable as stated
Powerful math reasoning models will arrive soon using Lean as a verifier
“And so if you can just use lean to kind of become the verifier, then it's easy. So I think we probably will see very powerful reasoning models in, in math very soon.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 19:37 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Jul 31, 2025 neutral
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Nathan Lambert Jul 31, 2025 ▶ 7:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 positive
Insight
Lambert: Scaling RL long enough requires a curriculum of increasing difficulty
“If you scale RL long enough, You're going to need a curriculum of things getting harder. And like, that's pretty obvious.”
Nathan Lambert Jul 31, 2025 ▶ 32:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 29, 2025 positive
Assertion Not checkable as stated
Morcos: Qwen is much easier to align than Llama due to pre-training
“It's much easier to RL Quen than it is to do Lama. Likely that has to do with the fact that Quen put a lot of synthetic reasoning traces into their training data.”
Ari Morcos Aug 29, 2025 ▶ 54:07 Better Data is All You Need — Ari Morcos, Datology
Oct 11, 2025 negative
Insight
Lenz: RL training wastes compute on saturated or impossible examples
“Once you've trained a few hundred steps of let's say GOP, Most of your training is just wasted on example that are either too hard for you and you didn't get any success on them or too easy and everything was a success.”
Barak Lenz Oct 11, 2025 ▶ 37:49 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Oct 16, 2025 bullish
Prediction Not checkable as stated
Corbitt: 55-60% chance RL becomes the standard pattern for deploying scale agents
“I think that the chances that like everyone should be, or, you know, everyone who's deploying an agent at scale should be doing RL with it, either as part of sort of like a, you know, like pre-deployment or even like continuously as it's deployed, that that's …”
Kyle Corbitt Oct 16, 2025 ▶ 18:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Oct 16, 2025 bearish
Opinion
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Kyle Corbitt Oct 16, 2025 ▶ 22:46 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Oct 16, 2025 positive
Insight
Corbitt: RL reward hacking is easily detected as models repeat the exploit
“Reward hacking is quite easy to detect once it starts happening, because once the model's found some hack, it just starts, like, doing it all the time.”
Kyle Corbitt Oct 16, 2025 ▶ 1:05:41 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Oct 16, 2025 neutral
Insight
Corbitt: Agent RL requires real runs inside highly realistic environments
“For RL to work, you have to be looking at real runs, ideally of your actual agent in its current state across within an environment as real as possible.”
Kyle Corbitt Oct 16, 2025 ▶ 30:10 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Dec 6, 2025 positive
Insight
Making billions of gameplay clips playable bridges imitation learning to reinforcement learning
“Actually making every single clip on the platform playable at billions of clips scale is how we go from imitation learning to RL.”
Pim de Witte Dec 6, 2025 ▶ 1:00:30 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Dec 6, 2025 positive
Assertion Not checkable as stated
General Intuition's foundation agent runs purely on vision without reinforcement learning
“This is just a base model. There's no RL, no fine tuning. This model sees no game states. It is purely capable, not sequence acceptance. It's purely predicting the actions from the phrase. That's it.”
Pim de Witte Dec 6, 2025 ▶ 4:04 World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
Dec 30, 2025 neutral
Opinion
Nair: Frontier AI labs have converged on similar reinforcement learning methods
“Well, it does seem like basically a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of like Frontier again”
Ashvin Nair Dec 30, 2025 ▶ 31:54 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 negative
Opinion
Nair: 2017–2022 academic RL breakthroughs failed because researchers overfit to benchmarks
“A lot of the methods that people were really excited about is, like you know, off policy learning, like, value functions, like, these kind of things, and somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why, but in the…”
Ashvin Nair Dec 30, 2025 ▶ 9:26 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 positive
Insight
Nair: Context integration, not model intelligence, bottlenecks useful automation
“A big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, …”
Ashvin Nair Dec 30, 2025 ▶ 13:09 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 30, 2025 neutral
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Ashvin Nair Dec 30, 2025 ▶ 12:26 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 31, 2025
Insight
McGrath: RL runs have far more infrastructure failure points than pre-training
“The issue with RL is, like, you're doing tasks, and each task could have, like, a different grading setup, and each one of those different grading setups, that's, like, more infrastructure, and so, You know, when I'm staying up late trying to figure out what's…”
Josh McGrath Dec 31, 2025 ▶ 2:12 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Dec 31, 2025 positive
Assertion Supported
Wang: GPU environments collect hundreds of millions of RL timesteps hourly
“With these, like, GPU accelerated environments, we can collect hundreds of millions of time steps of data within just a few hours”
Kevin Wang Dec 31, 2025 ▶ 17:43 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Dec 31, 2025 positive
Insight
Eysenbach: 1,000-layer RL requires reward-free objectives, not just architectural tricks
“I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective. This objective doesn't actually use rewards in it, and so there's another word in…”
Benjamin Eysenbach Dec 31, 2025 ▶ 8:08 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Dec 31, 2025 positive
Assertion Supported
Kevin Wang: Scaling RL network depth unlocks effective batch size scaling
“We notice that we see that scaling width actually also improves performance, and we also find that actually by scaling depth, we actually unlock the ability to scale along batch size as well.”
Kevin Wang Dec 31, 2025 ▶ 22:38 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Dec 31, 2025 neutral
Assertion Supported
Wang: 64 layers saturate performance in most reinforcement learning tasks
“Within our paper, like, for most environments we are able to, like, saturate, like, get to, like, almost perfect performance within just, you know, we don't even need to get to, like, a thousand layers. Like, maybe just 64 layers, for example, is sufficient.”
Kevin Wang Dec 31, 2025 ▶ 15:25 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Dec 31, 2025 neutral
Assertion Partly supported
Michał Zawalski: RL Scaling Gains Require Over 50 Million Transitions
“Going back to our paper, if you look at the plots, we only see this, like, huge performance increase When we cross, like, 50 millions of transitions gap.”
Michał Zawalski Dec 31, 2025 ▶ 17:06 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Dec 31, 2025 positive
Insight
Kevin Wang: Cross-entropy trajectory classification enables scalable deep reinforcement learning
“I think it's because we're fundamentally shifting the burden of learning from something like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're try…”
Kevin Wang Dec 31, 2025 ▶ 9:56 [NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Jan 17, 2026 bearish
Assertion Not checkable as stated
Reggio: Simple web research agents outperformed RL for credit underwriting at Brex
“We made this big investment. We were working with some outside, like the, like a company that specializes in this and the performance we ended up getting was inferior to just building a, like a web research agent.”
James Reggio Jan 17, 2026 ▶ 40:58 Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)
Jan 23, 2026 positive
Opinion
Yi Tay: Reinforcement learning is the primary AI modeling toolset today
“So I think RL is basically the main modeling tool set that we play around with these days.”
Yi Tay Jan 23, 2026 ▶ 4:10 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026 neutral
Insight
Yi Tay: ML and RL Knowledge Can Be Learned Easily by Engineers
“ML. ML can be learned easily. Our knowledge can be learned easily.”
Yi Tay Jan 23, 2026 ▶ 1:26:12 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026 neutral
Insight
Yi Tay: 'Reasoning' Technically Just Means Post-Training RL with Thinking Trajectories
“So I think the actual, like, technical definition of reasoning is making models better with thinking and post-training. Ok? Yeah. So basically, like, RL-ing the model to think better.”
Yi Tay Jan 23, 2026 ▶ 33:24 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026
Disclosure
Yi Tay: I had almost no RL background before returning to DeepMind
“I spent a lot of my past life, I call it the past art, working on like architectures and pre-training, but I think now I more, I have like transitioned more into RL. I'm not like old school RL, but the games RL and the old school RL, and to be honest, I had al…”
Yi Tay Jan 23, 2026 ▶ 3:30 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 28, 2026
Insight
White: Writing bulletproof RL verifiers is far harder than supervised training
“Pre-training or training transformers, you know on just data, like just supervised training where you just have the inputs and the outputs directly, very nice, relaxing, you know, like things are always robust, you know, things go pretty smoothly. When we do t…”
Andrew White Jan 28, 2026 ▶ 1:12:03 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Feb 12, 2026 positive
Insight
Dean: Applying RL to non-verifiable domains would dramatically improve AI models
“How do you get RL to work for non-verifiable domains? I think it's a pretty interesting open problem because I think that would broaden out the capabilities of the models, the improvements that you're seeing in both math and coding if we could apply those to o…”
Jeff Dean Feb 12, 2026 ▶ 42:58 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Feb 24, 2026 bearish
Prediction Not checkable as stated
O'Laughlin: Tech industry may face a CPU shortage from AI coding and RL
“You feel like we might actually be seeing a CPU shortage partially because of this refresh cycle, but partially also because like I legitimately believe the cloud code Cloud code is increasing software creation and then on top of that, there is real demand fro…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:56:07 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Feb 25, 2026 positive
Insight
Welling: Diffusion Models Share Exact Mathematics With Non-Equilibrium Stochastic Thermodynamics
“It turns out that the mathematics that we use for diffusion models, but even for reinforcement learning, for Schrodinger bridges, for MCMC sampling, has the same mathematics as this theory, this physical theory of non-equilibrium Systems.”
Max Welling Feb 25, 2026 ▶ 4:59 🔬Max Welling: Materials Underlie Everything
Mar 30, 2026 neutral
Insight
Lample: Long-horizon RL trajectories require new algorithms beyond GRPO
“GRPO, for instance, it doesn't really work with any bit of policy, which was okay initially, because you are solving math problems that can be solved in like a few thousand tokens, so the model can actually generate them pretty quickly, so when you do your upd…”
Guillaume Lample Mar 30, 2026 ▶ 45:43 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
May 21, 2026 bullish
Prediction Not checkable as stated
Burazin: RL workloads will reach 50% of Daytona's volume this month
“It will be this one 50%, yeah.”
Ivan Burazin May 21, 2026 ▶ 28:22 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
Jun 3, 2026 positive
Insight
Hong: Lean and Rust yield superior reinforcement learning convergence over Python
“If you want proof to be informal math, It's very annoying, because then that's, like, just makes objective function. Your code is something like Python, your proof is, say, natural language, math proof. You will not have very strong RL kind of performance, rig…”
Carina Hong Jun 3, 2026 ▶ 30:00 Scaling Past Informal AI - Carina Hong, Axiom Math
Jun 21, 2026 bearish
Opinion
Malde: Standard reinforcement learning is broken for continual learning
“RL, it's still taking all of this kind of Useful information from the real world, like I mentioned, all the corrections and everything, and putting it into just one number. Which is really broken.”
Ronak Malde Jun 21, 2026 ▶ 19:13 ⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
Jun 24, 2026 bullish
Prediction Not checkable as stated
Zaharia: Customizing AI models will get significantly easier over time
“My feeling is, like customizing models is actually going to get way easier over time. That's what we're finding, because The base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces, and the…”
Matei Zaharia Jun 24, 2026 ▶ 1:03:14 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Jun 25, 2026 neutral
Insight
OpenAI's Chen: Reinforcement learning struggles in subjective, hard-to-grade fields
“RLs traditionally had headwinds when it's come to fields that, you know, it's more kind of, Subjective than objective. So if you kind of think of, you know, one kind of, you know example of this is creative writing, where, you know, you could take two pieces o…”
Mark Chen Jun 25, 2026 ▶ 5:54 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jun 25, 2026
Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Mark Chen Jun 25, 2026 ▶ 14:03 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Jul 8, 2026
Insight
Bubna: RL Rollouts Are Extremely Bursty and Can Require 100,000 Sandboxes
“RL is insanely bursty. Like when you're doing rollouts you sometimes need a 100,000 sandboxes.”
Akshat Bubna Jul 8, 2026 ▶ 15:42 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 8, 2026
Insight
Bubna: Transferring RL weights is fundamentally an OS memory problem
“Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is, there's a lot of degrees of freedom, and it is basically a systems problem of Moving me…”
Akshat Bubna Jul 8, 2026 ▶ 31:40 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Jul 11, 2026 negative
Insight
Perszyk: Task-specific reinforcement learning fails to produce generalizable intelligence
“You can use things like reinforcement learning to get them really good at specific tasks that we might care about, but you do that for one task and you, it is not good at another task or it doesn't generalize.”
Danielle Perszyk Jul 11, 2026 ▶ 18:32 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Jul 16, 2026 neutral
Assertion Supported
Beam: Reinforcement learning achieves only 5% to 6% GPU FLOP utilization
“And for reinforcement learning, it's always somewhere, like, around five to, like, six percent. So, said differently, that means that we're getting, like, five percent of the actual GPU computing power that we're paying for.”
Andy Beam Jul 16, 2026 ▶ 1:38:42 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 16, 2026 bullish
Opinion
Beam: Nature and scientific experiments are ultimate verifiers for RL
“But what at Lilo we believe is that actually science running the scientific method and using nature and experiments as verifier is like the ultimate version of that. And so what we're building, we'll talk about these things that we call AI science factories. T…”
Andy Beam Jul 16, 2026 ▶ 6:59 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Jul 22, 2026 bullish
Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Eiso Kant Jul 22, 2026 ▶ 45:09 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 22, 2026 neutral
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 22, 2026 neutral
Insight
Kant: AI coding models perform best in their creators' proprietary harnesses
“No doubt it's going to be better in your own harness. And it's just because of like, where are you putting your reinforcement learning compute, right? You're putting your RL and your synthetic data. You're putting it to your own harness because it's the one th…”
Eiso Kant Jul 22, 2026 ▶ 1:01:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Jul 22, 2026 neutral
Insight
Kant: RL compute cannot scale like pre-training due to task batch constraints
“And RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the en…”
Eiso Kant Jul 22, 2026 ▶ 1:39:07 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.