reinforcement learning

also referred to as: rl

19 statements across 13 episodes · 12 bullish · 4 bearish · 13 people on the record · first statement Apr 25, 2023 by Noam Brown · across every show →

Everything said about reinforcement learning, oldest first

Apr 25, 2023 bearish
Insight
Reinforcement learning struggles in trading because financial markets are non-stationary
“I think the major challenge with Using things like reinforcement learning for trading is that it's a non-stationary environment. So you can have all this historical data, but it's not a stationary system and it's gonna like the markets respond to world events,…”
Noam Brown Apr 25, 2023 ▶ 30:54 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
Jun 1, 2023
Insight
Schimpf: Bounding non-deterministic AI is essential for military tech adoption
“A lot of these areas that are more non-deterministic, so, you know, things like reinforcement learning or, you know, potential applications of LOMs into the space, they are inherently non-deterministic, and that is a risk. And so kind of quantifying that, know…”
Brian Schimpf Jun 1, 2023 ▶ 20:05 No Priors Ep. 19 | With Anduril CEO Brian Schimpf
Jan 24, 2024 negative
Insight
Chen: Next-token LLM training is mimicry, imposing a natural capability ceiling
“For people that study reinforcement learning, we call it behavior cloning, which means you're just asking the AI to clone the behavior of another agent. And that is like one of the most primitive way possible to train this type of systems. Like, because if you…”
Peter Chen Jan 24, 2024 ▶ 39:24 No Priors Ep. 48 | With Covariant CEO Peter Chen
Mar 7, 2024 bullish
Prediction Not checkable as stated
Gil: New reinforcement learning AI agent products will emerge within 6-12 months
“And so I think that that purpose of knowledge is about to hit the world in the context of new products. And it'll take time for those products to emerge, you know, six months, 12 months, a year. But It does feel like that's another wave that's coming where you…”
Elad Gil Mar 7, 2024 ▶ 11:16 No Priors Ep. 54 | With Sarah Guo & Elad Gil
Aug 1, 2024 neutral
Assertion Supported
Vinyals: AlphaGo compute was mostly RL self-play, unlike modern LLM pre-training dominance
“Historically, if you look at AlphaGo, which actually followed quite closely the recipe of you pre-train your model on all human data, you then use RL to make it better, and then you do some search at inference time, the compute there was very skewed for the mi…”
Oriol Vinyals Aug 1, 2024 ▶ 23:43 No Priors Ep. 74 | With Google DeepMind VP of Research Oriol Vinyals
Aug 1, 2024 positive
Insight
Vinyals: LLM bootstrapping works if verification is easier than solution generation
“If checking that something is correct is easier than creating the solution, then we're in business because the language models will be able to evaluate their own samples more accurately than to generate them. And then we have a sort of reinforcement learning l…”
Oriol Vinyals Aug 1, 2024 ▶ 27:48 No Priors Ep. 74 | With Google DeepMind VP of Research Oriol Vinyals
Aug 1, 2024 bullish
Prediction Not checkable as stated
Vinyals predicts pre-training compute will drop to ~50% as RL expands
“So to me, that balance feels correct, like some on pre-training, and here we, we're trying to learn every task. So certainly that's going to be, you know, let's say it can be as high as 50%, not as high as over 90 like today. And then the rest mostly on reinfo…”
Oriol Vinyals Aug 1, 2024 ▶ 25:29 No Priors Ep. 74 | With Google DeepMind VP of Research Oriol Vinyals
Nov 14, 2024 bullish
Prediction Not checkable as stated
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Rishi Mehta Nov 14, 2024 ▶ 27:32 No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
Mar 20, 2025 bullish
Prediction Not checkable as stated
Chelsea Finn: Autonomous RL experience will play a huge role in robotics
“And then I also think that autonomous experience will play a huge role, just like we've seen in language models. After you get an initial language model, if you can use reinforcement learning to have the robot, the language model bootstrap on its own experienc…”
Chelsea Finn Mar 20, 2025 ▶ 30:54 No Priors Ep. 107 | With Physical Intelligence Co-Founder Chelsea Finn
Apr 10, 2025 positive
Insight
Foody: Model Improvement via RL Is Gated Entirely by Evaluation Benchmarks
“Reinforcement learning is becoming so effective that once you create evals, the models can learn them and how to you know, improve capabilities. And so for everything that we want alums to be good at, we need evals for those things.”
Brendan Foody Apr 10, 2025 ▶ 1:14 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
Apr 24, 2025 positive
Insight
RL models only need task and outcome definitions to learn research trajectories
“The cool thing with RL is that you don't necessarily need to Know the whole process of how the person would do the research. You just have to know what the task is and what the outcome should be, and the model will just learn during training how to get from th…”
Isa Fulford Apr 24, 2025 ▶ 9:32 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
May 1, 2025 positive
Assertion Supported
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
Brandon McKinzie May 1, 2025 ▶ 3:20 No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie
Jul 17, 2025 bullish
Prediction Not checkable as stated
Laskin: Scaling RL on LLMs is the final paradigm before ASI
“The next paradigm, and effectively the final paradigm that we need to have in place before a, you know, what people used to call AGI, or now I think the goalposts have shifted to ASI, is reached, is just figuring out how to scale reinforcement learning on top …”
Misha Laskin Jul 17, 2025 ▶ 7:06 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Jul 17, 2025 bullish
Insight
Laskin: RL requires far fewer FLOPs than pre-training for frontier models
“We're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs, but from a flops perspec…”
Misha Laskin Jul 17, 2025 ▶ 30:41 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Jul 17, 2025 negative
Insight
Laskin: Sensory and vision-language model rewards are far more hackable than LLM rewards
“The challenge is that if we, if you think that language model rewards are hackable vision language model rewards or, you know, like other sensory signal rewards are infinitely more hackable.”
Misha Laskin Jul 17, 2025 ▶ 51:50 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Jul 17, 2025 bullish
Insight
Laskin: Reinforcement learning is the only scalable path for synthetic data
“When we're generating synthetic data there is the only scalable path is really reinforcement learning.”
Misha Laskin Jul 17, 2025 ▶ 53:44 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Jul 24, 2025 negative
Insight
Chen: AI RL Environments Are Too Complex to Create Synthetically
“I think one of the things that people really underestimate is how it is, how complicated it is that you can't just synthetically generate it.”
Edwin Chen Jul 24, 2025 ▶ 14:20 No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen
Dec 5, 2025 positive
Insight
Pereyra: In Legal AI, the RL Environment Is a Client Matter
“And in legal, that RL environment is a client matter. So you have all of the context of a fund formation, an acquisition, a litigation, and the models are starting to learn. Let me go in the document management system and see if I can find this, go in the data…”
Gabe Pereyra Dec 5, 2025 ▶ 8:05 No Priors Ep. 142 | With Harvey Co-Founder and President Gabe Pereyra
Dec 5, 2025 neutral
Insight
Pereyra: Complex legal drafting lacks binary verifiability for AI reward functions
“For something like generate this merger agreement, it's really hard to just give some binary like this is good or this is bad. And I think this has been like a big research problem, like with all the labs we work with, and also internally, there is just this o…”
Gabe Pereyra Dec 5, 2025 ▶ 18:08 No Priors Ep. 142 | With Harvey Co-Founder and President Gabe Pereyra
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.