reinforcement learning
also referred to as: rl
19 statements across 13 episodes · 12 bullish · 4 bearish · 13 people on the record · first statement Apr 25, 2023 by Noam Brown · across every show →
Everything said about reinforcement learning, oldest first
Apr 25, 2023 bearish
Reinforcement learning struggles in trading because financial markets are non-stationary
“I think the major challenge with Using things like reinforcement learning for trading is that it's a non-stationary environment. So you can have all this historical data, but it's not a stationary system and it's gonna like the markets respond to world events,…”
Jun 1, 2023
Schimpf: Bounding non-deterministic AI is essential for military tech adoption
“A lot of these areas that are more non-deterministic, so, you know, things like reinforcement learning or, you know, potential applications of LOMs into the space, they are inherently non-deterministic, and that is a risk. And so kind of quantifying that, know…”
Jan 24, 2024 negative
Chen: Next-token LLM training is mimicry, imposing a natural capability ceiling
“For people that study reinforcement learning, we call it behavior cloning, which means you're just asking the AI to clone the behavior of another agent. And that is like one of the most primitive way possible to train this type of systems. Like, because if you…”
Mar 7, 2024 bullish
Gil: New reinforcement learning AI agent products will emerge within 6-12 months
“And so I think that that purpose of knowledge is about to hit the world in the context of new products. And it'll take time for those products to emerge, you know, six months, 12 months, a year. But It does feel like that's another wave that's coming where you…”
Aug 1, 2024 neutral
Vinyals: AlphaGo compute was mostly RL self-play, unlike modern LLM pre-training dominance
“Historically, if you look at AlphaGo, which actually followed quite closely the recipe of you pre-train your model on all human data, you then use RL to make it better, and then you do some search at inference time, the compute there was very skewed for the mi…”
Aug 1, 2024 positive
Vinyals: LLM bootstrapping works if verification is easier than solution generation
“If checking that something is correct is easier than creating the solution, then we're in business because the language models will be able to evaluate their own samples more accurately than to generate them. And then we have a sort of reinforcement learning l…”
Aug 1, 2024 bullish
Vinyals predicts pre-training compute will drop to ~50% as RL expands
“So to me, that balance feels correct, like some on pre-training, and here we, we're trying to learn every task. So certainly that's going to be, you know, let's say it can be as high as 50%, not as high as over 90 like today. And then the rest mostly on reinfo…”
Nov 14, 2024 bullish
Mehta: Specialist data seeds LLMs, but RL drives superhuman capability
“The specialist humans are gonna serve to, like, get the LLM from, like, just a bunch of weights that knows how to do nothing to, like, something that, like is, like, surprisingly strong. And then the RL is gonna take you from there to, like, something that's, …”
Mar 20, 2025 bullish
Chelsea Finn: Autonomous RL experience will play a huge role in robotics
“And then I also think that autonomous experience will play a huge role, just like we've seen in language models. After you get an initial language model, if you can use reinforcement learning to have the robot, the language model bootstrap on its own experienc…”
Apr 10, 2025 positive
Foody: Model Improvement via RL Is Gated Entirely by Evaluation Benchmarks
“Reinforcement learning is becoming so effective that once you create evals, the models can learn them and how to you know, improve capabilities. And so for everything that we want alums to be good at, we need evals for those things.”
Apr 24, 2025 positive
RL models only need task and outcome definitions to learn research trajectories
“The cool thing with RL is that you don't necessarily need to
Know the whole process of how the person would do the research.
You just have to know what the task is and what the outcome should be, and the model will just learn during training how to get from th…”
May 1, 2025 positive
McKinzie: Reinforcement learning is the key differentiator behind o3 reasoning
“I guess the short answer is reinforcement learning is, is the biggest one. So yeah, rather than just having to predict the next token and some large pre-training corpus from, you know you know, everywhere essentially now we have a more focused goal of the mode…”
Jul 17, 2025 bullish
Laskin: Scaling RL on LLMs is the final paradigm before ASI
“The next paradigm, and effectively the final paradigm that we need to have in place before a, you know, what people used to call AGI, or now I think the goalposts have shifted to ASI, is reached, is just figuring out how to scale reinforcement learning on top …”
Jul 17, 2025 bullish
Laskin: RL requires far fewer FLOPs than pre-training for frontier models
“We're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs, but from a flops perspec…”
Jul 17, 2025 negative
Jul 17, 2025 bullish
Jul 24, 2025 negative
Dec 5, 2025 positive
Pereyra: In Legal AI, the RL Environment Is a Client Matter
“And in legal, that RL environment is a client matter. So you have all of the context of a fund formation, an acquisition, a litigation, and the models are starting to learn. Let me go in the document management system and see if I can find this, go in the data…”
Dec 5, 2025 neutral
Pereyra: Complex legal drafting lacks binary verifiability for AI reward functions
“For something like generate this merger agreement, it's really hard to just give some binary like this is good or this is bad. And I think this has been like a big research problem, like with all the labs we work with, and also internally, there is just this o…”