Dwarkesh Patel discusses the technical limitations of using reinforcement learning to train autonomous AI agents to perform complex, multi-step real-world tasks.
“People have been talking about long horizon RL, which is the training method. You need to get something like this, where you go tell it to do something and then you reward it at the end for having achieved that outcome. But the difficulty with those kinds of approaches and the difficulty with RL in general is sparse reward and non-stationary distributions, which is like you know, like You failed to book me my right, the right appointments based on like reading out my inbox and like talking with me about it. Why did you fail? There's like so many different reasons you could have failed. That's hard to attribute to any one of them. You know what I mean? It's like hard to learn from that.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dwarkesh Patel
PredictionNot checkable as stated
Patel: Individuals will eventually train superintelligences in basements
“The cost of training, the systems is declining so fast that Literally you will be able to train a super intelligence in a basement at some point in the future.”
Dwarkesh: AGI requires further algorithmic progress, not just current model scaling
“I don't think we're just right around our corner from AGI and it's just a little additional dash of something. That's all it's going to take. I think, you know, people often ask if all AI progress stopped right now and all you could do is collect more data or …”
Dwarkesh: Lack of continual learning prevents LLMs from replacing human labor
“I think a big bottleneck these models have is their inability to learn on the job, to have continual learning. Their entire memory is extinguished at the end of a session. There's a bunch of reasons why I think this actually makes it really hard to get human-l…”
Patel: Online continual learning is not imminent for current AI architectures
“And the reason I don't think that's around the corner is just because there's not, there's no obvious way, at least as far as I can tell, to just slot in this online learning into the models as they exist right now.”
This entire site, over 300 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.