Insight certainty 3/5 debate potential 3/5

Patel: Delayed Reward Signals Will Slow AI Progress on Long-Horizon Tasks

Dwarkesh Patel · Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition · Jun 18, 2025 · at 20:41

Dwarkesh Patel discusses why the shift from instantaneous token prediction to multi-hour autonomous tasks slows reinforcement learning efficiency.

0:00 / 0:39exact quote · 39.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Now if we're getting to the world where you got to like do a project for seven hours and then at the end of those seven hours, then we tell you, Hey, did you did you get this right? Then like the progress just goes on a bunch. Cause you've gone from like getting signal within the matter of like microseconds to getting signal at the end of seven hours. And so the process of learning has just like become exponentially longer. And I think that might slow down how fast these models, like, you know, the next step now is like, not just being a chat bot, but actually doing real tasks in the world, like completing your taxes, coding, et cetera. And to these things, I think progress might be slower because of this dynamic where it takes a long time for you to learn whether you did the task right or wrong.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dwarkesh Patel

Prediction Not checkable as stated
Patel: Individuals will eventually train superintelligences in basements
“The cost of training, the systems is declining so fast that Literally you will be able to train a super intelligence in a basement at some point in the future.”
Dwarkesh Patel Jun 18, 2025 ▶ 56:49 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Opinion
Dwarkesh: AGI requires further algorithmic progress, not just current model scaling
“I don't think we're just right around our corner from AGI and it's just a little additional dash of something. That's all it's going to take. I think, you know, people often ask if all AI progress stopped right now and all you could do is collect more data or …”
Dwarkesh Patel Jun 18, 2025 ▶ 1:56 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Insight
Dwarkesh: Lack of continual learning prevents LLMs from replacing human labor
“I think a big bottleneck these models have is their inability to learn on the job, to have continual learning. Their entire memory is extinguished at the end of a session. There's a bunch of reasons why I think this actually makes it really hard to get human-l…”
Dwarkesh Patel Jun 18, 2025 ▶ 2:21 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Prediction Not checkable as stated
Patel: Reinforcement learning may not generalize beyond verifiable domains
“I still think I I'm like, I'm not confident that this will generalize to domains that are not so verifiable or text-based.”
Dwarkesh Patel Jun 18, 2025 ▶ 7:34 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Prediction Not checkable as stated
Patel: Online continual learning is not imminent for current AI architectures
“And the reason I don't think that's around the corner is just because there's not, there's no obvious way, at least as far as I can tell, to just slot in this online learning into the models as they exist right now.”
Dwarkesh Patel Jun 18, 2025 ▶ 8:11 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Assertion Not checkable as stated
Patel: Pre-training scaling is seeing diminishing returns
“Pre-training, which is this idea that you just make the model bigger that has had diminishing returns.”
Dwarkesh Patel Jun 18, 2025 ▶ 9:13 Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.