Insight certainty 4/5 debate potential 2/5

Laskin: Current RL algorithms lack atomic credit assignment, causing meandering reasoning

Misha Laskin · No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin · Jul 17, 2025 · at 36:33

Misha Laskin, co-founder and CEO of Reflection AI, breaks down why current reasoning models explore unnecessary reasoning paths during RL training.

0:00 / 0:40exact quote · 40.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The RL methods we have today are quite bad, I would say, exploration and credit assignment. Like they, they're sort of just like the fundamental algorithms are take the things that work and make them happen more frequently, and the things that don't work and have, and make them happen less frequently, but they don't discern at all along your, say, reasoning chain which part of the reasoning was correct and which part was incorrect, and so that's why you get these reasoning chains that are kind of garden path meandering, like they'll explore all sorts of things that are, you know, completely unnecessary and don't look at all Like the kind of structured thinking that a person would have. That's how the algorithm works. It doesn't it doesn't actually look at, there's no credit assignment step on any atomic level.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Misha Laskin

Opinion
Laskin: Enterprise AI coding tool productivity impact is negligible or negative
“Within enterprises, when you know, they're adopting coding tools and you see the impact that this is having on their actual productivity. And I think it's much lower than people expect. So it's in fact, it's sometimes negative, sometimes negligible.”
Misha Laskin Jul 17, 2025 ▶ 8:40 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Opinion
Laskin: Teaching AI agents to take action is mostly solved
“To me, it seems like really, 20% of the problem is teaching these agents how to act, and it's more or less solved.”
Misha Laskin Jul 17, 2025 ▶ 11:10 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Insight
Laskin: New frontier labs can succeed without cloud provider ownership
“Our thought was that this was the time where you can actually start a you know, a generational frontier lab that does not need to be coupled to a, you know, to a big cloud provider because if you do it right, you'll actually be able to generate you know, suffi…”
Misha Laskin Jul 17, 2025 ▶ 31:15 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Insight
Laskin: Machine learning generalization is just bringing test distribution into training
“There's no such thing as generalization. There's just bringing the test distribution into train.”
Misha Laskin Jul 17, 2025 ▶ 37:49 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Prediction Not checkable as stated
Laskin: Scaling RL on LLMs is the final paradigm before ASI
“The next paradigm, and effectively the final paradigm that we need to have in place before a, you know, what people used to call AGI, or now I think the goalposts have shifted to ASI, is reached, is just figuring out how to scale reinforcement learning on top …”
Misha Laskin Jul 17, 2025 ▶ 7:06 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Opinion
Laskin: Humanity's Last Exam Benchmark Barely Matters to End Users
“Now, that's great, but I think the downside of that is that does humanity's last exam actually matter in any meaningful way for an end user? And I would argue that some weak correlation, but the answer is most likely no.”
Misha Laskin Jul 17, 2025 ▶ 15:53 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.