Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
Misha Laskin · Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin] · Mar 7, 2025 · at 11:38
Reflection AI CEO Misha Laskin explains why supervised fine-tuning (SFT) and reinforcement learning must be combined when training autonomous coding agents.
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you don't really know what behavior to amplify. And so that's why it's really important in the kind of initial mixture for it to already be doing things out of the box that are somewhat sensible, even if they're unreliable.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →