supervised fine-tuning
3 statements across 3 episodes · 1 bullish · 0 bearish · 3 people on the record · first statement Mar 7, 2025 by Misha Laskin · across every show →
Everything said about supervised fine-tuning, oldest first
Mar 7, 2025 neutral
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”
Apr 29, 2025
Jin: Policy gradient algorithms function as weighted supervised fine-tuning
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinfo…”
Jan 23, 2026 positive
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”