supervised fine-tuning

3 statements across 3 episodes · 1 bullish · 0 bearish · 3 people on the record · first statement Mar 7, 2025 by Misha Laskin · across every show →

Everything said about supervised fine-tuning, oldest first

Mar 7, 2025 neutral
Insight
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”
Misha Laskin Mar 7, 2025 ▶ 11:38 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Apr 29, 2025
Insight
Jin: Policy gradient algorithms function as weighted supervised fine-tuning
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinfo…”
Roger Jin Apr 29, 2025 ▶ 5:59 What is an RL environment? w/ Nous Research's Roger Jin
Jan 23, 2026 positive
Insight
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”
Yi Tay Jan 23, 2026 ▶ 5:59 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.