on-policy reinforcement learning

1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 23, 2026 by Yi Tay · across every show →

Everything said about on-policy reinforcement learning, oldest first

Jan 23, 2026 positive
Insight
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”
Yi Tay Jan 23, 2026 ▶ 5:59 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.