on-policy reinforcement learning
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 23, 2026 by Yi Tay · across every show →
Everything said about on-policy reinforcement learning, oldest first
Jan 23, 2026 positive
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”