On Policy Distillation
topic on 1 show · 3 statements across 1 episodes
3 statements about On Policy Distillation, every show
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Agarwal: Google DeepMind Used On-Policy Distillation in Gemma Post-Training
“So this kind of thing was used in Jemma DuPo's training. It's mentioned briefly, like, there's no details there, but it was used there.”
Agarwal: Inverting KL Divergence Direction Derives On-Policy Distillation
“All you need to do, at least in math, if you just flip the direction of your KL divergence, so earlier we were minimizing KL between a teacher and student model, if we just swap that order, because KL is not symmetric, you will actually get this kind of distil…”