on-policy distillation

3 statements across 1 episodes · 3 bullish · 0 bearish · 1 people on the record · first statement Mar 23, 2025 by Rishabh Agarwal · across every show →

Everything said about on-policy distillation, oldest first

Mar 23, 2025 positive
Insight
Agarwal: Inverting KL Divergence Direction Derives On-Policy Distillation
“All you need to do, at least in math, if you just flip the direction of your KL divergence, so earlier we were minimizing KL between a teacher and student model, if we just swap that order, because KL is not symmetric, you will actually get this kind of distil…”
Rishabh Agarwal Mar 23, 2025 ▶ 24:37 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Disclosure
Agarwal: Google DeepMind Used On-Policy Distillation in Gemma Post-Training
“So this kind of thing was used in Jemma DuPo's training. It's mentioned briefly, like, there's no details there, but it was used there.”
Rishabh Agarwal Mar 23, 2025 ▶ 26:58 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Assertion Supported
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Rishabh Agarwal Mar 23, 2025 ▶ 38:45 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.