On Policy Distillation

topic on 1 show · 3 statements across 1 episodes

Latent Space

3 statements about On Policy Distillation, every show

LATENT SPACE Assertion Supported
Agarwal: RL literature shows on-policy distillation beats offline methods for agents
“The other thing is in the RL literature, there are, like, results which show that actually this kind of distillation is much more optimal for agentic tasks or really long horizon tasks.”
Rishabh Agarwal Mar 23, 2025 ▶ 38:45 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
LATENT SPACE Disclosure
Agarwal: Google DeepMind Used On-Policy Distillation in Gemma Post-Training
“So this kind of thing was used in Jemma DuPo's training. It's mentioned briefly, like, there's no details there, but it was used there.”
Rishabh Agarwal Mar 23, 2025 ▶ 26:58 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Agarwal: Inverting KL Divergence Direction Derives On-Policy Distillation
“All you need to do, at least in math, if you just flip the direction of your KL divergence, so earlier we were minimizing KL between a teacher and student model, if we just swap that order, because KL is not symmetric, you will actually get this kind of distil…”
Rishabh Agarwal Mar 23, 2025 ▶ 24:37 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.