KL Divergence

topic on 1 show · 3 statements across 1 episodes

Latent Space

3 statements about KL Divergence, every show

Agarwal: 50/50 mix of KL divergences usually works when goals are unclear
“The general recommendation I would give people is that maybe use a mixture of half and half. Like that's what some people have used, right? That's like saying, yeah, basically saying, I don't know what I want. I just want something to work well enough. I'll ju…”
Rishabh Agarwal Mar 23, 2025 ▶ 31:08 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Agarwal: Distillation KL direction dictates trade-off between diversity and performance
“There's a trade-off between diversity and performance. So on the y-axis, I'm showing performance. On the x-axis, I'm showing similarity or basically how, like one minus diversity. So more similar things are less diverse. And you can see, depending on the diver…”
Rishabh Agarwal Mar 23, 2025 ▶ 30:24 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Agarwal: Inverting KL Divergence Direction Derives On-Policy Distillation
“All you need to do, at least in math, if you just flip the direction of your KL divergence, so earlier we were minimizing KL between a teacher and student model, if we just swap that order, because KL is not symmetric, you will actually get this kind of distil…”
Rishabh Agarwal Mar 23, 2025 ▶ 24:37 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.