“This is something that people have found again and again, that basically you can train a model on some data, or you can train a bigger model on that data and distill that model to another model, and that distill model is better.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Rishabh Agarwal
AssertionSupported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh AgarwalMar 23, 2025▶ 17:41The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
PredictionNot checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh AgarwalMar 23, 2025▶ 42:39The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Distillation Drives Year-Over-Year AI Capability Cost Reductions
“The capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, it's not just because we are doing or figured out something magical. It is because distilla…”
Rishabh AgarwalMar 23, 2025▶ 4:47The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Rishabh AgarwalMar 23, 2025▶ 28:53The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Distillation KL direction dictates trade-off between diversity and performance
“There's a trade-off between diversity and performance. So on the y-axis, I'm showing performance. On the x-axis, I'm showing similarity or basically how, like one minus diversity. So more similar things are less diverse. And you can see, depending on the diver…”
Rishabh AgarwalMar 23, 2025▶ 30:24The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Synthetic data distillation achieves 80% to 90% of target gains
“Try the simplest thing first, which is synthetic data distillation. That already gets you to 80% of the job or 90% of the job.”
Rishabh AgarwalMar 23, 2025▶ 41:42The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.