Insight certainty 5/5 debate potential 1/5

Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches

Rishabh Agarwal · The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind · Mar 23, 2025 · at 14:26

Google DeepMind research scientist Rishabh Agarwal explains the practical advantages of synthetic sequence distillation compared to logit distillation on the Latent Space podcast.

0:00 / 0:21exact quote · 21.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another model. That's like a very good feature. Like, that's the cool thing about, that's why they were able to distill from DeepSeq model to Lama or Quen, and you don't have to even think about that tokenizer”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Rishabh Agarwal

Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Rishabh Agarwal Mar 23, 2025 ▶ 17:41 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Distillation Drives Year-Over-Year AI Capability Cost Reductions
“The capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, it's not just because we are doing or figured out something magical. It is because distilla…”
Rishabh Agarwal Mar 23, 2025 ▶ 4:47 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Distilling a Large Model Outperforms Direct Training on the Same Data
“This is something that people have found again and again, that basically you can train a model on some data, or you can train a bigger model on that data and distill that model to another model, and that distill model is better.”
Rishabh Agarwal Mar 23, 2025 ▶ 5:54 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Rishabh Agarwal Mar 23, 2025 ▶ 28:53 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Distillation KL direction dictates trade-off between diversity and performance
“There's a trade-off between diversity and performance. So on the y-axis, I'm showing performance. On the x-axis, I'm showing similarity or basically how, like one minus diversity. So more similar things are less diverse. And you can see, depending on the diver…”
Rishabh Agarwal Mar 23, 2025 ▶ 30:24 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.