distillation

14 statements across 5 episodes · 8 bullish · 2 bearish · 6 people on the record · first statement Oct 13, 2024 by Vibhu Sapra · across every show →

Everything said about distillation, oldest first

Oct 13, 2024 negative
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Nov 2, 2024 neutral
Assertion Supported
Consistency Models Still Underperform Diffusion Teacher Models on Standard Benchmarks
“The distillation doesn't do quite as well as the training. And, but neither of them do as well as the diffusion teacher. Including the one that was trained from scratch.”
RJ Honicky Nov 2, 2024 ▶ 48:15 [Paper Club] Intro to Diffusion Models and OpenAI sCM: Simple, Stable, Scalable Consistency Models
Mar 23, 2025 positive
Insight
Agarwal: Optimal Post-Training Pipeline Combines Heavy Distillation Followed by RL
“So, so I would think maybe an optimal pipeline would look like you do distillation heavily, but then you still do some RL afterwards, because maybe there's still something you can get out of your reward functions or whatever your post-training stack is.”
Rishabh Agarwal Mar 23, 2025 ▶ 8:05 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: Standard RLHF Infrastructure Can Be Repurposed for Model Distillation
“First step is you go to an RLHF, RLXF, whatever framework you have. You turn off the reward term. So you delete the reward part of it. All you are left is some KL term. And now you swap your KL, the anchor policy, to a bigger teacher policy. And there you go. …”
Rishabh Agarwal Mar 23, 2025 ▶ 31:54 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Assertion Supported
Agarwal: Gemma 2 Used Soft-Label Logit Distillation During Pre-Training
“GemRTool used distillation for pre-training, where they used logits, or these soft labels, which is rather than having hard zero, one tokens, which is, I want to predict this next token, they have like soft labels for all possible tokens.”
Rishabh Agarwal Mar 23, 2025 ▶ 10:05 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 neutral
Insight
Swyx: The Cost-Performance Pareto Frontier Slope Reflects Distillation SOTA
“The slope of the Pareto Frontier is basically the state of the art of distillation, and the gentler the slope, the better distillation is.”
Shawn Wang Mar 23, 2025 ▶ 5:11 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: 50/50 mix of KL divergences usually works when goals are unclear
“The general recommendation I would give people is that maybe use a mixture of half and half. Like that's what some people have used, right? That's like saying, yeah, basically saying, I don't know what I want. I just want something to work well enough. I'll ju…”
Rishabh Agarwal Mar 23, 2025 ▶ 31:08 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 neutral
Insight
Agarwal: Distillation KL direction dictates trade-off between diversity and performance
“There's a trade-off between diversity and performance. So on the y-axis, I'm showing performance. On the x-axis, I'm showing similarity or basically how, like one minus diversity. So more similar things are less diverse. And you can see, depending on the diver…”
Rishabh Agarwal Mar 23, 2025 ▶ 30:24 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: Distillation Drives Year-Over-Year AI Capability Cost Reductions
“The capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, it's not just because we are doing or figured out something magical. It is because distilla…”
Rishabh Agarwal Mar 23, 2025 ▶ 4:47 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: RLHF and LLM distillation can be combined simultaneously
“You can actually combine RLHF and distillation together, because now you're doing two things at the same time.”
Rishabh Agarwal Mar 23, 2025 ▶ 33:17 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: Distillation accelerates speculative decoding for large models
“Now, the thing is, the effectiveness of this method depends on how close the sampler, the small model is to the bigger model that we want to speed up, and actually distillation exactly fixes that, which is, by distillation, you can make things closer to each o…”
Rishabh Agarwal Mar 23, 2025 ▶ 34:57 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Oct 20, 2025 negative
Opinion
Bakouch: Open-source AI is not really doing distillation yet
“Actually, distillation is like quite a big thing that people in the open are not really Really doing yet, I think”
Elie Bakouch Oct 20, 2025 ▶ 37:15 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Feb 12, 2026 positive
Insight
Jeff Dean: Teacher model logits enable small models to learn from multi-pass training
“One of the key advantages of distillation is that you can have a much smaller model And you can have a very large you know, training data set and you can get utility out of making many passes over that data set because you're now getting the logits from the mu…”
Jeff Dean Feb 12, 2026 ▶ 6:02 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Feb 12, 2026
Insight
Jeff Dean: Capable small models require first building frontier models
“Through distillation, which is a key technique for making the smaller models more capable, you know, you have to have the frontier model in order to then distill it into your smaller model. So it's not like an either or choice. You sort of need that in order t…”
Jeff Dean Feb 12, 2026 ▶ 3:06 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.