teacher model

also referred to as: teacher models

3 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Mar 23, 2025 by Rishabh Agarwal · across every show →

Everything said about teacher model, oldest first

Mar 23, 2025 bullish
Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 neutral
Insight
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Rishabh Agarwal Mar 23, 2025 ▶ 28:53 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Jun 1, 2026 positive
Insight
He: Distillation works because teacher models are simpler than the internet
“I guess the, from the modeling perspective, the strong model, the teacher model is trying to model The image and videos of entire internet. And that distribution is extremely complex. As a step distilled model is just trying to learn from the teacher. The teac…”
Ethan He Jun 1, 2026 ▶ 39:48 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.