Teacher Model

topic on 1 show · 3 statements across 2 episodes

Latent Space

3 statements about Teacher Model, every show

He: Distillation works because teacher models are simpler than the internet
“I guess the, from the modeling perspective, the strong model, the teacher model is trying to model The image and videos of entire internet. And that distribution is extremely complex. As a step distilled model is just trying to learn from the teacher. The teac…”
Ethan He Jun 1, 2026 ▶ 39:48 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
LATENT SPACE Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Rishabh Agarwal Mar 23, 2025 ▶ 28:53 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.