teacher model
also referred to as: teacher models
3 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Mar 23, 2025 by Rishabh Agarwal · across every show →
Everything said about teacher model, oldest first
Mar 23, 2025 bullish
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Mar 23, 2025 neutral
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”
Jun 1, 2026 positive
He: Distillation works because teacher models are simpler than the internet
“I guess the, from the modeling perspective, the strong model, the teacher model is trying to model The image and videos of entire internet. And that distribution is extremely complex. As a step distilled model is just trying to learn from the teacher. The teac…”