Teacher Model
topic on 1 show · 3 statements across 2 episodes
3 statements about Teacher Model, every show
He: Distillation works because teacher models are simpler than the internet
“I guess the, from the modeling perspective, the strong model, the teacher model is trying to model The image and videos of entire internet. And that distribution is extremely complex. As a step distilled model is just trying to learn from the teacher. The teac…”
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Agarwal: Standard distillation KL loss places mass where teacher has none
“The typically what we use is this mode covering KL, the one on the left. That is the standard distillation loss that everyone uses. But you can see the behavior. And you can already see what's weird about it. It's putting a lot of mass on places where there's …”