model training efficiency
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jul 23, 2024 by Thomas Scialom · across every show →
Everything said about model training efficiency, oldest first
Jul 23, 2024 positive
Scialom: Larger tokenizers allow models to see more text per compute unit
“With a bigger vocabulary, for the same text, you have less tokens, right? And so you can train your model on the same amount of knowledge with fewer steps. So for the same compute, you can see more knowledge if you don't epoch.”