Pre Training Data
topic on 4 shows · 5 statements across 5 episodes
the Y Combinator Startup Podcast
Latent Space
Lenny's Podcast
the MAD Podcast
5 statements about Pre Training Data, every show
Yi Tay: AI Labs Prioritize Data Efficiency Because the World Lacks Tokens
“I think in general, the, like learning more, like extracting more from varied data points is definitely valuable, but I think that's what related to the fact that we're like running out of Tokens in the world.”
Soldaini: Mid-training requires re-mixing pre-training data to avoid model forgetting
“When you do that, you also need to make sure that The model doesn't forget stuff that I've seen during pre-training, so that's why, like, you mix some of the best data from pre-training, you do carry over.”
Foody: Efficient post-training datasets, not 10x pre-training, drive AI progress
“And it's not going to be, you know, 10 X more pre-training data that gets those capabilities. It's much more going to be all of the post-training data sets that are far more data efficient and thoughtful that help us get there.”
Musk: AI developers have run out of high-quality human pre-training data
“We, we've kind of run out of pre-training data or human generated pre like human generated data. You run out of tokens pretty fast certainly of high quality tokens.”
Huang: Adding one billion tokens cannot teach trillion-token models new knowledge
“All models these days are now double-digit trillions, right? So it's kind of a drop in the bucket if you really think I can just put, you know, a billion tokens in there, and I actually think that the model's gonna truly learn new Information.”