Pre Training Data

topic on 4 shows · 5 statements across 5 episodes

the Y Combinator Startup Podcast Latent Space Lenny's Podcast the MAD Podcast

5 statements about Pre Training Data, every show

LATENT SPACE Assertion Not checkable as stated
Yi Tay: AI Labs Prioritize Data Efficiency Because the World Lacks Tokens
“I think in general, the, like learning more, like extracting more from varied data points is definitely valuable, but I think that's what related to the fact that we're like running out of Tokens in the world.”
Yi Tay Jan 23, 2026 ▶ 57:07 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
MAD Insight
Soldaini: Mid-training requires re-mixing pre-training data to avoid model forgetting
“When you do that, you also need to make sure that The model doesn't forget stuff that I've seen during pre-training, so that's why, like, you mix some of the best data from pre-training, you do carry over.”
Luca Soldaini Nov 20, 2025 ▶ 50:58 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Foody: Efficient post-training datasets, not 10x pre-training, drive AI progress
“And it's not going to be, you know, 10 X more pre-training data that gets those capabilities. It's much more going to be all of the post-training data sets that are far more data efficient and thoughtful that help us get there.”
Brendan Foody Sep 18, 2025 ▶ 58:51 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
Y COMBINATOR Assertion Not checkable as stated
Musk: AI developers have run out of high-quality human pre-training data
“We, we've kind of run out of pre-training data or human generated pre like human generated data. You run out of tokens pretty fast certainly of high quality tokens.”
Elon Musk Jun 19, 2025 ▶ 32:56 Elon Musk: Digital Superintelligence, Multiplanetary Life, How to Be Useful · Y Combinator
Huang: Adding one billion tokens cannot teach trillion-token models new knowledge
“All models these days are now double-digit trillions, right? So it's kind of a drop in the bucket if you really think I can just put, you know, a billion tokens in there, and I actually think that the model's gonna truly learn new Information.”
Mark Huang May 31, 2024 ▶ 31:49 How to train a Million Context LLM — with Mark Huang of Gradient.ai

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.