Luca Soldaini, AI research scientist at Ai2, details the pre-training data pool and algorithmic sampling strategy for the Dolma 3 dataset used in OLMo 3.
“There's like a pool of about 10 trillion tokens from which we have like an algorithm also fully open source. To like sample about six trillion tokens that we use during training.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Luca Soldaini
AssertionNot checkable as stated
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Luca SoldainiNov 20, 2025▶ 10:52Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Insight
Soldaini: AI scaffolding allows people outside frontier labs to drive capabilities
“If the scaffolding is what really moves a lot of like from, you know, broad capability model to like something that actually has meaningful impact, that scaffolding is not just like, oh, only the labs of people are trained models can do it. Like the number of …”
Luca SoldainiNov 20, 2025▶ 1:26:12Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Disclosure
Ai2 releases OLMo 3 with full training recipes, data, and intermediate checkpoints
“We're not just releasing the final models. We're releasing, you know, the entire recipe we followed to get this model. So the data, the intermediate states, the evaluation frameworks, all the details, all the bits that people need to know to make models like O…”
Luca SoldainiNov 20, 2025▶ 1:46Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
AssertionNot checkable as stated
Soldaini: Frontier AI labs limit final pre-training runs to two months
“I think it's standard practice among the frontier labs to try to cap your big final pre-training run to two months not more than that.”
Luca SoldainiNov 20, 2025▶ 47:21Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Insight
Soldaini: Flawed long-context model architecture cannot be saved by good data
“But they're like technical decisions in how you set up your model that you can have the best data in the world. And your model will not be able to reason over many, many tokens. So it doesn't matter in the sense that you can't train the model on bad data, but …”
Luca SoldainiNov 20, 2025▶ 53:40Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
AssertionNot checkable as stated
Soldaini: 95% of web pages are under 3,000 tokens
“Like 95% web pages are below 3000 tokens.”
Luca SoldainiNov 20, 2025▶ 7:31Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.