“There is this paper from actually from people at HuginFace at the time that is saying that you basically can repeat your data Like, up to three epochs before seeing diminishing gain before seeing diminishing return.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Elie Bakouch
AssertionNot checkable as stated
Bakouch: Novel optimizer speedups are exaggerated due to undertuned AdamW baselines
“And what they find is that the speed up is greatly, greatly exaggerated. And mostly because often people like undertone the Adam W baseline.”
Elie BakouchOct 20, 2025▶ 16:06⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Disclosure
Bakouch primarily uses closed AI models for deep research out of convenience
“Yeah, I think even I, to be honest, I'm mainly using a closed model because it's easier, basically. It's like the same interface when I'm like, for example, researching very deep on papers.”
Elie BakouchOct 20, 2025▶ 57:55⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
AssertionNot checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie BakouchOct 20, 2025▶ 59:13⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
AssertionSupported
Bakouch: Megatron's local batching hindered MoE expert specialization
“And they find that, for example, in the Megatron item code base this wasn't the case. This was done as the local batch I think. So it basically means that everyone that is using Megatron at the time of the paper I mean, it wasn't even a paper. It was just a bl…”
Elie BakouchOct 20, 2025▶ 24:46⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Opinion
Bakouch: Open-source AI is not really doing distillation yet
“Actually, distillation is like quite a big thing that people in the open are not really
Really doing yet, I think”
Elie BakouchOct 20, 2025▶ 37:15⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
AssertionSupported
Bakouch: SmolLM 2 scored random on MMLU until 6.5 trillion tokens
“And this is basically until like 6.5 trillion of tokens, which is a lot, to be honest. Until this amount of token, the MMLU in the QA format, meaning that the model have to select which answer is, the model have to output, for example, the right answer is A, o…”
Elie BakouchOct 20, 2025▶ 43:42⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.