Oct 20, 2025 · 1h 3m · latent-space

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF

Elie Bakouch · 43m spoken Shawn Wang · 8m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, Hugging Face researcher Elie Bakouch delivers a comprehensive masterclass on open-source LLM pretraining. He covers end-to-end dataset curation, advanced optimizers like Muon, Mixture of Experts (MoE) routing architectures, and the deployment of compact, on-device language models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.6% of the talking time here. How this is scored →

The hosts as informed peer 5.8 Guest teaching 6.3 Guest disagreement 1.7 The hosts pushing back 2.5
05100:0015:0030:0045:001:00:004:08–10:19 · The hosts as informed peer 6/10 A Unified Framework for LLM Pretraining Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters.10:19–20:57 · The hosts as informed peer 5/10 Deep Dive into Optimizers: AdamW, Muon, and Stability Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers.20:58–39:37 · The hosts as informed peer 6/10 Mixture of Experts (MoE) Architectures and Routing Dynamics The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models.39:43–53:00 · The hosts as informed peer 7/10 Dataset Engineering: Rephrasing, Formats, and Epoch Scaling Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits.53:01–57:16 · The hosts as informed peer 5/10 Open Science with SmolLM, Open-Source Tooling, and Training Realities The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations.57:17–1:03:32 · The hosts as informed peer 6/10 On-Device AI, Small Model Ecosystem, Modular, and Conclusion Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction.4:08–10:19 · Guest teaching 6/10 A Unified Framework for LLM Pretraining Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters.10:19–20:57 · Guest teaching 7/10 Deep Dive into Optimizers: AdamW, Muon, and Stability Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers.20:58–39:37 · Guest teaching 8/10 Mixture of Experts (MoE) Architectures and Routing Dynamics The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models.39:43–53:00 · Guest teaching 6/10 Dataset Engineering: Rephrasing, Formats, and Epoch Scaling Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits.53:01–57:16 · Guest teaching 6/10 Open Science with SmolLM, Open-Source Tooling, and Training Realities The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations.57:17–1:03:32 · Guest teaching 5/10 On-Device AI, Small Model Ecosystem, Modular, and Conclusion Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction.4:08–10:19 · Guest disagreement 1/10 A Unified Framework for LLM Pretraining Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters.10:19–20:57 · Guest disagreement 2/10 Deep Dive into Optimizers: AdamW, Muon, and Stability Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers.20:58–39:37 · Guest disagreement 2/10 Mixture of Experts (MoE) Architectures and Routing Dynamics The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models.39:43–53:00 · Guest disagreement 2/10 Dataset Engineering: Rephrasing, Formats, and Epoch Scaling Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits.53:01–57:16 · Guest disagreement 1/10 Open Science with SmolLM, Open-Source Tooling, and Training Realities The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations.57:17–1:03:32 · Guest disagreement 2/10 On-Device AI, Small Model Ecosystem, Modular, and Conclusion Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction.4:08–10:19 · The hosts pushing back 2/10 A Unified Framework for LLM Pretraining Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters.10:19–20:57 · The hosts pushing back 2/10 Deep Dive into Optimizers: AdamW, Muon, and Stability Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers.20:58–39:37 · The hosts pushing back 3/10 Mixture of Experts (MoE) Architectures and Routing Dynamics The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models.39:43–53:00 · The hosts pushing back 4/10 Dataset Engineering: Rephrasing, Formats, and Epoch Scaling Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits.53:01–57:16 · The hosts pushing back 1/10 Open Science with SmolLM, Open-Source Tooling, and Training Realities The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations.57:17–1:03:32 · The hosts pushing back 3/10 On-Device AI, Small Model Ecosystem, Modular, and Conclusion Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 40.5% · guest 59.5%0:00 · the hosts 40.5% · guest 59.5%3:00 · the hosts 14.1% · guest 85.9%3:00 · the hosts 14.1% · guest 85.9%6:00 · the hosts 9% · guest 91%6:00 · the hosts 9% · guest 91%9:00 · the hosts 41.7% · guest 58.3%9:00 · the hosts 41.7% · guest 58.3%12:00 · the hosts 6% · guest 94%12:00 · the hosts 6% · guest 94%15:00 · the hosts 11.3% · guest 88.7%15:00 · the hosts 11.3% · guest 88.7%18:00 · the hosts 18.9% · guest 81.1%18:00 · the hosts 18.9% · guest 81.1%21:00 · the hosts 7.7% · guest 92.3%21:00 · the hosts 7.7% · guest 92.3%24:00 · the hosts 13.8% · guest 86.2%24:00 · the hosts 13.8% · guest 86.2%27:00 · the hosts 22.3% · guest 77.7%27:00 · the hosts 22.3% · guest 77.7%30:00 · the hosts 31.6% · guest 68.4%30:00 · the hosts 31.6% · guest 68.4%33:00 · the hosts 12.9% · guest 87.1%33:00 · the hosts 12.9% · guest 87.1%36:00 · the hosts 30.5% · guest 69.5%36:00 · the hosts 30.5% · guest 69.5%39:00 · the hosts 26.8% · guest 73.2%39:00 · the hosts 26.8% · guest 73.2%42:00 · the hosts 0.3% · guest 99.7%42:00 · the hosts 0.3% · guest 99.7%45:00 · the hosts 32.6% · guest 67.4%45:00 · the hosts 32.6% · guest 67.4%48:00 · the hosts 24.1% · guest 75.9%48:00 · the hosts 24.1% · guest 75.9%51:00 · the hosts 17.3% · guest 82.7%51:00 · the hosts 17.3% · guest 82.7%54:00 · the hosts 22.8% · guest 77.2%54:00 · the hosts 22.8% · guest 77.2%57:00 · the hosts 20.6% · guest 79.4%57:00 · the hosts 20.6% · guest 79.4%1:00:00 · the hosts 64.9% · guest 35.1%1:00:00 · the hosts 64.9% · guest 35.1%1:03:00 · the hosts 73.2% · guest 26.8%1:03:00 · the hosts 73.2% · guest 26.8%
Sharpest disagreement ▶ 16:04 Elie dismisses hyped optimizer speedups

Elie rejects claims of massive speedups in recent optimizer papers, noting that authors frequently undertune the AdamW baseline to exaggerate their gains.

Hardest push from the hosts ▶ 46:20 Swyx pushes back against over-sanitized datasets

Swyx challenges the trend of web rephrasing, arguing that scrubbing all typos and grammatical imperfections leaves models unprepared for real user prompts.

Biggest teaching moment ▶ 23:15 Elie details Megatron's micro-batch load balancing issue

Elie demonstrates how Megatron's micro-batch routing calculation prevented expert domain specialization, educating the hosts on a crucial distributed MoE training detail.

The host holds their own ▶ 1:00:10 Alessio on the broken economics of hosting sub-1B models

Alessio explains the economic and infrastructure bottleneck of serving sub-billion parameter models under existing hourly Hugging Face API pricing.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
A Unified Framework for LLM Pretraining 6612 Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters.
Deep Dive into Optimizers: AdamW, Muon, and Stability 5722 Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers.
Mixture of Experts (MoE) Architectures and Routing Dynamics 6823 The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models.
Dataset Engineering: Rephrasing, Formats, and Epoch Scaling 7624 Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits.
Open Science with SmolLM, Open-Source Tooling, and Training Realities 5611 The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations.
On-Device AI, Small Model Ecosystem, Modular, and Conclusion 6523 Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction.

Statements from this episode (16)

Opinion
Swix: SmolLM 3 is the best open-source AI paper in two years
“I think that, you know, honestly, SmallLM three, one of the single best open source AI research papers I've read, I think probably in like a year, maybe two years.”
Shawn Wang Oct 20, 2025 ▶ 0:50
Disclosure
Hugging Face's pre-training research team averages 30 people
“We also have this team working on pre-training and training models such as small LM. And we are basically a very small team of We have 30 people in average”
Elie Bakouch Oct 20, 2025 ▶ 1:31
Insight
Bakouch: Multi-token prediction is particularly effective for coding models
“Multi-token prediction, which is very good for example, for coding.”
Elie Bakouch Oct 20, 2025 ▶ 6:11
Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Elie Bakouch Oct 20, 2025 ▶ 8:54
Prediction Not checkable as stated
Swix: Frontier model sizes have likely peaked around 2 trillion parameters
“I wonder if we've hit the peak big model craze because now I do expect, you know, 10 trillion model releases, you know, 100 trillion model releases. Probably not. I think we might have peaked at two.”
Shawn Wang Oct 20, 2025 ▶ 9:22
Assertion Not checkable as stated
Bakouch: Novel optimizer speedups are exaggerated due to undertuned AdamW baselines
“And what they find is that the speed up is greatly, greatly exaggerated. And mostly because often people like undertone the Adam W baseline.”
Elie Bakouch Oct 20, 2025 ▶ 16:06
Insight
Bakouch: Optimizer ablations must wait for full learning rate decay
“Especially when doing optimizer ablation, you never want to look at early curve, and you want to, like, wait for the long rate to have fully decay, and, like, do anything to be to zero, or whatever value.”
Elie Bakouch Oct 20, 2025 ▶ 19:01
Assertion Supported
Bakouch: Megatron's local batching hindered MoE expert specialization
“And they find that, for example, in the Megatron item code base this wasn't the case. This was done as the local batch I think. So it basically means that everyone that is using Megatron at the time of the paper I mean, it wasn't even a paper. It was just a bl…”
Elie Bakouch Oct 20, 2025 ▶ 24:46
Disclosure
Bakouch: Hugging Face plans to train an MoE model soon
“For example, we tried we are training MOE currently at TargetFace. I mean, we'll train soon. We start the training soon. And we tried with Megatron and we benchmarked, like, for example, the Mistral architecture with the Queen's three this one.”
Elie Bakouch Oct 20, 2025 ▶ 34:59
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Shawn Wang Oct 20, 2025 ▶ 36:32
Opinion
Bakouch: Open-source AI is not really doing distillation yet
“Actually, distillation is like quite a big thing that people in the open are not really Really doing yet, I think”
Elie Bakouch Oct 20, 2025 ▶ 37:15
Assertion Supported
Bakouch: SmolLM 2 scored random on MMLU until 6.5 trillion tokens
“And this is basically until like 6.5 trillion of tokens, which is a lot, to be honest. Until this amount of token, the MMLU in the QA format, meaning that the model have to select which answer is, the model have to output, for example, the right answer is A, o…”
Elie Bakouch Oct 20, 2025 ▶ 43:42
Assertion Supported
Bakouch: Hugging Face research showed diminishing returns after 3 data epochs
“There is this paper from actually from people at HuginFace at the time that is saying that you basically can repeat your data Like, up to three epochs before seeing diminishing gain before seeing diminishing return.”
Elie Bakouch Oct 20, 2025 ▶ 48:51
Disclosure
Bakouch: Hugging Face introduces code and math in 1T+ token phase
“We are basically introducing higher quality data like I don't know where it is, but if you look at the different phase, we are basically adding more code and math through the train and less web data. But we're only doing it like one point trillion token.”
Elie Bakouch Oct 20, 2025 ▶ 52:36
Disclosure
Bakouch primarily uses closed AI models for deep research out of convenience
“Yeah, I think even I, to be honest, I'm mainly using a closed model because it's easier, basically. It's like the same interface when I'm like, for example, researching very deep on papers.”
Elie Bakouch Oct 20, 2025 ▶ 57:55
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.