Oct 20, 2025 · 1h 3m · latent-space
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space Podcast, Hugging Face researcher Elie Bakouch delivers a comprehensive masterclass on open-source LLM pretraining. He covers end-to-end dataset curation, advanced optimizers like Muon, Mixture of Experts (MoE) routing architectures, and the deployment of compact, on-device language models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Elie rejects claims of massive speedups in recent optimizer papers, noting that authors frequently undertune the AdamW baseline to exaggerate their gains.
Hardest push from the hosts ▶ 46:20 Swyx pushes back against over-sanitized datasetsSwyx challenges the trend of web rephrasing, arguing that scrubbing all typos and grammatical imperfections leaves models unprepared for real user prompts.
Biggest teaching moment ▶ 23:15 Elie details Megatron's micro-batch load balancing issueElie demonstrates how Megatron's micro-batch routing calculation prevented expert domain specialization, educating the hosts on a crucial distributed MoE training detail.
The host holds their own ▶ 1:00:10 Alessio on the broken economics of hosting sub-1B modelsAlessio explains the economic and infrastructure bottleneck of serving sub-billion parameter models under existing hourly Hugging Face API pricing.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| A Unified Framework for LLM Pretraining | 6 | 6 | 1 | 2 | Alessio and Swyx ask targeted questions about pretraining stages and model scaling frontiers, while Elie outlines his five-pillar framework for LLM training and points out anomalies like DeepSeek using Llama 2 Adam parameters. | |
| Deep Dive into Optimizers: AdamW, Muon, and Stability | 5 | 7 | 2 | 2 | Swyx probes into the mechanics of newer optimizers like Muon and Defazio's schedule-free approach, prompting Elie to explain Newton-Schulz orthogonalization and debunk exaggerated speedup claims in recent papers. | |
| Mixture of Experts (MoE) Architectures and Routing Dynamics | 6 | 8 | 2 | 3 | The hosts discuss MoE sparsity, distillation, and routing mechanics. Elie dives deep into global vs. local batch load-balancing statistics in Megatron and explains shared vs. zero-communication experts in cutting-edge models. | |
| Dataset Engineering: Rephrasing, Formats, and Epoch Scaling | 7 | 6 | 2 | 4 | Swyx actively pushes on the pitfalls of synthetic rephrasing making datasets 'too clean' to handle typos and user errors. Elie agrees with the premise and shares empirical findings on conversational format impact and epoch repetition limits. | |
| Open Science with SmolLM, Open-Source Tooling, and Training Realities | 5 | 6 | 1 | 1 | The conversation turns to open science practices, Datatrove, LightEval, and training war stories, with Elie showing loss curve spikes from bugged optimizer ablations. | |
| On-Device AI, Small Model Ecosystem, Modular, and Conclusion | 6 | 5 | 2 | 3 | Alessio and Swyx probe into deployment economics and whether Elie has evaluated Modular's stack. Elie candidly admits he hasn't looked closely into Mojo/Modular because changing pretraining codebases is high friction. |