Elie Bakouch

Pre-training Lead, Hugging Face · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

Elie Bakouch leads pre-training efforts at Hugging Face and is an architect behind SmolLM. He specializes in large language model architecture design, training stability, and dataset optimization.

13statements → 6claims → 4claims resolved → 100%fully supported → 3.92/5average certainty → 2.08/5average debate potential → ≈4.0/5argument clarity, estimated →

4 supported 0 partly supported 0 contradicted 2 not checkable as stated how the 6 claims stand · each chip opens the sources

6 assertions · 1 opinion · 2 insights · 4 disclosures · every statement was checked. The predictions and assertions are the 6 claims: statements the public record can support or contradict. 4 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Elie argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Bakouch: Megatron's local batching hindered MoE expert specialization
“And they find that, for example, in the Megatron item code base this wasn't the case. This was done as the local batch I think. So it basically means that everyone that is using Megatron at the time of the paper I mean, it wasn't even a paper. It was just a bl…”
Elie Bakouch Oct 20, 2025 ▶ 24:46 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Elie Bakouch on measured tape to publish a rate. This says nothing about how they speak.

Everything Elie Bakouch said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Bakouch: Novel optimizer speedups are exaggerated due to undertuned AdamW baselines
“And what they find is that the speed up is greatly, greatly exaggerated. And mostly because often people like undertone the Adam W baseline.”
Elie Bakouch Oct 20, 2025 ▶ 16:06 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Disclosure
Bakouch primarily uses closed AI models for deep research out of convenience
“Yeah, I think even I, to be honest, I'm mainly using a closed model because it's easier, basically. It's like the same interface when I'm like, for example, researching very deep on papers.”
Elie Bakouch Oct 20, 2025 ▶ 57:55 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Bakouch: Megatron's local batching hindered MoE expert specialization
“And they find that, for example, in the Megatron item code base this wasn't the case. This was done as the local batch I think. So it basically means that everyone that is using Megatron at the time of the paper I mean, it wasn't even a paper. It was just a bl…”
Elie Bakouch Oct 20, 2025 ▶ 24:46 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Opinion
Bakouch: Open-source AI is not really doing distillation yet
“Actually, distillation is like quite a big thing that people in the open are not really Really doing yet, I think”
Elie Bakouch Oct 20, 2025 ▶ 37:15 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Bakouch: SmolLM 2 scored random on MMLU until 6.5 trillion tokens
“And this is basically until like 6.5 trillion of tokens, which is a lot, to be honest. Until this amount of token, the MMLU in the QA format, meaning that the model have to select which answer is, the model have to output, for example, the right answer is A, o…”
Elie Bakouch Oct 20, 2025 ▶ 43:42 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Insight
Bakouch: Multi-token prediction is particularly effective for coding models
“Multi-token prediction, which is very good for example, for coding.”
Elie Bakouch Oct 20, 2025 ▶ 6:11 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Elie Bakouch Oct 20, 2025 ▶ 8:54 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Insight
Bakouch: Optimizer ablations must wait for full learning rate decay
“Especially when doing optimizer ablation, you never want to look at early curve, and you want to, like, wait for the long rate to have fully decay, and, like, do anything to be to zero, or whatever value.”
Elie Bakouch Oct 20, 2025 ▶ 19:01 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Disclosure
Bakouch: Hugging Face plans to train an MoE model soon
“For example, we tried we are training MOE currently at TargetFace. I mean, we'll train soon. We start the training soon. And we tried with Megatron and we benchmarked, like, for example, the Mistral architecture with the Queen's three this one.”
Elie Bakouch Oct 20, 2025 ▶ 34:59 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Bakouch: Hugging Face research showed diminishing returns after 3 data epochs
“There is this paper from actually from people at HuginFace at the time that is saying that you basically can repeat your data Like, up to three epochs before seeing diminishing gain before seeing diminishing return.”
Elie Bakouch Oct 20, 2025 ▶ 48:51 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Disclosure
Hugging Face's pre-training research team averages 30 people
“We also have this team working on pre-training and training models such as small LM. And we are basically a very small team of We have 30 people in average”
Elie Bakouch Oct 20, 2025 ▶ 1:31 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Disclosure
Bakouch: Hugging Face introduces code and math in 1T+ token phase
“We are basically introducing higher quality data like I don't know where it is, but if you look at the different phase, we are basically adding more code and math through the train and less web data. But we're only doing it like one point trillion token.”
Elie Bakouch Oct 20, 2025 ▶ 52:36 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF

Appearances (1)

EpisodeDateSpeaking time
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePD Oct 20, 2025 43m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.