Dec 24, 2024 · 28m · latent-space

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]

Loubna Ben Allal · 24m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At NeurIPS 2024, Loubna Ben Allal of Hugging Face presents a comprehensive overview of synthetic data generation pipelines and the rapid advancement of small, on-device language models. She demonstrates how synthetic pre-training, intelligent filtering, and compute-efficient over-training enable compact architectures like SmolLM2 to deliver high-performance, private, and cost-effective edge AI solutions.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.0 Guest teaching 0.0 Guest disagreement 0.5 The hosts pushing back 0.0
05100:0010:0020:000:06–2:47 · The hosts as informed peer 0/10 Opening Remarks and Talk Agenda Solo presentation opening outlining the evolution of synthetic data from post-training to pre-training and evaluation. Host scores remain zero as this is a monologue presentation.2:49–5:17 · The hosts as informed peer 0/10 Demystifying Model Collapse and Web Data Pollution Loubna refutes conventional media and academic panic regarding model collapse and web pollution, presenting empirical data from FineWeb showing newer dumps perform better.5:18–9:13 · The hosts as informed peer 0/10 Synthetic Pre-Training and the Cosmopedia Approach Loubna discusses reproducing Microsoft's Phi methodology with Cosmopedia, explaining prompt seeding strategies to ensure synthetic data diversity.9:15–11:31 · The hosts as informed peer 0/10 Web Rephrasing and Programmatic Data Refinement Technical breakdown of web rephrasing techniques using small models and programmatic refinement from Nemotron CC and ProX.11:33–13:49 · The hosts as informed peer 0/10 LLM-Based Quality Classifiers for Large-Scale Filtering Detailed walkthrough of FineWeb-Edu and DCLM classifier-based quality filtering pipelines.13:51–16:52 · The hosts as informed peer 0/10 Post-Training Advancements: AgentInstruct, TÜLU 3, and SmolTalk Review of post-training datasets including AgentInstruct, TÜLU 3, SmolTalk, and multilingual data arbitrage.16:52–19:11 · The hosts as informed peer 0/10 The Emergence of High-Performing Smol and Edge Models Demonstration of competitive small models running on consumer hardware like iPhones via PocketPal.19:12–24:01 · The hosts as informed peer 0/10 Rethinking Scaling Laws: Inference Economics and Long Training Loubna challenges the 'bigger is always better' scaling paradigm, arguing for overtraining smaller architectures due to inference economics.24:03–26:12 · The hosts as informed peer 0/10 On-Device Deployment, Tooling, and Structured Generation Discussion of on-device advantages, privacy, and JSON-constrained structured text generation in browser.26:14–28:05 · The hosts as informed peer 0/10 Future Trends, The Return of Fine-Tuning, and Q&A Loubna concludes with a hot take predicting the industry will cycle back to fine-tuning specialized small models over prompt engineering.0:06–2:47 · Guest teaching 0/10 Opening Remarks and Talk Agenda Solo presentation opening outlining the evolution of synthetic data from post-training to pre-training and evaluation. Host scores remain zero as this is a monologue presentation.2:49–5:17 · Guest teaching 0/10 Demystifying Model Collapse and Web Data Pollution Loubna refutes conventional media and academic panic regarding model collapse and web pollution, presenting empirical data from FineWeb showing newer dumps perform better.5:18–9:13 · Guest teaching 0/10 Synthetic Pre-Training and the Cosmopedia Approach Loubna discusses reproducing Microsoft's Phi methodology with Cosmopedia, explaining prompt seeding strategies to ensure synthetic data diversity.9:15–11:31 · Guest teaching 0/10 Web Rephrasing and Programmatic Data Refinement Technical breakdown of web rephrasing techniques using small models and programmatic refinement from Nemotron CC and ProX.11:33–13:49 · Guest teaching 0/10 LLM-Based Quality Classifiers for Large-Scale Filtering Detailed walkthrough of FineWeb-Edu and DCLM classifier-based quality filtering pipelines.13:51–16:52 · Guest teaching 0/10 Post-Training Advancements: AgentInstruct, TÜLU 3, and SmolTalk Review of post-training datasets including AgentInstruct, TÜLU 3, SmolTalk, and multilingual data arbitrage.16:52–19:11 · Guest teaching 0/10 The Emergence of High-Performing Smol and Edge Models Demonstration of competitive small models running on consumer hardware like iPhones via PocketPal.19:12–24:01 · Guest teaching 0/10 Rethinking Scaling Laws: Inference Economics and Long Training Loubna challenges the 'bigger is always better' scaling paradigm, arguing for overtraining smaller architectures due to inference economics.24:03–26:12 · Guest teaching 0/10 On-Device Deployment, Tooling, and Structured Generation Discussion of on-device advantages, privacy, and JSON-constrained structured text generation in browser.26:14–28:05 · Guest teaching 0/10 Future Trends, The Return of Fine-Tuning, and Q&A Loubna concludes with a hot take predicting the industry will cycle back to fine-tuning specialized small models over prompt engineering.0:06–2:47 · Guest disagreement 0/10 Opening Remarks and Talk Agenda Solo presentation opening outlining the evolution of synthetic data from post-training to pre-training and evaluation. Host scores remain zero as this is a monologue presentation.2:49–5:17 · Guest disagreement 2/10 Demystifying Model Collapse and Web Data Pollution Loubna refutes conventional media and academic panic regarding model collapse and web pollution, presenting empirical data from FineWeb showing newer dumps perform better.5:18–9:13 · Guest disagreement 1/10 Synthetic Pre-Training and the Cosmopedia Approach Loubna discusses reproducing Microsoft's Phi methodology with Cosmopedia, explaining prompt seeding strategies to ensure synthetic data diversity.9:15–11:31 · Guest disagreement 0/10 Web Rephrasing and Programmatic Data Refinement Technical breakdown of web rephrasing techniques using small models and programmatic refinement from Nemotron CC and ProX.11:33–13:49 · Guest disagreement 0/10 LLM-Based Quality Classifiers for Large-Scale Filtering Detailed walkthrough of FineWeb-Edu and DCLM classifier-based quality filtering pipelines.13:51–16:52 · Guest disagreement 0/10 Post-Training Advancements: AgentInstruct, TÜLU 3, and SmolTalk Review of post-training datasets including AgentInstruct, TÜLU 3, SmolTalk, and multilingual data arbitrage.16:52–19:11 · Guest disagreement 0/10 The Emergence of High-Performing Smol and Edge Models Demonstration of competitive small models running on consumer hardware like iPhones via PocketPal.19:12–24:01 · Guest disagreement 1/10 Rethinking Scaling Laws: Inference Economics and Long Training Loubna challenges the 'bigger is always better' scaling paradigm, arguing for overtraining smaller architectures due to inference economics.24:03–26:12 · Guest disagreement 0/10 On-Device Deployment, Tooling, and Structured Generation Discussion of on-device advantages, privacy, and JSON-constrained structured text generation in browser.26:14–28:05 · Guest disagreement 1/10 Future Trends, The Return of Fine-Tuning, and Q&A Loubna concludes with a hot take predicting the industry will cycle back to fine-tuning specialized small models over prompt engineering.0:06–2:47 · The hosts pushing back 0/10 Opening Remarks and Talk Agenda Solo presentation opening outlining the evolution of synthetic data from post-training to pre-training and evaluation. Host scores remain zero as this is a monologue presentation.2:49–5:17 · The hosts pushing back 0/10 Demystifying Model Collapse and Web Data Pollution Loubna refutes conventional media and academic panic regarding model collapse and web pollution, presenting empirical data from FineWeb showing newer dumps perform better.5:18–9:13 · The hosts pushing back 0/10 Synthetic Pre-Training and the Cosmopedia Approach Loubna discusses reproducing Microsoft's Phi methodology with Cosmopedia, explaining prompt seeding strategies to ensure synthetic data diversity.9:15–11:31 · The hosts pushing back 0/10 Web Rephrasing and Programmatic Data Refinement Technical breakdown of web rephrasing techniques using small models and programmatic refinement from Nemotron CC and ProX.11:33–13:49 · The hosts pushing back 0/10 LLM-Based Quality Classifiers for Large-Scale Filtering Detailed walkthrough of FineWeb-Edu and DCLM classifier-based quality filtering pipelines.13:51–16:52 · The hosts pushing back 0/10 Post-Training Advancements: AgentInstruct, TÜLU 3, and SmolTalk Review of post-training datasets including AgentInstruct, TÜLU 3, SmolTalk, and multilingual data arbitrage.16:52–19:11 · The hosts pushing back 0/10 The Emergence of High-Performing Smol and Edge Models Demonstration of competitive small models running on consumer hardware like iPhones via PocketPal.19:12–24:01 · The hosts pushing back 0/10 Rethinking Scaling Laws: Inference Economics and Long Training Loubna challenges the 'bigger is always better' scaling paradigm, arguing for overtraining smaller architectures due to inference economics.24:03–26:12 · The hosts pushing back 0/10 On-Device Deployment, Tooling, and Structured Generation Discussion of on-device advantages, privacy, and JSON-constrained structured text generation in browser.26:14–28:05 · The hosts pushing back 0/10 Future Trends, The Return of Fine-Tuning, and Q&A Loubna concludes with a hot take predicting the industry will cycle back to fine-tuning specialized small models over prompt engineering.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 3:50 Refuting model collapse narratives

Loubna explicitly challenges prevalent media and Nature paper warnings about catastrophic model collapse, arguing empirical Common Crawl evidence shows synthetic-rich web data improves quality rather than degrades it.

Hardest push from the hosts ▶ 19:10 Challenging large model obsession

Loubna pushes back against the prevailing industry narrative that pursuing ever-larger models is optimal, emphasizing the steep recurring inference costs and showing that training smaller models on massive token counts yields superior trade-offs.

Biggest teaching moment ▶ 4:30 Explaining why small-scale collapse studies fail to generalize

Loubna breaks down why small-scale iterative self-training studies observe collapse whereas distilling knowledge from large, capable teacher models avoids quality degradation.

The host holds their own ▶ 27:00 Cyclical shift back to fine-tuning

Loubna synthesizes the trajectory of modern NLP, articulating that after swinging from fine-tuning to massive model prompt engineering, cost pressures are driving engineering practice back toward domain fine-tuning of small models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Opening Remarks and Talk Agenda 0000 Solo presentation opening outlining the evolution of synthetic data from post-training to pre-training and evaluation. Host scores remain zero as this is a monologue presentation.
Demystifying Model Collapse and Web Data Pollution 0020 Loubna refutes conventional media and academic panic regarding model collapse and web pollution, presenting empirical data from FineWeb showing newer dumps perform better.
Synthetic Pre-Training and the Cosmopedia Approach 0010 Loubna discusses reproducing Microsoft's Phi methodology with Cosmopedia, explaining prompt seeding strategies to ensure synthetic data diversity.
Web Rephrasing and Programmatic Data Refinement 0000 Technical breakdown of web rephrasing techniques using small models and programmatic refinement from Nemotron CC and ProX.
LLM-Based Quality Classifiers for Large-Scale Filtering 0000 Detailed walkthrough of FineWeb-Edu and DCLM classifier-based quality filtering pipelines.
Post-Training Advancements: AgentInstruct, TÜLU 3, and SmolTalk 0000 Review of post-training datasets including AgentInstruct, TÜLU 3, SmolTalk, and multilingual data arbitrage.
The Emergence of High-Performing Smol and Edge Models 0000 Demonstration of competitive small models running on consumer hardware like iPhones via PocketPal.
Rethinking Scaling Laws: Inference Economics and Long Training 0010 Loubna challenges the 'bigger is always better' scaling paradigm, arguing for overtraining smaller architectures due to inference economics.
On-Device Deployment, Tooling, and Structured Generation 0000 Discussion of on-device advantages, privacy, and JSON-constrained structured text generation in browser.
Future Trends, The Return of Fine-Tuning, and Q&A 0010 Loubna concludes with a hot take predicting the industry will cycle back to fine-tuning specialized small models over prompt engineering.

Statements from this episode (21)

Assertion Partly supported
Ben Allal: LLMs can be trained with entirely synthetic pipelines
“Today you can train an LLM with like an entirely synthetic pipeline. For example, you can use our Cosmopedia data sets and you can train a one B model on like a hundred and fifty billion tokens. Those are a hundred percent synthetic, and those are also of good…”
Loubna Ben Allal Dec 24, 2024 ▶ 1:35
Prediction Not checkable as stated
Ben Allal: Properly curated synthetic data prevents model collapse
“And I think there's a lot of concerns about model collapse, and I'm going to talk about that later, but we'll see that like, if we use synthetic data properly and we curate it carefully that shouldn't happen.”
Loubna Ben Allal Dec 24, 2024 ▶ 2:09
Assertion Supported
Hugging Face finds LLM proxy words jumped in Common Crawl after ChatGPT
“For example, here we measured like these words ratio in different dumps of common crawl, and we can see that like the ratio really increased after chat GPT's release.”
Loubna Ben Allal Dec 24, 2024 ▶ 3:50
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12
Opinion
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Loubna Ben Allal Dec 24, 2024 ▶ 4:35
Insight
Ben Allal: Distilling from large models avoids small-model self-training collapse
“I think if you do that approach, it's normal to observe this kind of behavior because the quality is going to be worse because the model is already small, and then if you train it just on these generations, you shouldn't expect it to become better. But what we…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:56
Insight
Ben Allal: Prompt diversity is essential for scaling synthetic training data
“The key ingredient to getting a good data set that is synthetic is trying as much as possible to keep it diverse, because if you just throw the same prompts as your model, like generate, like, a textbook about linear algebra, and even if you change the tempera…”
Loubna Ben Allal Dec 24, 2024 ▶ 6:08
Assertion Supported
Synthetic college text boosts MMLU; middle school text boosts OpenBookQA
“College textbooks are really good for MLU or middle school textbooks are good for benchmarks like open book, UA and Pico.”
Loubna Ben Allal Dec 24, 2024 ▶ 7:44
Assertion Supported
Ben Allal: NVIDIA generated 1.9 trillion synthetic tokens for Nemotron-CC
“This is a recent paper from NVIDIA, Mnemotron CC. They took things a bit further and they generated not a few billion tokens, but 1.9 trillion tokens, which is huge.”
Loubna Ben Allal Dec 24, 2024 ▶ 8:45
Insight
Ben Allal: Small models can generate synthetic data by rephrasing web pages
“The interesting thing in this approach is that you can use a model that is small Because it doesn't, rewriting doesn't require knowledge. It's just rewriting a page into a different style. So the model doesn't need to have like knowledge that is like extensive…”
Loubna Ben Allal Dec 24, 2024 ▶ 9:39
Assertion Supported
Ben Allal: Pre-training on rewritten C4 web data outperforms raw C4
“They rewrite some samples from C four into Q and A into Wikipedia, and they find that doing this works better than training just on C four.”
Loubna Ben Allal Dec 24, 2024 ▶ 9:59
Disclosure
Hugging Face Filtered 15T Token Dataset Down to 1.5T Educational Tokens
“And then we run this classifier on all of fine web, which is a 15 trillion tokens data set. And then we only keep the pages that have like a score that's higher than three. So for example, in our case, we went from 15 trillion tokens to just 1.5 trillion token…”
Loubna Ben Allal Dec 24, 2024 ▶ 12:21
Assertion Supported
Ben Allal: FineWeb-Edu Outperforms All Other Public Web Datasets
“And as you can see here FineWebEDU outperforms all the other public web datasets by a larger margin on a couple of benchmarks.”
Loubna Ben Allal Dec 24, 2024 ▶ 12:37
Insight
Ben Allal: Pooling multiple teacher models produces superior synthetic datasets
“Synthetic data, it doesn't have to come from a single model. And because we have so many good models now, you could like pull these models together and get like a dataset that's over really high quality and that's diverse and that's covers all your needs.”
Loubna Ben Allal Dec 24, 2024 ▶ 16:26
Assertion Supported
Llama 3.2 1B matches Llama 2 13B performance on LMSYS Arena
“For example, Lama 3.21 B it matches Lama two 13 B from that was the release last year on the LMSS arena, which is basically the default go to leaderboard for evaluating models using human evaluation.”
Loubna Ben Allal Dec 24, 2024 ▶ 17:04
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Loubna Ben Allal Dec 24, 2024 ▶ 20:33
Insight
Meta research: Sub-1B language models benefit more from depth than width
“For example, they find that depth is more important than width, so it's more important to have models that have, like, more layers than just making them more wide. They also find that GQA helps, that tie-in the embedding helps, so I think it's a nice study ove…”
Loubna Ben Allal Dec 24, 2024 ▶ 21:11
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Loubna Ben Allal Dec 24, 2024 ▶ 22:31
Insight
Ben Allal: Small models continue improving when trained on 11T tokens
“For example, smaller than one was trained only on one trillion tokens, but this model is trained on 11 trillion tokens. And we saw that the performance kept improving. The models didn't really plateau me training. Which I think is really interesting. It shows …”
Loubna Ben Allal Dec 24, 2024 ▶ 23:03
Insight
Ben Allal: Small models make more sense than large models for text extraction
“So I think text extraction is like one use case where small models can be really performant, and it makes sense to use them instead of just using larger models.”
Loubna Ben Allal Dec 24, 2024 ▶ 25:08
Prediction Not checkable as stated
Ben Allal: AI industry will shift to fine-tuning over prompt engineering
“And I think we're going back to fine tuning where we realize these models are really cosplay. It's better to use just a small model. We try to specialize it. So I think it's a little bit of a cycle and we're going to start to see like more of fine tuning and l…”
Loubna Ben Allal Dec 24, 2024 ▶ 27:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.