synthetic data

19 statements across 12 episodes · 11 bullish · 0 bearish · 12 people on the record · first statement Oct 20, 2023 by Jeremy Howard · across every show →

Everything said about synthetic data, oldest first

Oct 20, 2023 neutral
Assertion Supported
Howard: Microsoft's phi-1.5 lacks world knowledge due to synthetic training data
“Fi-one-point-five has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know it doesn't know who anybody is, he doesn't know about any movies, it doesn't really know anything about anything, like, because it was never, it's never rea…”
Jeremy Howard Oct 20, 2023 ▶ 1:07:21 The End of Finetuning — with Jeremy Howard of Fast.ai
Dec 17, 2023 bullish
Opinion
Liu: Synthetic data and task-specific fine-tuning provide alpha for code automation
“I feel like most models today, they still use, like, combination of, like, the stack and the pile as, like their training corpus but you can only stretch that so far. At some point, we need more data and I don't know. I think there's still more alpha in, like,…”
Beyang Liu Dec 17, 2023 ▶ 1:11:39 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Mar 6, 2024 neutral
Insight
Chintala: Synthetic data only works where humans already have symbolic models
“Outside of this, like, where we don't have good symbolic models, like, synthetic data obviously, like, doesn't make any sense. So synthetic data is not a magic wand where it'll work in all cases, in every case, you know, whatever. It's just where we as humans …”
Soumith Chintala Mar 6, 2024 ▶ 35:52 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Aug 22, 2024 positive
Opinion
Wang: Natural language-to-code translation is ripe for synthetic data generation
“I think that translation between natural language, English versus code and back and forth, I think is actually actually a really ripe source of synthetic data and Lama three specifically called out that, that they trained on that.”
Shawn Wang Aug 22, 2024 ▶ 48:12 Is finetuning GPT4o worth it?
Oct 13, 2024 neutral
Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Nov 25, 2024 neutral
Opinion
Pre-training on human data is hitting limits; synthetic data is required
“So I think on the data side, we're approaching the limit and the only data to increase that is synthetic generated data.”
Lin Qiao Nov 25, 2024 ▶ 35:05 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Dec 24, 2024 positive
Insight
Ben Allal: Pooling multiple teacher models produces superior synthetic datasets
“Synthetic data, it doesn't have to come from a single model. And because we have so many good models now, you could like pull these models together and get like a dataset that's over really high quality and that's diverse and that's covers all your needs.”
Loubna Ben Allal Dec 24, 2024 ▶ 16:26 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 positive
Assertion Supported
Ben Allal: Pre-training on rewritten C4 web data outperforms raw C4
“They rewrite some samples from C four into Q and A into Wikipedia, and they find that doing this works better than training just on C four.”
Loubna Ben Allal Dec 24, 2024 ▶ 9:59 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024
Assertion Supported
Synthetic college text boosts MMLU; middle school text boosts OpenBookQA
“College textbooks are really good for MLU or middle school textbooks are good for benchmarks like open book, UA and Pico.”
Loubna Ben Allal Dec 24, 2024 ▶ 7:44 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 positive
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 positive
Prediction Not checkable as stated
Ben Allal: Properly curated synthetic data prevents model collapse
“And I think there's a lot of concerns about model collapse, and I'm going to talk about that later, but we'll see that like, if we use synthetic data properly and we curate it carefully that shouldn't happen.”
Loubna Ben Allal Dec 24, 2024 ▶ 2:09 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024
Insight
Ben Allal: Prompt diversity is essential for scaling synthetic training data
“The key ingredient to getting a good data set that is synthetic is trying as much as possible to keep it diverse, because if you just throw the same prompts as your model, like generate, like, a textbook about linear algebra, and even if you change the tempera…”
Loubna Ben Allal Dec 24, 2024 ▶ 6:08 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 bullish
Opinion
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Loubna Ben Allal Dec 24, 2024 ▶ 4:35 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 positive
Insight
Ben Allal: Small models can generate synthetic data by rephrasing web pages
“The interesting thing in this approach is that you can use a model that is small Because it doesn't, rewriting doesn't require knowledge. It's just rewriting a page into a different style. So the model doesn't need to have like knowledge that is like extensive…”
Loubna Ben Allal Dec 24, 2024 ▶ 9:39 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Jul 31, 2025 bullish
Prediction Held up
Lambert: Labs will surely use parallel-compute models to generate synthetic data
“Well, I bet people, I mean, they surely will use these for synthetic data. It's just like the marginal gain on synthetic data is always very high.”
Nathan Lambert Jul 31, 2025 ▶ 50:22 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 29, 2025 positive
Insight
Morcos: Filtering synthetic data between generation cycles prevents model collapse
“If you filter the data after each point, that's now information injection, and that can break all of this and I think can prevent model collapse.”
Ari Morcos Aug 29, 2025 ▶ 43:47 Better Data is All You Need — Ari Morcos, Datology
Apr 2, 2026 bullish
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Fan-yun Sun Apr 2, 2026 ▶ 2:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Jun 1, 2026
Insight
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Ethan He Jun 1, 2026 ▶ 11:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 3, 2026 neutral
Insight
Hong: Accumulating synthetic AI data is not a moat, just buffer
“I think everyone is trying to accumulate like a data, which is not a mode. It's just time and time mode. It's all about like, you know, whether you can execute fast enough to make sure that you have like a certain buffer because of say your data set, you know,…”
Carina Hong Jun 3, 2026 ▶ 1:03:05 Scaling Past Informal AI - Carina Hong, Axiom Math
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.