Common Crawl

4 statements across 3 episodes · 1 bullish · 1 bearish · 3 people on the record · first statement Dec 23, 2024 by Luca Soldani · said 26 times in 13 episodes since 2023 · across every show →

Mentions by year

brought up most by Michael Royzen (5), Shawn Wang (3), Simon Eskildsen (2), Pim de Witte (2), Luca Soldani (2), Loubna Ben Allal (2), Ari Morcos (2), Jeremy Howard (1)

tap a year for its mentions
00531052023202420252026episodesmentions
0352023202420252026episodes it came up in
002.52.5552023202420252026episodesmentions per episode
2026 3 mentions in 2 episodes 2 per episode
2025 10 mentions in 5 episodes 2 per episode
2024 8 mentions in 5 episodes 2 per episode
2023 5 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Common Crawl, oldest first

Dec 23, 2024 negative
Assertion Supported
Soldani: Content owners blanket block crawling due to closed AI models
“What they found is, as a reaction to, like, the close like, of the existence of closed models, like OpenAI or Cloud GPT or Cloud a lot of content owners have blanket blocked any type of crawling to their website.”
Luca Soldani Dec 23, 2024 ▶ 18:34 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Dec 24, 2024 positive
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Dec 24, 2024 neutral
Assertion Supported
Hugging Face finds LLM proxy words jumped in Common Crawl after ChatGPT
“For example, here we measured like these words ratio in different dumps of common crawl, and we can see that like the ratio really increased after chat GPT's release.”
Loubna Ben Allal Dec 24, 2024 ▶ 3:50 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Sep 23, 2025 neutral
Assertion Supported
Bachman: 90% of documents in web pre-training datasets are under 2k context
“If you take like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short. There's some long context documents, but the …”
Diego Bachman Sep 23, 2025 ▶ 30:30 ⚡️ Beyond Transformers with Power Retention
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.