Common Crawl

product on 11 shows · 7 statements across 6 episodes · said 62 times in 31 episodes since 2022

Latent Space 26 All-In 18 Lenny's Podcast 4 the a16z Podcast 4 No Priors 2 the MAD Podcast 2 20VC 2 American Optimist 1 the Knowledge Project 1 Sourcery 1 Big Technology 1

Mentions by year, every show

tap a year for its mentions
00158301520222023202420252026episodesmentions
081520222023202420252026episodes it came up in
001.57.531520222023202420252026episodesmentions per episode

Latent Space 26All-In 18Lenny's Podcast 4the a16z Podcast 420VC 2the MAD Podcast 2No Priors 2American Optimist 13 more shows

2026 3 mentions in 2 episodes 2 per episode
2025 27 mentions in 13 episodes 2 per episode
2024 15 mentions in 9 episodes 2 per episode
2023 16 mentions in 6 episodes 3 per episode
2022 1 mention in 1 episode

every mention on every show, scene by scene, with the transcript →

7 statements about Common Crawl, every show

LATENT SPACE Assertion Supported
Bachman: 90% of documents in web pre-training datasets are under 2k context
“If you take like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short. There's some long context documents, but the …”
Diego Bachman Sep 23, 2025 ▶ 30:30 ⚡️ Beyond Transformers with Power Retention
LENNY'S PODCAST Assertion Partly supported
Smith: There is more AI-generated content online than human-created content
“We found that there's more AI generated content on the internet than human generated content. So back to the common crawl study, we looked at a 100,000 different URLs over the past five years. And then you can see this curve where AI generated is now higher th…”
Ethan Smith Sep 14, 2025 ▶ 55:05 The ultimate guide to AEO: How to get ChatGPT to recommend your product | Ethan Smith (Graphite)
ALL-IN Prediction Open · timeframe Aug 2028
Palihapitiya: Grok 5 and Grok 6 will not use internet data
“Grok five and for sure grok six will not use common crawl. It will not use the internet.”
Chamath Palihapitiya Aug 1, 2025 ▶ 50:36 Trump AI Speech & Action Plan, DC Summit Recap, Hot GDP Print, Trade Deals, Altman Warns No Privacy
LATENT SPACE Assertion Supported
Hugging Face finds LLM proxy words jumped in Common Crawl after ChatGPT
“For example, here we measured like these words ratio in different dumps of common crawl, and we can see that like the ratio really increased after chat GPT's release.”
Loubna Ben Allal Dec 24, 2024 ▶ 3:50 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Soldani: Content owners blanket block crawling due to closed AI models
“What they found is, as a reaction to, like, the close like, of the existence of closed models, like OpenAI or Cloud GPT or Cloud a lot of content owners have blanket blocked any type of crawling to their website.”
Luca Soldani Dec 23, 2024 ▶ 18:34 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
ALL-IN Assertion Not checkable as stated
Friedberg: YouTube's data repo is 300x larger than Common Crawl
“In YouTube growing by one to two petabytes per day, which makes YouTube's data repository 300 times larger than common crawl, which makes it bigger than anything else that anyone else has.”
David Friedberg Feb 9, 2024 ▶ 1:05:58 E165: Vision Pro: use or lose? Meta vs Snap, SaaS recovery, AI investing, rolling real estate crisis

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.