Common Crawl, every mention

17 scenes · ← back to Common Crawl

tap a year for its mentions
00531052023202420252026episodesmentions
0352023202420252026episodes it came up in
002.52.5552023202420252026episodesmentions per episode

every year anyone Michael Royzen 5Shawn Wang 3Simon Eskildsen 2Pim de Witte 2Luca Soldani 2Loubna Ben Allal 2Ari Morcos 2Jeremy Howard 1Elie Bakouch 1

Verbatim, from the transcripts: the passages where Common Crawl comes up

loading…

Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer Mar 12, 2026 · 2 mentions

  • ▶ 51:46 Simon Eskildsen Like someone searching for a very long text string on all of common crawl. 2 times in the scene

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay Jan 23, 2026 · 1 mention

  • ▶ 36:00 unnamed speaker But I would say that percentage on Common Crawl that has COT tokens in there went from zero to 0.001%.

World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI Dec 6, 2025 · 2 mentions

  • ▶ 25:35 Pim de Witte What if we predict action tokens on essentially what is the equivalent of the common crawl data set, uh, but for interactivity?
  • ▶ 28:04 Pim de Witte We essentially have sort of the internet or like common crawl, if you will, and every single lab is trying to simulate that, right, in order to get similar data, in order to train their agents.

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 1 mention

  • ▶ 3:38 Elie Bakouch Yeah, I think this is the, so like the way it works is that, yeah, so sometimes like I think the PDF from Common Crawl are not really well extracted, basically.

⚡️ Beyond Transformers with Power Retention Sep 23, 2025 · 1 mention

  • ▶ 30:30 unnamed speaker If you take, uh, like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short.

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 5 mentions

The Utility of Interpretability — Emmanuel Amiesen Jun 6, 2025 · 1 mention

  • ▶ 53:33 unnamed speaker Somehow C-Four, the, the common, the colossal clean corpus, did much better than, ah, common crawl, even though it filtered out most of this, like, it was very prudish.

Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024] Dec 24, 2024 · 2 mentions

  • ▶ 3:50 Loubna Ben Allal For example, here we measured like these words ratio in different dumps of common crawl, and we can see that like the ratio really increased after chat GPT's release.
  • ▶ 10:14 Loubna Ben Allal Uh, they rewrite some pages from common crawl, uh, for two reasons.

Best of 2024: Open Models [LS LIVE! at NeurIPS 2024] Dec 23, 2024 · 2 mentions

  • ▶ 18:10 Luca Soldani So what they did is they, um, went through every snapshot of Common Crawl. 2 times in the scene

Answer.ai & AI Magic with Jeremy Howard Aug 17, 2024 · 1 mention

A Comprehensive Overview of Large Language Models - Latent Space Paper Club Mar 15, 2024 · 1 mention

  • ▶ 37:51 unnamed speaker Uh, we've got, these are things that we've seen before, Wikipedia datasets, C-Four dataset, Common Crawl, uh, which is used for your, I would say, more general purpose models.

The Four Wars of the AI Stack - Dec 2023 Recap Jan 26, 2024 · 2 mentions

  • ▶ 8:47 unnamed speaker Instead now, anything we're going to get on common crawl updates and things like that, you never know. 2 times in the scene

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind Nov 3, 2023 · 5 mentions

  • ▶ 10:33 Michael Royzen So in the demo, in the notebook, uh, it was, um, there were instructions for how to make an elastic search index just for Wikipedia, and I was like, why not do all of Common Crawl? 5 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.