Common Crawl, every mention
17 scenes · ← back to Common Crawl
tap a year for its mentions
every year anyone Michael Royzen 5Shawn Wang 3Simon Eskildsen 2Pim de Witte 2Luca Soldani 2Loubna Ben Allal 2Ari Morcos 2Jeremy Howard 1Elie Bakouch 1
Verbatim, from the transcripts: the passages where Common Crawl comes up
Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
- ▶ 51:46 Simon Eskildsen Like someone searching for a very long text string on all of common crawl. 2 times in the scene
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
- ▶ 36:00 unnamed speaker But I would say that percentage on Common Crawl that has COT tokens in there went from zero to 0.001%.
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
- ▶ 25:35 Pim de Witte What if we predict action tokens on essentially what is the equivalent of the common crawl data set, uh, but for interactivity?
- ▶ 28:04 Pim de Witte We essentially have sort of the internet or like common crawl, if you will, and every single lab is trying to simulate that, right, in order to get similar data, in order to train their agents.
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 3:38 Elie Bakouch Yeah, I think this is the, so like the way it works is that, yeah, so sometimes like I think the PDF from Common Crawl are not really well extracted, basically.
⚡️ Beyond Transformers with Power Retention
- ▶ 30:30 unnamed speaker If you take, uh, like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 17:29 Ari Morcos Really wonderful effort to kind of curate, uh, common crawl style datasets.
- ▶ 23:01 Shawn Wang And I think everyone starts a common crawl. 3 times in the scene
- ▶ 1:08:46 Ari Morcos Um, like we didn't need to go to common crawl to get those tokens.
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 53:33 unnamed speaker Somehow C-Four, the, the common, the colossal clean corpus, did much better than, ah, common crawl, even though it filtered out most of this, like, it was very prudish.
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 3:50 Loubna Ben Allal For example, here we measured like these words ratio in different dumps of common crawl, and we can see that like the ratio really increased after chat GPT's release.
- ▶ 10:14 Loubna Ben Allal Uh, they rewrite some pages from common crawl, uh, for two reasons.
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 18:10 Luca Soldani So what they did is they, um, went through every snapshot of Common Crawl. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 31:47 Jeremy Howard He, like, went back to Common Crawl and did everything.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 37:51 unnamed speaker Uh, we've got, these are things that we've seen before, Wikipedia datasets, C-Four dataset, Common Crawl, uh, which is used for your, I would say, more general purpose models.
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 8:47 unnamed speaker Instead now, anything we're going to get on common crawl updates and things like that, you never know. 2 times in the scene
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 10:33 Michael Royzen So in the demo, in the notebook, uh, it was, um, there were instructions for how to make an elastic search index just for Wikipedia, and I was like, why not do all of Common Crawl? 5 times in the scene