C4, every mention
6 scenes across 2 shows · ← back to C4
tap a year for its mentions
Latent Space 6
Sourcery 1
every year every show
Latent Space 6
Sourcery 1
Verbatim, from the transcripts: passages where C4 comes up on Latent Space, Sourcery
Inside The $2.2B AI Research Accelerator | Turing · Sourcery with Molly O'Shea
- ▶ 18:47 Jonathan Siddharth Uh, there are these data sets like Common Crawl, C-for, GitHub, Archive.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 1:11:00 Ari Morcos The first thing I would just say is, like, if you are one of these people that keeps on finding yourself, just, like, staring at the data, you keep on going into the data set, if you can tell me what your, you know, favorite and, and least…
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 53:33 unnamed speaker Somehow C-Four, the, the common, the colossal clean corpus, did much better than, ah, common crawl, even though it filtered out most of this, like, it was very prudish.
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 9:15 Loubna Ben Allal Um, so for this rephrasing the web, uh, this approach was, uh, suggested in this paper by Pratyush, where basically in this paper, they take, uh, some samples from C-IVOR datasets, and then they use an LLM to rewrite these samples into a… 2 times in the scene
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
- ▶ 39:38 unnamed speaker He did this on Mistral-Seven-B, um, and OpenWebMap in C-IV.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 37:51 unnamed speaker Uh, we've got, these are things that we've seen before, Wikipedia datasets, C-Four dataset, Common Crawl, uh, which is used for your, I would say, more general purpose models.