Wikipedia, every mention
47 scenes · ← back to Wikipedia
tap a year for its mentions
every year anyone Shawn Wang 12Jeremy Howard 6Jack Morris 5William Beauchamp 4Usama Shafqat 4Loubna Ben Allal 4Tri Dao 3Michael Royzen 2Heather Kulik 2Harrison Chase 2
Verbatim, from the transcripts: the passages where Wikipedia comes up
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
- ▶ 5:56 Joon Sung Park It's the, it's social media, Wikipedia, all these kind of data.
- ▶ 1:03:42 Shawn Wang Uh, it's very, very famous, like, you know, to the point of having a Wikipedia page about this kind of like really in-depth understanding and interview of people as they, about their, about their lives, which seems mundane, but is told in…
The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
- ▶ 22:39 Dan Biderman Um, so the examples we like to give is that, uh, if you take a Lama, a 70 B model, and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this, uh,
Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
- ▶ 10:09 Danielle Perszyk We, we depend upon using tools to be able to do daily things, and we can look up on Wikipedia or AI tools that are just organizing information in new ways.
- ▶ 31:39 unnamed speaker I don't know, like, you know, I think there's like a trend of like Wikipedia's wikis, LLM wikis that like, you know, encode some kind of knowledge that can be passed between agents.
The Blueprint for Autonomous Work Agents | Gavriel Cohen, NanoClaw
- ▶ 5:59 Gavriel Cohen Is the second brain use case, where you're just kind of dumping in information, and you're not expecting it to give you ready-made output, but it's just collecting that information, building up its internal memory, or knowledge graph, or…
The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
- ▶ 46:18 Reynold Xin Uh, by the way, you can search on Wikipedia.
🔬There Is No AlphaFold for Materials — AI for Materials Discovery with Heather Kulik
- ▶ 0:14 Heather Kulik CHAT GPT is super good at Wikipedia level chemistry knowledge.
- ▶ 13:40 Heather Kulik My personal experience is that, um, and this will date itself immediately, is, is that, is that, uh, ChatGPT is super good at Wikipedia-level chemistry knowledge, but one of my favorite things to actually throw at GPT, as, as an anecdote,…
🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
- ▶ 47:05 Andrew White I think, you know, the first set of answers in twenty-twenty-three, I think, was basically no, is that, like, you know, you can go find the synthesis route for many dangerous compounds on Wikipedia. 2 times in the scene
Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
- ▶ 5:31 Steve Yegge Who are much better engineers than I am, ok, I mean, world class, maybe some of the best in the whole world, ok, have built technologies that you've heard of, and they're not using AI yet, except the occasional, I'll ask Cursor a chat…
⚡️ Building the AI Hardware Engineer with Matthias Wagner, Co-founder of Flux
- ▶ 21:37 unnamed speaker Deep wiki is just kind of like a Wikipedia entry on it.
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 31:01 Shawn Wang A Docker container that has like a clone of Reddit, a clone of Wikipedia, a clone of GitLab, a clone of CMS, and a clone of an e-commerce place.
The antidote to AI fatigue — Answer.ai Solveit
Better Data is All You Need — Ari Morcos, Datology
- ▶ 47:23 Shawn Wang And I feel like that's a little bit of cargo culting of like, oh, just cause you like write like Wikipedia or write like textbooks, the models learn better.
Information Theory for Language Models: Jack Morris
- ▶ 20:37 Shawn Wang You can compare that to Wikipedia. 3 times in the scene
- ▶ 22:33 Jack Morris Actually, let's go to your Wikipedia, uh, numbers if, if you still have access to that. 4 times in the scene
- ▶ 34:34 Jack Morris And they're probably all trained on like MS Marco, which is a really popular dataset and pre-trained maybe on Wikipedia.
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 31:25 Will Brown And so for, for these experiments we were doing, the trick was like, okay, does the, like some string matching thing involving the ground truth answer and the return search results from Wikipedia?
Why is everyone cloning Deep Research?
- ▶ 3:06 Arush Sehgal No Wikipedia page already about this topic or something like that, right?
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 8:27 William Beauchamp But I would consider Wikipedia to be a platform where instead of the Britannica encyclopedia, which is this, it's like a monolithic, you get all the, the researchers together, you get all the data together, and you combine it in this, in… 2 times in the scene
- ▶ 53:11 William Beauchamp Now, with Wikipedia, Wikipedia crowdsources knowledge. 2 times in the scene
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 4:42 Loubna Ben Allal Um, and the issue with, like, uh, model collapse is that, for example, those studies, they were done at, like, a small scale, and you would ask the model to complete, for example, a Wikipedia paragraph, and then you would train it on these…
- ▶ 9:32 Loubna Ben Allal For example, they ask an LLM to rewrite the sample into a Wikipedia passage or into a Q&A page. 3 times in the scene
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 11:41 unnamed speaker One was a, all of Wikipedia, the English version, which was like eight hundred million words, and then I think a set of books, which was 2.5 billion words. 2 times in the scene
Agents @ Work: Lindy.ai (with live demo!)
- ▶ 22:01 Florent Crivello You go through Google to search Wikipedia.
How NotebookLM Was Made
- ▶ 29:29 Usama Shafqat I think another cool one is just like any Wikipedia article. 2 times in the scene
- ▶ 40:37 Usama Shafqat Like, maybe you are going from, let's say, the Wikipedia article of, like, one of the History of Mysteries, maybe, episodes. 2 times in the scene
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 3:23 Harrison Chase And, and I think in the paper, you mostly deal with Wikipedia and I think there's some other datasets as well, but the outside world is the outside world. 2 times in the scene
This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
- ▶ 37:39 Rob Haisfield Like, one game that I like to play with WebSim, sometimes with co-op, is like, I'll open a page, so like, one of the first ones that I did was I tried to go to Wikipedia in a universe where octopus were sapient and not humans, right?
- ▶ 1:22:45 unnamed speaker Why Liquid AI is challenging the Perceptron, and why you should not donate to Wikipedia.
- ▶ 1:52:44 Shawn Wang Um, okay, and then, and then the last thing, which I, which I didn't know, a lot of LLMs rely on Wikipedia for data. 6 times in the scene
Breaking down the OG GPT Paper by Alec Radford
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 37:51 unnamed speaker Uh, we've got, these are things that we've seen before, Wikipedia datasets, C-Four dataset, Common Crawl, uh, which is used for your, I would say, more general purpose models.
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
- ▶ 41:32 Erik Bernhardsson So, so let's say, you know, you have like a, actually we're like writing a blog post, like we, where we, we take all of Wikipedia and like, uh, parallelize embeddings in 15 minutes and, and, and produce vectors for each article.
Building an open AI company - with Ce and Vipul of Together AI
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 10:11 unnamed speaker Um, you have, uh, you, you created this like Wikipedia, like, 3 times in the scene
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 8:15 Michael Royzen Um, the demo itself, it used, I think, BART as the model, and in the notebook, it had support for both, um, an elastic search, um, index of Wikipedia, as well as a dense index, um, powered by Facebook's face. 2 times in the scene
Powering your Copilot for Data - with Artem Keydunov from Cube.dev
- ▶ 8:55 Artem Keydunov Uh, I'll, I'll try, you know, like it's really like a lot of like, uh, Wikipedia pages and like a lot of like a blog post trying to go into academics of it.
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 11:40 Jeremy Howard You know, with that background and this kind of like interest in transfer learning, you know, I'd been thinking about this thing for kind of 30 years, and I thought like, oh, I wonder if we're there yet, you know, because we have a lot of… 3 times in the scene
- ▶ 16:10 Jeremy Howard And similar with, with Stephen, you know, I asked Stephen Merity, like, why don't we just find, you know, take your AWD, ASTLEM, and like, train it on all of Wikipedia, and fine-tune it, and he's kind of like, I don't think that's gonna…
- ▶ 35:55 Jeremy Howard Whittaker, um, basically had been, um, playing around with this fun Kaggle competition, which is actually still running as we speak, which is, um, can you create a model which can answer multiple choice questions about anything that's in…
- ▶ 1:07:21 Jeremy Howard Um, the thing, ok, so, like, Fi-one-point-five, uh, has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know, um, it doesn't know who anybody is, he doesn't know about any movies, it doesn't, doesn't really…
- ▶ 1:20:56 unnamed speaker How much is Wikipedia?
FlashAttention-2: Making Transformers 800% faster AND exact
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 53:33 George Hotz right and there's a question of how long this thesis is going to continue for it's a cool thesis and look i think um i would be lying along with everybody else i was into language models like way back in the day for the hotter price i got…