arXiv, every mention
40 scenes · ← back to arXiv
tap a year for its mentions
every year anyone Ben Firshman 11Shawn Wang 6Eugene Yan 6Alex Lupsasca 6Alessio Fanelli 4Yi Tay 3Sander Schulhoff 2Nathan Lambert 2Will Brown 1Tri Dao 1
Verbatim, from the transcripts: the passages where arXiv comes up
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
- ▶ 3:31 Shawn Wang Archive gives you something, right?
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- ▶ 1:02:28 Evan Feinberg You know, back in the day when AI could still be in peer reviewed journals, not just sort of like random white papers on, on archive.
🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
- ▶ 1:11:09 unnamed speaker Now we have actually seen this, and you can go read the publication, it's on archive, where the public data set that we used actually is showing improvements, like five to 16%, I believe, on general scientific reasoning. 2 times in the scene
🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
- ▶ 14:42 Alex Lupsasca So we put this on the archive a little over a month ago now.
- ▶ 51:29 Alex Lupsasca Finally, it just, we, we say, okay, write up the paper and you can see the paper that it writes out and it's very close to the final thing that we actually put on the archive.
- ▶ 1:13:01 Alex Lupsasca I liked it very much, and it came out in June on the archive, and in August GPT-V came out, and the cutoff date for its training set
- ▶ 1:20:24 Alex Lupsasca They submit that to archive, and this is a problem that the academic community is dealing, is trying to come to grips with now, which is this problem of AI sloppy for science. 3 times in the scene
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- ▶ 1:17 unnamed speaker And it was basically like, uh, like really jank, like view comment button next to like every paragraph of like an archive paper. 4 times in the scene
- ▶ 4:18 unnamed speaker So if you're trying to discover papers on one end of the spectrum is, like, the archives sort by new.
- ▶ 7:03 unnamed speaker I think like, like kind of the analogy of like archive existed and we came in and like built like an intelligent layer of our archive, you know,
- ▶ 25:23 unnamed speaker I guess when it, when it comes to novelty, it's concerned, like, Archive has had this problem, like, even early on, where people would use it for, like, I forget the exact term, but people would post a paper, even when it's not fully… 3 times in the scene
- ▶ 30:23 unnamed speaker I mean, yeah, this, just going back onto sort of the whole, like, the whole paper discussion, I think, like, you just have to look at a graph of, like, archive submissions in the last, like, 10 years. 3 times in the scene
[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
- ▶ 11:36 Anastasios Angelopoulos So you can go look at the first version of the paper on archive yourself, and you'll see the claims.
⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs
- ▶ 16:21 Jessica Rumbelow I'm actually really worried about this because I think we're already seeing archive and other online repositories and-
Better Data is All You Need — Ari Morcos, Datology
- ▶ 24:17 Shawn Wang Uh, archive is, which is, you know, GitHub for papers, books.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 35:26 Nathan Lambert So I, I told you this at lunch, which is like, um, deep research, but only archive papers. 2 times in the scene
⚡️The Future of Notebooks - with Akshay Agrawal of Marimo
- ▶ 14:02 Akshay Agrawal So this demo here, we have like a bunch of papers from archive and we've indexed them using chroma DB into like a basic vector database.
Information Theory for Language Models: Jack Morris
- ▶ 31:38 Jack Morris And then we recently had a paper come out on archive, which will hopefully be published at some point.
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
- ▶ 1:05:10 Spooks (Swyx) Whatever hits archive, literally you do the same as the rest of us.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 29:46 Chris Lattner To me, what I look at is not just the things that people have done and put into VLM, for example, but the continuous stream of archive papers, right?
- ▶ 1:13:25 Alessio Fanelli You mentioned some of the research on inference and the archive papers. 3 times in the scene
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 1:46:55 Emmanuel Ameisen If somebody's going to be able to read this, like, if we give them an archive paper with a bunch of equation and some, like, random plot, they'd be like, that's not for me, but they see this and they're like, hey, like, this is really…
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 34:07 Will Brown So yeah, so like a lot of models naturally will like think in LaTeX because they've been trained on a lot of archive-like tech.
Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
- ▶ 30:57 Alessio Fanelli Uh, uh, apparently the paper is going live in three hours on archive.
Building Manus AI (first ever Manus Meetup)
- ▶ 16:25 unnamed speaker And also, uh, with the AI industry, we read papers every day on AXE, so you can just, on AXE, you can just summarize the, uh, uh, or expand this paper just inside the AXE.
[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
- ▶ 0:57 Eugene Yan Right, we just say, hey, you know, LM, you know, given this archive, here are these 10 archive papers, find me everything that's relevant to hallucinations. 2 times in the scene
- ▶ 8:08 Eugene Yan So for example, imagine you are summarizing an archive paper.
- ▶ 14:55 Eugene Yan So let's, again, take our example of this archive PDF. 2 times in the scene
- ▶ 44:50 Eugene Yan If you search it up on Google, I'm sure you will leave it to the archive.
- ▶ 50:27 Shreya Shankar Yeah, we, um, we do measure precision in the updated eval that is not on Archive.
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 6:17 Sander Schulhoff And then we put it on archive and the response was amazing. 2 times in the scene
- ▶ 8:15 Shawn Wang It's like been inflated into like a 10 page PDF that's posted on Archive, and then you've done the reverse of compressing it into like one paragraph each of each paper. 4 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 1:03:35 Mark Huang Like, people have already have it, dropped it on archive, or they're just openly talking about it.
Breaking down the OG GPT Paper by Alec Radford
- ▶ 2:46 unnamed speaker Because basically you can utilize Wikipedia, which is very big, or even all the papers published on Archive, and so on.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 9:01 Ben Firshman So maths, physics, computer science, machine learning, notably, is all published on the archive, which is, um, actually a surprisingly old institution. 9 times in the scene
- ▶ 13:13 Ben Firshman One thing that's really interesting about these machine learning papers is that these machine learning papers are published on, on the archive, and a lot of them are actual fundamental research, so, like, should be, like, prose describing…
- ▶ 27:52 Ben Firshman So we were, we were creating an accompanying material for their models, basically.
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 1:03:48 Jeremy Howard You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as the paragraphs, which obviously is not going…