The Stack, every mention
5 scenes · ← back to The Stack
tap a year for its mentions
every year anyone Varun Mohan 2Diego Bachman 1Beyang Liu 1
Verbatim, from the transcripts: the passages where The Stack comes up
⚡️ Beyond Transformers with Power Retention
- ▶ 15:51 Diego Bachman And this is the baseline level of loss that it achieves on like a very standard, they call it the stack data set and specialized to Python.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 13:15 Varun Mohan I think maybe after we started, there was the stock that was also came to be, but for us, we had a model for ourselves even before that.
- ▶ 32:58 Varun Mohan And I think one of the cool things that the stack showed actually was they did a, like a, I think they did some ablation studies where they were like, Hey, what happens if we do, if we do decontamination of our data, what happens if we do…
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
- ▶ 1:11:39 Beyang Liu It's just, like, I feel like most models today, they still use, like, combination of, like, the stack and the pile, uh, as, like, uh, their, their training corpus, um, but you can only stretch that so far.
FlashAttention-2: Making Transformers 800% faster AND exact
- ▶ 59:02 unnamed speaker And like the PTA models are great, but like the pile and like the stack, like, you know, everybody uses them, you know?