FlashAttention, every mention
32 scenes · ← back to FlashAttention
tap a year for its mentions
every year anyone Diego Bachman 9Quentin Anthony 6Tri Dao 5Chris Lattner 3Alessio Fanelli 3Shawn Wang 2Jeremy Howard 2Andrej Karpathy 2Stefano Ermon 1Mark Huang 1
Verbatim, from the transcripts: the passages where FlashAttention comes up
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 2:35 Quentin Anthony So I ported a flash attention over to those very painful.
- ▶ 6:32 Quentin Anthony So when I want to make an optimization to say flash attention, I don't go from the top down. 3 times in the scene
- ▶ 11:39 Quentin Anthony So I don't want to reinvent the wheel when I'm writing like a flash attention kernel.
- ▶ 16:04 Quentin Anthony I have some sort of like new model architecture that changes the tension and I can just fuse an existing flash attention kernel onto some epilogue.
⚡️ Beyond Transformers with Power Retention
- ▶ 7:26 Diego Bachman And just to give you some sense of how much at training time, the sort of speed up that we get purely from the flop reductions and efficient hardware level implementation of flash of, of, of, of power retention over flash attention at a…
- ▶ 9:19 Diego Bachman It's actually, in many ways, significantly more involved than flash attention or a lot of other kernels that people, uh, like to write. 6 times in the scene
- ▶ 15:59 Diego Bachman And what happens when you switch it from flash attention to power retention?
- ▶ 27:20 Diego Bachman There's things like flash attention from TreeDAO, which produced these massive, massive improvements in performance.
⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
- ▶ 22:37 Stefano Ermon I wasn't the flash attention paper.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 20:44 Chris Lattner Well, that's a very fancy compiler technology that allows you to say, okay, you will write one version of, uh, flash attention and then cool.
- ▶ 30:51 Chris Lattner And so you can go look at how we brought up H-one hundred, built flash attention from scratch in a few weeks, built like all the stuff. 2 times in the scene
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
[Paper Club] Upcycling Large Language Models into Mixture of Experts
- ▶ 14:08 unnamed speaker I'm just curious, like, is, are these available to, for everyone to use in the way that Fletcher Tension is, or are these specific to the Megatron implementation?
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 6:19 Alessio Fanelli We had Trida on the podcast that you mentioned, uh, he was really inspired working with like systems people to think about flash attention.
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 14:20 Andrej Karpathy Step two, we don't want to write our own flash retention. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 42:17 Jeremy Howard And even as we did so, new regressions were appearing in, like, Transformers and stuff, that Benjamin then had to go away and figure out, like, oh, how come flash attention doesn't work in this version of Transformers anymore with this set… 2 times in the scene
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 24:54 Mark Huang So we combined flash attention and ring attention together, uh, with like a really specific, uh, network topology on our GPUs to be able to maximize the, the memory bandwidth.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 25:55 unnamed speaker Um, other kinds of tricks that we are using, uh, things like flash attention, where, um, it's a very smart way of utilizing, uh, memory.
Building an open AI company - with Ce and Vipul of Together AI
- ▶ 6:33 Ce Zhang When you do, for example, flash attention, it's about data movement at a different essentially memory hierarchy, right?
- ▶ 35:07 Alessio Fanelli Because we at ThreeDAO, obviously, flash attention to is open source. 2 times in the scene
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 13:46 unnamed speaker Because when we had three DAO on the pockets, it's a flash of tension came to be because at AZ, they have so much overlap between systems engineering, like, uh, deep learning engineers.
- ▶ 1:16:21 unnamed speaker Uh, I always say, like, it does remind me of Flesh Attention a little bit, in the sense that, like, it's, like, kind of an equivalent, uh, thing to the thing it's replacing, and it's just faster, cheaper.
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 1:15:32 unnamed speaker As we had, um, three DAO on the podcast who created Flash Attention. 4 times in the scene
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 20:42 unnamed speaker We also did an episode on Flash Attention, which helps to make part of it sub-linear at least.
FlashAttention-2: Making Transformers 800% faster AND exact
- ▶ 0:18 unnamed speaker Um, you might not remember his name, but he's one of the main authors in the flash attention paper, which is one of the seminal work in, uh, the transformers era.
- ▶ 1:39 unnamed speaker Um, so flesh attention is definitely 4 times in the scene
- ▶ 10:09 Tri Dao Yeah, I think mostly on the, kind of on the practical side is that, um, when you do kernel fusion is a little bit, uh, you lose a little bit of flexibility in the sense that, hey, now you have, um, for example, is, uh, flash attention is…
- ▶ 15:16 unnamed speaker Um, how do you think about the future of like flash attention? 2 times in the scene
- ▶ 17:58 unnamed speaker Um, did you think flash attention would be this popular?
- ▶ 32:23 Tri Dao Um, by the end, and like I just ended up working with the code a whole lot more, and I realized that, hey, there are these inefficiencies still in flash attention. 3 times in the scene
- ▶ 47:30 Tri Dao So for a really long sequences, um, when you train with transformer, you know, with flash retention and so on, it's still, you know, the computation is still quadratic in the sequence length.
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 51:34 Shawn Wang Because you're finding algorithms, like flash attention. 2 times in the scene