FlashAttention

also referred to as: flash attention

7 statements across 4 episodes · 5 bullish · 1 bearish · 4 people on the record · first statement Aug 3, 2023 by Tri Dao · said 53 times in 17 episodes since 2023 · across every show →

Mentions by year

brought up most by Diego Bachman (9), Quentin Anthony (6), Tri Dao (5), Chris Lattner (3), Alessio Fanelli (3), Shawn Wang (2), Jeremy Howard (2), Andrej Karpathy (2)

tap a year for its mentions
001052010202320242025episodesmentions
0510202320242025episodes it came up in
002.55510202320242025episodesmentions per episode
2025 19 mentions in 4 episodes 5 per episode
2024 14 mentions in 9 episodes 2 per episode
2023 20 mentions in 4 episodes 5 per episode

every mention, scene by scene, with the transcript →

Everything said about FlashAttention, oldest first

Aug 3, 2023 positive
Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 positive
Insight
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 negative
Insight
Kernel fusion sacrifices flexibility for researchers experimenting with attention
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exa…”
Tri Dao Aug 3, 2023 ▶ 10:09 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 bullish
Prediction Not checkable as stated
FlashAttention techniques generalize across accelerators with asymmetric memory
“I expect the idea to be broadly these ideas to be broadly applicable to different hardware. As long as, I think the main idea is you have, like, asymmetry in, in memory hierarchy, which tends to be everywhere, you know, in, in a lot of a lot of accelerators.”
Tri Dao Aug 3, 2023 ▶ 34:11 FlashAttention-2: Making Transformers 800% faster AND exact
Oct 20, 2023 bullish
Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai
Feb 8, 2024
Insight
Zhang: ML systems optimization is fundamentally always about data movement
“Fundamentally, the thing we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like computing, it's about communication across different machines. When y…”
Ce Zhang Feb 8, 2024 ▶ 6:19 Building an open AI company - with Ce and Vipul of Together AI
Jun 13, 2025 positive
Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.