FlashAttention-2

2 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Aug 3, 2023 by Tri Dao · said 9 times in 5 episodes since 2023 · across every show →

Mentions by year

brought up most by Quentin Anthony (2), Alessio Fanelli (1)

tap a year for its mentions
002142202320242025episodesmentions
012202320242025episodes it came up in
002142202320242025episodesmentions per episode
2025 3 mentions in 2 episodes 2 per episode
2024 2 mentions in 2 episodes 1 per episode
2023 4 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about FlashAttention-2, oldest first

Aug 3, 2023 positive
Assertion Supported
FlashAttention-2 is twice as fast as FlashAttention-1
“We managed to make it to X faster. And now it's pretty close to probably the efficiency of things like matrix multiply, which probably this, the most optimized subroutine on the planet.”
Tri Dao Aug 3, 2023 ▶ 33:01 FlashAttention-2: Making Transformers 800% faster AND exact
Nov 3, 2025 bullish
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Quentin Anthony Nov 3, 2025 ▶ 3:19 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.