FlashAttention

product on 6 shows · 8 statements across 5 episodes · said 63 times in 24 episodes since 2023

Latent Space 53 No Priors 3 the MAD Podcast 3 Invest Like the Best 2 the Y Combinator Startup Podcast 1 Sourcery 1

Mentions by year, every show

tap a year for its mentions
0013525102023202420252026episodesmentions
05102023202420252026episodes it came up in
002.555102023202420252026episodesmentions per episode

Latent Space 53the MAD Podcast 3No Priors 3Invest Like the Best 2Sourcery 1the Y Combinator Startup Podcast 1

2026 6 mentions in 4 episodes 2 per episode
2025 19 mentions in 4 episodes 5 per episode
2024 15 mentions in 10 episodes 2 per episode
2023 23 mentions in 6 episodes 4 per episode

every mention on every show, scene by scene, with the transcript →

8 statements about FlashAttention, every show

LATENT SPACE Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)
SOURCERY Assertion Not checkable as stated
Isford: Together AI Has the Market's Fastest Inference Tech
“They have the author of Flash Attention, Sri Dao, who is brilliant, as their chief scientist. He is, has a proprietary version of the fastest, like, inference technology on the market that's low cost to serve.”
Grace Isford Aug 23, 2024 ▶ 34:33 Computer Science is 'So Hot Right Now' | Grace Isford, Lux Capital · Sourcery with Molly O'Shea
Zhang: ML systems optimization is fundamentally always about data movement
“Fundamentally, the thing we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like computing, it's about communication across different machines. When y…”
Ce Zhang Feb 8, 2024 ▶ 6:19 Building an open AI company - with Ce and Vipul of Together AI
LATENT SPACE Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai
LATENT SPACE Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact
Kernel fusion sacrifices flexibility for researchers experimenting with attention
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exa…”
Tri Dao Aug 3, 2023 ▶ 10:09 FlashAttention-2: Making Transformers 800% faster AND exact
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29 FlashAttention-2: Making Transformers 800% faster AND exact
LATENT SPACE Prediction Not checkable as stated
FlashAttention techniques generalize across accelerators with asymmetric memory
“I expect the idea to be broadly these ideas to be broadly applicable to different hardware. As long as, I think the main idea is you have, like, asymmetry in, in memory hierarchy, which tends to be everywhere, you know, in, in a lot of a lot of accelerators.”
Tri Dao Aug 3, 2023 ▶ 34:11 FlashAttention-2: Making Transformers 800% faster AND exact

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.