Aug 3, 2023 · 1h 4m · latent-space

FlashAttention-2: Making Transformers 800% faster AND exact

Tri Dao · 49m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this technical interview, researcher Tri Dao discusses the architectural principles behind FlashAttention and FlashAttention-2, highlighting how memory-aware algorithmic design, hardware-software co-design, and open-source collaboration are reshaping transformer efficiency and the future of AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.0 Guest teaching 3.5 Guest disagreement 0.2 The hosts pushing back 0.7
05100:0015:0030:0045:001:00:001:39–6:05 · The hosts as informed peer 4/10 FlashAttention: Exact Attention vs. Approximation Methods The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly.6:05–11:53 · The hosts as informed peer 5/10 I/O Awareness, Kernel Fusion, and Online Softmax The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA.11:53–17:43 · The hosts as informed peer 7/10 GPU Memory Hierarchy: HBM, SRAM, and Future Scaling The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms.17:44–22:35 · The hosts as informed peer 4/10 Code Artifacts, Impact, and Hazy Research Group The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption.22:36–31:09 · The hosts as informed peer 5/10 Navigating Academia vs. Industry and Evaluation Benchmarks The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets.31:09–34:37 · The hosts as informed peer 5/10 FlashAttention-2 Improvements and Hardware Implementation The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency.34:37–40:41 · The hosts as informed peer 6/10 The Hardware Lottery and Dedicated AI Accelerators The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers.40:42–44:13 · The hosts as informed peer 4/10 PhD Research Strategy: Balancing Fundamentals with Fast ML Trends The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research.44:17–50:48 · The hosts as informed peer 5/10 Evaluating Transformer Alternatives: State Space Models and RNNs The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck.50:48–1:00:51 · The hosts as informed peer 7/10 Open Source AI Models, Licensing, and Dataset Incentives The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data.1:00:51–1:02:22 · The hosts as informed peer 3/10 Joining Together AI as Chief Scientist The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round.1:39–6:05 · Guest teaching 6/10 FlashAttention: Exact Attention vs. Approximation Methods The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly.6:05–11:53 · Guest teaching 5/10 I/O Awareness, Kernel Fusion, and Online Softmax The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA.11:53–17:43 · Guest teaching 4/10 GPU Memory Hierarchy: HBM, SRAM, and Future Scaling The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms.17:44–22:35 · Guest teaching 2/10 Code Artifacts, Impact, and Hazy Research Group The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption.22:36–31:09 · Guest teaching 3/10 Navigating Academia vs. Industry and Evaluation Benchmarks The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets.31:09–34:37 · Guest teaching 3/10 FlashAttention-2 Improvements and Hardware Implementation The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency.34:37–40:41 · Guest teaching 4/10 The Hardware Lottery and Dedicated AI Accelerators The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers.40:42–44:13 · Guest teaching 2/10 PhD Research Strategy: Balancing Fundamentals with Fast ML Trends The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research.44:17–50:48 · Guest teaching 5/10 Evaluating Transformer Alternatives: State Space Models and RNNs The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck.50:48–1:00:51 · Guest teaching 3/10 Open Source AI Models, Licensing, and Dataset Incentives The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data.1:00:51–1:02:22 · Guest teaching 1/10 Joining Together AI as Chief Scientist The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round.1:39–6:05 · Guest disagreement 2/10 FlashAttention: Exact Attention vs. Approximation Methods The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly.6:05–11:53 · Guest disagreement 0/10 I/O Awareness, Kernel Fusion, and Online Softmax The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA.11:53–17:43 · Guest disagreement 0/10 GPU Memory Hierarchy: HBM, SRAM, and Future Scaling The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms.17:44–22:35 · Guest disagreement 0/10 Code Artifacts, Impact, and Hazy Research Group The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption.22:36–31:09 · Guest disagreement 0/10 Navigating Academia vs. Industry and Evaluation Benchmarks The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets.31:09–34:37 · Guest disagreement 0/10 FlashAttention-2 Improvements and Hardware Implementation The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency.34:37–40:41 · Guest disagreement 0/10 The Hardware Lottery and Dedicated AI Accelerators The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers.40:42–44:13 · Guest disagreement 0/10 PhD Research Strategy: Balancing Fundamentals with Fast ML Trends The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research.44:17–50:48 · Guest disagreement 0/10 Evaluating Transformer Alternatives: State Space Models and RNNs The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck.50:48–1:00:51 · Guest disagreement 0/10 Open Source AI Models, Licensing, and Dataset Incentives The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data.1:00:51–1:02:22 · Guest disagreement 0/10 Joining Together AI as Chief Scientist The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round.1:39–6:05 · The hosts pushing back 1/10 FlashAttention: Exact Attention vs. Approximation Methods The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly.6:05–11:53 · The hosts pushing back 2/10 I/O Awareness, Kernel Fusion, and Online Softmax The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA.11:53–17:43 · The hosts pushing back 1/10 GPU Memory Hierarchy: HBM, SRAM, and Future Scaling The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms.17:44–22:35 · The hosts pushing back 0/10 Code Artifacts, Impact, and Hazy Research Group The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption.22:36–31:09 · The hosts pushing back 1/10 Navigating Academia vs. Industry and Evaluation Benchmarks The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets.31:09–34:37 · The hosts pushing back 1/10 FlashAttention-2 Improvements and Hardware Implementation The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency.34:37–40:41 · The hosts pushing back 1/10 The Hardware Lottery and Dedicated AI Accelerators The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers.40:42–44:13 · The hosts pushing back 0/10 PhD Research Strategy: Balancing Fundamentals with Fast ML Trends The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research.44:17–50:48 · The hosts pushing back 0/10 Evaluating Transformer Alternatives: State Space Models and RNNs The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck.50:48–1:00:51 · The hosts pushing back 1/10 Open Source AI Models, Licensing, and Dataset Incentives The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data.1:00:51–1:02:22 · The hosts pushing back 0/10 Joining Together AI as Chief Scientist The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 2:18 Gentle correction on linear attention complexity

Tri reframes the host's premise that FlashAttention is computationally linear, pointing out that computation remains quadratic while memory is made linear and hardware-aware.

Hardest push from the hosts ▶ 9:50 Challenging kernel fusion from database perspectives

The host challenges the universality of kernel fusion by drawing on database concepts of observability and atomic rollback.

Biggest teaching moment ▶ 2:18 Linear memory scaling vs quadratic compute

Tri clearly educates the host on the distinction between FLOP count and wall-clock IO bottlenecks, correcting the misunderstanding that exact attention computation was reduced to linear.

The host holds their own ▶ 14:33 Deep dive into SRAM and HBM hardware specs

The host demonstrates strong technical knowledge by quoting exact memory capacities (40GB vs 20MB) and transfer bandwidths (1.5 TB/s vs 19 TB/s) between HBM and SRAM.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
FlashAttention: Exact Attention vs. Approximation Methods 4621 The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly.
I/O Awareness, Kernel Fusion, and Online Softmax 5502 The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA.
GPU Memory Hierarchy: HBM, SRAM, and Future Scaling 7401 The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms.
Code Artifacts, Impact, and Hazy Research Group 4200 The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption.
Navigating Academia vs. Industry and Evaluation Benchmarks 5301 The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets.
FlashAttention-2 Improvements and Hardware Implementation 5301 The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency.
The Hardware Lottery and Dedicated AI Accelerators 6401 The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers.
PhD Research Strategy: Balancing Fundamentals with Fast ML Trends 4200 The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research.
Evaluating Transformer Alternatives: State Space Models and RNNs 5500 The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck.
Open Source AI Models, Licensing, and Dataset Incentives 7301 The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data.
Joining Together AI as Chief Scientist 3100 The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round.

Statements from this episode (17)

Disclosure
Tri Dao joins Together AI as Chief Scientist
“Yeah, yeah, so I just joined this week actually, and it's been really exciting.”
Tri Dao Aug 3, 2023 ▶ 0:51
Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07
Insight
FLOP counts do not necessarily correlate with wall-clock runtime
“Flops or floating point operations don't necessarily correlate with runtime. There are other factors like memory reading and writing, parallelism, and so on.”
Tri Dao Aug 3, 2023 ▶ 5:30
Assertion Supported
Memory read/write dominates standard attention computation time
“We ended up focusing a lot more on Memory reading and writing, because that turned out to be the majority of time when you're doing attention is reading and writing memory.”
Tri Dao Aug 3, 2023 ▶ 5:54
Insight
Kernel fusion sacrifices flexibility for researchers experimenting with attention
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exa…”
Tri Dao Aug 3, 2023 ▶ 10:09
Prediction Partly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri Dao Aug 3, 2023 ▶ 11:39
Prediction Not checkable as stated
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”
Tri Dao Aug 3, 2023 ▶ 16:07
Insight
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29
Assertion Supported
FlashAttention-2 is twice as fast as FlashAttention-1
“We managed to make it to X faster. And now it's pretty close to probably the efficiency of things like matrix multiply, which probably this, the most optimized subroutine on the planet.”
Tri Dao Aug 3, 2023 ▶ 33:01
Prediction Not checkable as stated
FlashAttention techniques generalize across accelerators with asymmetric memory
“I expect the idea to be broadly these ideas to be broadly applicable to different hardware. As long as, I think the main idea is you have, like, asymmetry in, in memory hierarchy, which tends to be everywhere, you know, in, in a lot of a lot of accelerators.”
Tri Dao Aug 3, 2023 ▶ 34:11
Insight
Hardware and software co-evolve to favor dominant AI architectures
“There is this feedback loop where somehow The model architectures that take advantage of hardware become popular, and the hardware will also kind of evolve to optimize a little bit for that kind of architecture, and software framework software frameworks will …”
Tri Dao Aug 3, 2023 ▶ 36:22
Insight
Multi-year hardware cycles make betting on future AI architectures difficult
“Hardware has, my understanding is has a kind of a longer time scale. So you need to design hardware, you need to manufacture it, you know, maybe on the order of three to five years or something like that. So you know, people are taking different bets but the, …”
Tri Dao Aug 3, 2023 ▶ 40:02
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri Dao Aug 3, 2023 ▶ 46:51
Prediction Not checkable as stated
RNNs will outperform Transformers in batch generation and long sequences
“I am personally bullish on, on, on RNNs. I think RNNs they don't, they essentially summarize the past into a state vector. They have fixed size, so the size doesn't grow with the history. So that means that you don't need as much memory to keep around all the …”
Tri Dao Aug 3, 2023 ▶ 49:14
Prediction Not checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Tri Dao Aug 3, 2023 ▶ 54:14
Insight
High human labeling costs keep instruction datasets closed-source
“These companies still, they do pay for human labelers, right? To annotate these instruction tuning data set. And that is expensive. Right. And maybe, you know, they will see that as their competitive advantage. And so it's harder to incentivize these companies…”
Tri Dao Aug 3, 2023 ▶ 57:46
Prediction Not checkable as stated
Future capable AI models will require explicit reasoning modules
“And in the future, I think we can, we will need to design architecture that kind of explicitly have some kind of Reasoning module in it if we want to have much more capable models.”
Tri Dao Aug 3, 2023 ▶ 1:03:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.