Aug 3, 2023 · 1h 4m · latent-space
FlashAttention-2: Making Transformers 800% faster AND exact
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this technical interview, researcher Tri Dao discusses the architectural principles behind FlashAttention and FlashAttention-2, highlighting how memory-aware algorithmic design, hardware-software co-design, and open-source collaboration are reshaping transformer efficiency and the future of AI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Tri reframes the host's premise that FlashAttention is computationally linear, pointing out that computation remains quadratic while memory is made linear and hardware-aware.
Hardest push from the hosts ▶ 9:50 Challenging kernel fusion from database perspectivesThe host challenges the universality of kernel fusion by drawing on database concepts of observability and atomic rollback.
Biggest teaching moment ▶ 2:18 Linear memory scaling vs quadratic computeTri clearly educates the host on the distinction between FLOP count and wall-clock IO bottlenecks, correcting the misunderstanding that exact attention computation was reduced to linear.
The host holds their own ▶ 14:33 Deep dive into SRAM and HBM hardware specsThe host demonstrates strong technical knowledge by quoting exact memory capacities (40GB vs 20MB) and transfer bandwidths (1.5 TB/s vs 19 TB/s) between HBM and SRAM.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| FlashAttention: Exact Attention vs. Approximation Methods | 4 | 6 | 2 | 1 | The host assumes FlashAttention reduces algorithmic complexity from quadratic to linear computation, which Tri politely clarifies by explaining memory is linear while compute remains quadratic but hardware-friendly. | |
| I/O Awareness, Kernel Fusion, and Online Softmax | 5 | 5 | 0 | 2 | The host draws analogies to database transactions and atomic operations to ask about kernel fusion downsides. Tri explains the online softmax trick and compiler tradeoffs in CUDA. | |
| GPU Memory Hierarchy: HBM, SRAM, and Future Scaling | 7 | 4 | 0 | 1 | The host cites specific hardware throughput numbers (1.5 TB/s HBM vs 19 TB/s SRAM) and TSMC SRAM scaling limits. Tri responds with physical silicon constraints and parallels to 1980s disk sorting algorithms. | |
| Code Artifacts, Impact, and Hazy Research Group | 4 | 2 | 0 | 0 | The host asks about research impact and lab culture at Hazy Research. Tri explains why prioritizing code artifacts over purely theoretical papers drove adoption. | |
| Navigating Academia vs. Industry and Evaluation Benchmarks | 5 | 3 | 0 | 1 | The host asks how incoming researchers should evaluate academia versus industry labs and whether evals distort modeling. Tri outlines academia's role in long-tail evaluation and higher-risk bets. | |
| FlashAttention-2 Improvements and Hardware Implementation | 5 | 3 | 0 | 1 | The host inquires about FlashAttention-2 and hardware portability beyond Nvidia. Tri explains building atop Cutlass v3 primitives and optimizing kernel execution to achieve near matrix-multiply efficiency. | |
| The Hardware Lottery and Dedicated AI Accelerators | 6 | 4 | 0 | 1 | The host references Sara Hooker's hardware lottery paper and wafer-scale hardware like Cerebras. Tri extends the concept to a software framework lottery favoring Transformers. | |
| PhD Research Strategy: Balancing Fundamentals with Fast ML Trends | 4 | 2 | 0 | 0 | The host asks how PhD students avoid getting scooped or invalidated by fast-moving industrial trends. Tri shares his early work on matrix-vector fundamentals and balancing explore-exploit research. | |
| Evaluating Transformer Alternatives: State Space Models and RNNs | 5 | 5 | 0 | 0 | The host asks about alternatives to Transformers like SSMs and modern RNNs. Tri breaks down state-space models and why RNNs can achieve higher generation throughput by eliminating the KV cache memory bottleneck. | |
| Open Source AI Models, Licensing, and Dataset Incentives | 7 | 3 | 0 | 1 | The host articulates a taxonomy of open source AI across open weights, restricted licenses like Llama 2, and open datasets. Tri analyzes corporate incentives around releasing pretraining and instruction tuning data. | |
| Joining Together AI as Chief Scientist | 3 | 1 | 0 | 0 | The host asks Tri about joining Together AI as Chief Scientist, leading into the wrap-up lightning round. |