Tri Dao

Assistant Professor, Princeton University · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistfounderexecutiveinvestor@tri_dao ↗tridao.me ↗

He specializes in hardware-aware machine learning systems and developed FlashAttention and the Mamba architecture to accelerate deep learning on GPUs. He leads research as Chief Scientist at Together AI and is an Assistant Professor of Computer Science at Princeton University.

17statements → 9claims → 4claims resolved → 75%fully supported → 3.47/5average certainty → 2.06/5average debate potential → 17said about them ↓

3 supported 1 partly supported 0 contradicted 5 not checkable as stated how the 9 claims stand · each chip opens the sources

6 predictions · 3 assertions · 1 opinion · 6 insights · 1 disclosure · every statement was checked. The predictions and assertions are the 9 claims: statements the public record can support or contradict. 4 are resolved, and 5 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Tri argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact

Expressed certainty vs assessment result

none yet certainty 1
50% certainty 2
none yet certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Tri Dao said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Prediction Partly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri Dao Aug 3, 2023 ▶ 11:39 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
RNNs will outperform Transformers in batch generation and long sequences
“I am personally bullish on, on, on RNNs. I think RNNs they don't, they essentially summarize the past into a state vector. They have fixed size, so the size doesn't grow with the history. So that means that you don't need as much memory to keep around all the …”
Tri Dao Aug 3, 2023 ▶ 49:14 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Tri Dao Aug 3, 2023 ▶ 54:14 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
FLOP counts do not necessarily correlate with wall-clock runtime
“Flops or floating point operations don't necessarily correlate with runtime. There are other factors like memory reading and writing, parallelism, and so on.”
Tri Dao Aug 3, 2023 ▶ 5:30 FlashAttention-2: Making Transformers 800% faster AND exact
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri Dao Aug 3, 2023 ▶ 46:51 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
Future capable AI models will require explicit reasoning modules
“And in the future, I think we can, we will need to design architecture that kind of explicitly have some kind of Reasoning module in it if we want to have much more capable models.”
Tri Dao Aug 3, 2023 ▶ 1:03:19 FlashAttention-2: Making Transformers 800% faster AND exact
Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Kernel fusion sacrifices flexibility for researchers experimenting with attention
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exa…”
Tri Dao Aug 3, 2023 ▶ 10:09 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”
Tri Dao Aug 3, 2023 ▶ 16:07 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29 FlashAttention-2: Making Transformers 800% faster AND exact
Assertion Supported
FlashAttention-2 is twice as fast as FlashAttention-1
“We managed to make it to X faster. And now it's pretty close to probably the efficiency of things like matrix multiply, which probably this, the most optimized subroutine on the planet.”
Tri Dao Aug 3, 2023 ▶ 33:01 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
FlashAttention techniques generalize across accelerators with asymmetric memory
“I expect the idea to be broadly these ideas to be broadly applicable to different hardware. As long as, I think the main idea is you have, like, asymmetry in, in memory hierarchy, which tends to be everywhere, you know, in, in a lot of a lot of accelerators.”
Tri Dao Aug 3, 2023 ▶ 34:11 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Hardware and software co-evolve to favor dominant AI architectures
“There is this feedback loop where somehow The model architectures that take advantage of hardware become popular, and the hardware will also kind of evolve to optimize a little bit for that kind of architecture, and software framework software frameworks will …”
Tri Dao Aug 3, 2023 ▶ 36:22 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Multi-year hardware cycles make betting on future AI architectures difficult
“Hardware has, my understanding is has a kind of a longer time scale. So you need to design hardware, you need to manufacture it, you know, maybe on the order of three to five years or something like that. So you know, people are taking different bets but the, …”
Tri Dao Aug 3, 2023 ▶ 40:02 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
High human labeling costs keep instruction datasets closed-source
“These companies still, they do pay for human labelers, right? To annotate these instruction tuning data set. And that is expensive. Right. And maybe, you know, they will see that as their competitive advantage. And so it's harder to incentivize these companies…”
Tri Dao Aug 3, 2023 ▶ 57:46 FlashAttention-2: Making Transformers 800% faster AND exact
Disclosure
Tri Dao joins Together AI as Chief Scientist
“Yeah, yeah, so I just joined this week actually, and it's been really exciting.”
Tri Dao Aug 3, 2023 ▶ 0:51 FlashAttention-2: Making Transformers 800% faster AND exact
Assertion Supported
Memory read/write dominates standard attention computation time
“We ended up focusing a lot more on Memory reading and writing, because that turned out to be the majority of time when you're doing attention is reading and writing memory.”
Tri Dao Aug 3, 2023 ▶ 5:54 FlashAttention-2: Making Transformers 800% faster AND exact

The other half of the tape: Tri Dao's own voice is left out of every number here. Other people bring the name up 16 times in 12 episodes on Latent Space. 1 statement on the record names them. every mention, with the transcript →

Who brings them up most Alessio Fanelli 6Ali Taha 2Quentin Anthony 1Diego Bachman 1Chris Lattner 1

Statements about Tri Dao, by other people (1)

Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)

Every mention by year

tap a year for its mentions
0043852023202420252026episodesmentions
0352023202420252026episodes it came up in
0012.5252023202420252026episodesmentions per episode
2026 2 mentions in 1 episode
2025 5 mentions in 5 episodes 1 per episode
2024 7 mentions in 5 episodes 1 per episode
2023 2 mentions in 1 episode

Appearances (1)

EpisodeDateSpeaking time
FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 49m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.