Tri Dao

6 statements across 2 episodes · 5 bullish · 0 bearish · 2 people on the record · first statement Aug 3, 2023 by Tri Dao · said 16 times in 12 episodes since 2023 · across every show →

On the record as a speaker too: Tri Dao's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Alessio Fanelli (6), Ali Taha (2), Quentin Anthony (1), Diego Bachman (1), Chris Lattner (1)

tap a year for its mentions
0043852023202420252026episodesmentions
0352023202420252026episodes it came up in
0012.5252023202420252026episodesmentions per episode
2026 2 mentions in 1 episode
2025 5 mentions in 5 episodes 1 per episode
2024 7 mentions in 5 episodes 1 per episode
2023 2 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Tri Dao, oldest first

Aug 3, 2023 positive
Prediction Not checkable as stated
Future capable AI models will require explicit reasoning modules
“And in the future, I think we can, we will need to design architecture that kind of explicitly have some kind of Reasoning module in it if we want to have much more capable models.”
Tri Dao Aug 3, 2023 ▶ 1:03:19 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 positive
Disclosure
Tri Dao joins Together AI as Chief Scientist
“Yeah, yeah, so I just joined this week actually, and it's been really exciting.”
Tri Dao Aug 3, 2023 ▶ 0:51 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 neutral
Prediction Not checkable as stated
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”
Tri Dao Aug 3, 2023 ▶ 16:07 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 positive
Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 positive
Insight
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29 FlashAttention-2: Making Transformers 800% faster AND exact
Jun 13, 2025 positive
Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.