The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 7 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59 The Shape of Compute (Chris Lattner of Modular)
Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai
Insight
Zhang: ML systems optimization is fundamentally always about data movement
“Fundamentally, the thing we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like computing, it's about communication across different machines. When y…”
Ce Zhang Feb 8, 2024 ▶ 6:19 Building an open AI company - with Ce and Vipul of Together AI
Assertion Supported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri Dao Aug 3, 2023 ▶ 3:07 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Releasing highly optimized code mattered more than the FlashAttention paper
“I think when we were writing the paper, I remember sending an email to one of my advisors, like hey, I'm excited about this paper but I think the most important thing will be the artifact, which is the code. So I knew that, like, the code will be valuable and,…”
Tri Dao Aug 3, 2023 ▶ 18:29 FlashAttention-2: Making Transformers 800% faster AND exact
Insight
Kernel fusion sacrifices flexibility for researchers experimenting with attention
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exa…”
Tri Dao Aug 3, 2023 ▶ 10:09 FlashAttention-2: Making Transformers 800% faster AND exact
Prediction Not checkable as stated
FlashAttention techniques generalize across accelerators with asymmetric memory
“I expect the idea to be broadly these ideas to be broadly applicable to different hardware. As long as, I think the main idea is you have, like, asymmetry in, in memory hierarchy, which tends to be everywhere, you know, in, in a lot of a lot of accelerators.”
Tri Dao Aug 3, 2023 ▶ 34:11 FlashAttention-2: Making Transformers 800% faster AND exact
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.