Tri Dao, creator of FlashAttention, explains the trade-offs of kernel fusion in machine learning architectures during a discussion on GPU memory optimization.
“When you do kernel fusion is a little bit you lose a little bit of flexibility in the sense that, hey, now you have for example, is flash attention is just a subroutine that you would call to do attention. But as a researcher, let's say you don't want that exact thing, right? You don't want just attention. Let's say you want some modification to attention... And so kernel fusion just means that, okay, and we have a subroutine that does the entire thing, but if you want to experiment with things, you won't be able to use that that fused kernel.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Tri Dao
PredictionPartly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri DaoAug 3, 2023▶ 11:39FlashAttention-2: Making Transformers 800% faster AND exact
PredictionNot checkable as stated
RNNs will outperform Transformers in batch generation and long sequences
“I am personally bullish on, on, on RNNs. I think RNNs they don't, they essentially summarize the past into a state vector. They have fixed size, so the size doesn't grow with the history. So that means that you don't need as much memory to keep around all the …”
Tri DaoAug 3, 2023▶ 49:14FlashAttention-2: Making Transformers 800% faster AND exact
PredictionNot checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Tri DaoAug 3, 2023▶ 54:14FlashAttention-2: Making Transformers 800% faster AND exact
Insight
FLOP counts do not necessarily correlate with wall-clock runtime
“Flops or floating point operations don't necessarily correlate with runtime. There are other factors like memory reading and writing, parallelism, and so on.”
Tri DaoAug 3, 2023▶ 5:30FlashAttention-2: Making Transformers 800% faster AND exact
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri DaoAug 3, 2023▶ 46:51FlashAttention-2: Making Transformers 800% faster AND exact
PredictionNot checkable as stated
Future capable AI models will require explicit reasoning modules
“And in the future, I think we can, we will need to design architecture that kind of explicitly have some kind of Reasoning module in it if we want to have much more capable models.”
Tri DaoAug 3, 2023▶ 1:03:19FlashAttention-2: Making Transformers 800% faster AND exact
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.