“And in the future, I think we can, we will need to design architecture that kind of explicitly have some kind of Reasoning module in it if we want to have much more capable models.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Tri Dao
PredictionPartly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri DaoAug 3, 2023▶ 11:39FlashAttention-2: Making Transformers 800% faster AND exact
PredictionNot checkable as stated
RNNs will outperform Transformers in batch generation and long sequences
“I am personally bullish on, on, on RNNs. I think RNNs they don't, they essentially summarize the past into a state vector. They have fixed size, so the size doesn't grow with the history. So that means that you don't need as much memory to keep around all the …”
Tri DaoAug 3, 2023▶ 49:14FlashAttention-2: Making Transformers 800% faster AND exact
PredictionNot checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Tri DaoAug 3, 2023▶ 54:14FlashAttention-2: Making Transformers 800% faster AND exact
Insight
FLOP counts do not necessarily correlate with wall-clock runtime
“Flops or floating point operations don't necessarily correlate with runtime. There are other factors like memory reading and writing, parallelism, and so on.”
Tri DaoAug 3, 2023▶ 5:30FlashAttention-2: Making Transformers 800% faster AND exact
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri DaoAug 3, 2023▶ 46:51FlashAttention-2: Making Transformers 800% faster AND exact
AssertionSupported
FlashAttention achieves 2x to 4x wall-clock speedup with linear memory
“So in the end, we ended up being, the memory is linear in sequence length. In terms of computation, it's still quadratic, but we managed to make it much more hardware friendly, and as a result, we do get wall clock speed up on the order of two to four X which …”
Tri DaoAug 3, 2023▶ 3:07FlashAttention-2: Making Transformers 800% faster AND exact
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.