“Nobody is actually putting in a two million context prompt into these models.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dan Fu
AssertionNot checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan FuDec 24, 2024▶ 33:002024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Modern GPU compute primitives should be matrices, not floats
“We basically built a whole library just around this basic idea that all your basic compute primitives should not be a float, but it should be a matrix and everything should just be matrix compute.”
Dan FuDec 24, 2024▶ 30:292024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
PredictionOpen · timeframe Dec 2029
Fu: Real-time long-context video generation cannot use quadratic attention
“You're certainly not going to do a giant quadratic attention computation to try to run that.”
Dan FuDec 24, 2024▶ 31:332024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Efficient AI Architectures Are Dead on Arrival Without Hardware Co-Design
“Even if your model is theoretically more efficient, if somebody goes and runs it and it's two times slower one of the things that, that we've learned is that if you're in that situation, it's just going to be dead on arrival. So you want to be designing your a…”
Dan FuDec 24, 2024▶ 15:512024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
AssertionNot checkable as stated
Fu: AI21's Jamba Is the State of the Art Non-Transformer Model
“AI-II trained this hybrid MOE called Jamba that, that, that seems, that is currently the state of the art for these non-transformer architectures.”
Dan FuDec 24, 2024▶ 16:502024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Changing one PyTorch line requires a week of CUDA development
“If we decided to change one thing in PyTorch, like one line of PyTorch code is like a week of CUDA code at least.”
Dan FuDec 24, 2024▶ 29:382024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.