Assertion certainty 4/5 debate potential 2/5

Fu: AI21's Jamba Is the State of the Art Non-Transformer Model

Dan Fu · 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024] · Dec 24, 2024 · at 16:50

Dan Fu reviews the leading post-transformer implementations at NeurIPS 2024.

0:00 / 0:08exact quote · 8.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“AI-II trained this hybrid MOE called Jamba that, that, that seems, that is currently the state of the art for these non-transformer architectures.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dan Fu

Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan Fu Dec 24, 2024 ▶ 33:00 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Modern GPU compute primitives should be matrices, not floats
“We basically built a whole library just around this basic idea that all your basic compute primitives should not be a float, but it should be a matrix and everything should just be matrix compute.”
Dan Fu Dec 24, 2024 ▶ 30:29 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Prediction Open · timeframe Dec 2029
Fu: Real-time long-context video generation cannot use quadratic attention
“You're certainly not going to do a giant quadratic attention computation to try to run that.”
Dan Fu Dec 24, 2024 ▶ 31:33 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Efficient AI Architectures Are Dead on Arrival Without Hardware Co-Design
“Even if your model is theoretically more efficient, if somebody goes and runs it and it's two times slower one of the things that, that we've learned is that if you're in that situation, it's just going to be dead on arrival. So you want to be designing your a…”
Dan Fu Dec 24, 2024 ▶ 15:51 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Opinion
Dan Fu: Nobody is actually submitting 2M token prompts into LLMs
“Nobody is actually putting in a two million context prompt into these models.”
Dan Fu Dec 24, 2024 ▶ 37:33 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Fu: Changing one PyTorch line requires a week of CUDA development
“If we decided to change one thing in PyTorch, like one line of PyTorch code is like a week of CUDA code at least.”
Dan Fu Dec 24, 2024 ▶ 29:38 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.