Dec 24, 2024 · 42m · latent-space

2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]

Dan Fu · 24m spoken Eugene Cheah · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At NeurIPS 2024, Dan and Eugene examine the emergence of post-transformer architectures, detailing how State Space Models (SSMs) and RWKV overcome quadratic attention bottlenecks through subquadratic scaling, hardware-model co-design, and hybrid systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.7 Guest teaching 0.5 Guest disagreement 0.5 The hosts pushing back 0.5
05100:0015:0030:004:04–18:16 · The hosts as informed peer 0/10 Evolution of State Space Models and Subquadratic Primitives This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero.18:16–24:15 · The hosts as informed peer 0/10 RWKV Architecture Mechanics and Open Source Origins Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension.24:16–28:50 · The hosts as informed peer 0/10 Transformer Conversion and Hybrid Architecture Synergy Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation.28:52–32:35 · The hosts as informed peer 0/10 Hardware-Model Co-Design and Next-Generation Modalities Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment.32:35–36:27 · The hosts as informed peer 0/10 Perspectives on Retrieval-Augmented Generation and Fixed-State Memory The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers.36:28–42:43 · The hosts as informed peer 4/10 Audience Q&A on Context Scaling, VRAM, and Streaming The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks.4:04–18:16 · Guest teaching 0/10 Evolution of State Space Models and Subquadratic Primitives This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero.18:16–24:15 · Guest teaching 0/10 RWKV Architecture Mechanics and Open Source Origins Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension.24:16–28:50 · Guest teaching 0/10 Transformer Conversion and Hybrid Architecture Synergy Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation.28:52–32:35 · Guest teaching 0/10 Hardware-Model Co-Design and Next-Generation Modalities Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment.32:35–36:27 · Guest teaching 0/10 Perspectives on Retrieval-Augmented Generation and Fixed-State Memory The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers.36:28–42:43 · Guest teaching 3/10 Audience Q&A on Context Scaling, VRAM, and Streaming The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks.4:04–18:16 · Guest disagreement 0/10 Evolution of State Space Models and Subquadratic Primitives This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero.18:16–24:15 · Guest disagreement 0/10 RWKV Architecture Mechanics and Open Source Origins Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension.24:16–28:50 · Guest disagreement 0/10 Transformer Conversion and Hybrid Architecture Synergy Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation.28:52–32:35 · Guest disagreement 0/10 Hardware-Model Co-Design and Next-Generation Modalities Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment.32:35–36:27 · Guest disagreement 1/10 Perspectives on Retrieval-Augmented Generation and Fixed-State Memory The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers.36:28–42:43 · Guest disagreement 2/10 Audience Q&A on Context Scaling, VRAM, and Streaming The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks.4:04–18:16 · The hosts pushing back 0/10 Evolution of State Space Models and Subquadratic Primitives This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero.18:16–24:15 · The hosts pushing back 0/10 RWKV Architecture Mechanics and Open Source Origins Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension.24:16–28:50 · The hosts pushing back 0/10 Transformer Conversion and Hybrid Architecture Synergy Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation.28:52–32:35 · The hosts pushing back 0/10 Hardware-Model Co-Design and Next-Generation Modalities Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment.32:35–36:27 · The hosts pushing back 0/10 Perspectives on Retrieval-Augmented Generation and Fixed-State Memory The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers.36:28–42:43 · The hosts pushing back 3/10 Audience Q&A on Context Scaling, VRAM, and Streaming The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 37:49 Eugene polling the audience on real 2M context usage

Eugene playfully challenges the practical utility of extreme context windows by polling the audience on whether anyone actually uses multi-million token prompts.

Hardest push from the hosts ▶ 36:48 Host challenges practical necessity of massive context scaling

The host directly questions the premise of extreme context benchmarks, asking why anyone would submit a two-million token query when RAG exists.

Biggest teaching moment ▶ 38:35 Eugene explains the VRAM backpropagation bottleneck during training

Eugene educates the room on the reality of training recurrent architectures on long contexts, explaining that state unrolling still consumes massive VRAM during backpropagation.

The host holds their own ▶ 38:08 Host asserts frequent real-world use of massive context windows

When Eugene suggests almost nobody uses multi-million token windows, the host immediately counters from personal technical experience, asserting he uses it frequently.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Evolution of State Space Models and Subquadratic Primitives 0000 This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero.
RWKV Architecture Mechanics and Open Source Origins 0000 Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension.
Transformer Conversion and Hybrid Architecture Synergy 0000 Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation.
Hardware-Model Co-Design and Next-Generation Modalities 0000 Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment.
Perspectives on Retrieval-Augmented Generation and Fixed-State Memory 0010 The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers.
Audience Q&A on Context Scaling, VRAM, and Streaming 4323 The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks.

Statements from this episode (15)

Insight
Fu: Efficient AI Architectures Are Dead on Arrival Without Hardware Co-Design
“Even if your model is theoretically more efficient, if somebody goes and runs it and it's two times slower one of the things that, that we've learned is that if you're in that situation, it's just going to be dead on arrival. So you want to be designing your a…”
Dan Fu Dec 24, 2024 ▶ 15:51
Assertion Not checkable as stated
Fu: AI21's Jamba Is the State of the Art Non-Transformer Model
“AI-II trained this hybrid MOE called Jamba that, that, that seems, that is currently the state of the art for these non-transformer architectures.”
Dan Fu Dec 24, 2024 ▶ 16:50
Assertion Supported
Fu: Stanford and Arc Institute's DNA SSM Made the Cover of Science
“One of those gated, SSM gated states-based models ended up on the cover of science because a great group of folks went and trained some DNA models. So that's Michael Polley, Eric Yuen from Stanford and the Arc Institute.”
Dan Fu Dec 24, 2024 ▶ 17:31
Disclosure
Cheah: RWKV Trains Models on Over 100 Languages Targeting 200
“We actually train our models primarily on over a hundred language, which is another topic altogether and our goal is to train to even 200 languages to cover all languages in the world.”
Eugene Cheah Dec 24, 2024 ▶ 19:51
Insight
Cheah: RWKV Innovations Are Found Empirically Before Academic Rationalization
“Officially in the paper, I'll say we had this idea and we wrote it this way. The reality is someone came in the code, we tested it worked, and then we rationalized it.”
Eugene Cheah Dec 24, 2024 ▶ 22:46
Disclosure
Cheah: RWKV organization has less compute than a single Google researcher
“So our entire organization has less compute than a single researcher in Google.”
Eugene Cheah Dec 24, 2024 ▶ 24:32
Assertion Not checkable as stated
Cheah: Most enterprise AI workloads use 70B models under 32k context
“Majority of enterprise workload today is just on Senti B at under 32 K context line.”
Eugene Cheah Dec 24, 2024 ▶ 27:23
Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15
Insight
Fu: Changing one PyTorch line requires a week of CUDA development
“If we decided to change one thing in PyTorch, like one line of PyTorch code is like a week of CUDA code at least.”
Dan Fu Dec 24, 2024 ▶ 29:38
Insight
Fu: Modern GPU compute primitives should be matrices, not floats
“We basically built a whole library just around this basic idea that all your basic compute primitives should not be a float, but it should be a matrix and everything should just be matrix compute.”
Dan Fu Dec 24, 2024 ▶ 30:29
Prediction Open · timeframe Dec 2029
Fu: Real-time long-context video generation cannot use quadratic attention
“You're certainly not going to do a giant quadratic attention computation to try to run that.”
Dan Fu Dec 24, 2024 ▶ 31:33
Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan Fu Dec 24, 2024 ▶ 33:00
Assertion Supported
Cheah: RWKV operates with a 40-megabyte fixed state size
“Like, we, like, RWKV is running at 40 megabytes for its state.”
Eugene Cheah Dec 24, 2024 ▶ 34:34
Opinion
Dan Fu: Nobody is actually submitting 2M token prompts into LLMs
“Nobody is actually putting in a two million context prompt into these models.”
Dan Fu Dec 24, 2024 ▶ 37:33
Insight
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Eugene Cheah Dec 24, 2024 ▶ 41:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.