Dec 24, 2024 · 42m · latent-space
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At NeurIPS 2024, Dan and Eugene examine the emergence of post-transformer architectures, detailing how State Space Models (SSMs) and RWKV overcome quadratic attention bottlenecks through subquadratic scaling, hardware-model co-design, and hybrid systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Eugene playfully challenges the practical utility of extreme context windows by polling the audience on whether anyone actually uses multi-million token prompts.
Hardest push from the hosts ▶ 36:48 Host challenges practical necessity of massive context scalingThe host directly questions the premise of extreme context benchmarks, asking why anyone would submit a two-million token query when RAG exists.
Biggest teaching moment ▶ 38:35 Eugene explains the VRAM backpropagation bottleneck during trainingEugene educates the room on the reality of training recurrent architectures on long contexts, explaining that state unrolling still consumes massive VRAM during backpropagation.
The host holds their own ▶ 38:08 Host asserts frequent real-world use of massive context windowsWhen Eugene suggests almost nobody uses multi-million token windows, the host immediately counters from personal technical experience, asserting he uses it frequently.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Evolution of State Space Models and Subquadratic Primitives | 0 | 0 | 0 | 0 | This is an uninterrupted slide presentation by Dan Fu tracing the history of state space models, linear attention, and subquadratic primitives. Because it is a solo technical lecture without host participation, all host interaction and combativeness scores are zero. | |
| RWKV Architecture Mechanics and Open Source Origins | 0 | 0 | 0 | 0 | Eugene Cheah delivers a presentation segment outlining RWKV's open-source origin, GPU cascading mechanics, and separation into time-mix and channel-mix. There is no host involvement or adversarial tension. | |
| Transformer Conversion and Hybrid Architecture Synergy | 0 | 0 | 0 | 0 | Eugene explains model weight conversion methods (QRWKV) and the counterintuitive performance boost seen in hybrid transformer-SSM models. The segment remains a pure monologue presentation. | |
| Hardware-Model Co-Design and Next-Generation Modalities | 0 | 0 | 0 | 0 | Dan Fu discusses CUDA co-design library ThunderKittens, matrix-level primitives, and real-time streaming video generation. No host interaction occurs during this segment. | |
| Perspectives on Retrieval-Augmented Generation and Fixed-State Memory | 0 | 0 | 1 | 0 | The guests discuss their hot takes on RAG, comparing fixed-state recurrent memory to human biological constraints and external database lookups. The dynamic is cooperative between co-speakers. | |
| Audience Q&A on Context Scaling, VRAM, and Streaming | 4 | 3 | 2 | 3 | The host opens Q&A with a provocative pushback asking who actually needs 2-million-token contexts when RAG exists, and notes his own frequent usage of Gemini's large window. The guests acknowledge the challenge and provide nuanced answers on VRAM scaling, streaming stability, and long-range benchmarks. |