Aug 31, 2023 · 1h 56m · latent-space
RWKV: Reinventing RNNs for the Transformer Era
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this extensive discussion, UIlicious CTO Eugene Cheah details the architecture, grassroots open-source origins, and multilingual capabilities of RWKV, demonstrating how linear-time RNNs can match Transformer performance while drastically reducing compute and memory overhead.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Eugene dismisses academic rationalizations and bluntly admits that the RWKV community simply sucks at marketing compared to rival model creators.
Hardest push from the hosts ▶ 1:31:00 Host challenges text diffusion latencyThe host rejects Eugene's text diffusion concept, arguing that streaming autoregressive tokens is essential for acceptable user experience.
Biggest teaching moment ▶ 33:20 Educating the host on Attention-Free TransformersEugene introduces Apple's Attention-Free Transformer paper to explain how RWKV eliminates multi-head attention entirely, catching the host completely off guard.
The host holds their own ▶ 1:42:50 Host cites Justine Tunney's GGML mmap breakthroughThe host immediately identifies and contextualizes Justine Tunney's contribution of memory mapping to llama.cpp/GGML to prove his grasp of systems-level ML optimization.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Eugene Cheah's Background and the Creation of GPU.js | 6 | 5 | 1 | 2 | The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs. | |
| Building UIlicious and Applying Code Models for Automation | 6 | 4 | 1 | 2 | Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning. | |
| The Context Window Bottleneck and Need for Alternatives | 6 | 6 | 2 | 3 | Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits. | |
| Grassroots AI Communities and Multilingual RWKV Adoption | 6 | 5 | 1 | 2 | Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities. | |
| Deconstructing RWKV Architecture and Attention-Free Mechanisms | 5 | 7 | 2 | 3 | Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations. | |
| RWKV Model Lineage and Multilingual World Datasets | 5 | 6 | 1 | 2 | Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios. | |
| Human Evals and Redesigning Tokenization with Tries | 5 | 7 | 2 | 2 | Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities. | |
| Multimodal Capabilities and the Genesis of Blink's Research | 4 | 7 | 2 | 2 | Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability. | |
| Parallel Training Mechanics: Time-Mix and Channel-Mix Explained | 6 | 7 | 1 | 2 | Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity. | |
| Linear Computational Complexity and Low-Resource Edge Deployments | 6 | 6 | 1 | 2 | Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints. | |
| Adoption Roadblocks, Context Scaling Limits, and RWKV v5 | 7 | 6 | 2 | 4 | The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions. | |
| Decentralized Open-Source Governance and the RWKV Foundation | 6 | 5 | 1 | 2 | Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources. | |
| Text Diffusion Models and Tackling the Token Crisis | 6 | 7 | 2 | 4 | Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements. | |
| Engineering Pathways in AI: From Prompting to Systems | 7 | 6 | 2 | 2 | Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment. | |
| Character Alignment and Optimization in Persona-Driven AI | 6 | 7 | 2 | 2 | Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness. |