Aug 31, 2023 · 1h 56m · latent-space

RWKV: Reinventing RNNs for the Transformer Era

Eugene Cheah · 1h 15m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this extensive discussion, UIlicious CTO Eugene Cheah details the architecture, grassroots open-source origins, and multilingual capabilities of RWKV, demonstrating how linear-time RNNs can match Transformer performance while drastically reducing compute and memory overhead.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.8 Guest teaching 6.1 Guest disagreement 1.5 The hosts pushing back 2.4
05100:0020:0040:001:00:001:20:001:40:000:01–5:45 · The hosts as informed peer 6/10 Eugene Cheah's Background and the Creation of GPU.js The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs.5:45–13:29 · The hosts as informed peer 6/10 Building UIlicious and Applying Code Models for Automation Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning.13:29–20:00 · The hosts as informed peer 6/10 The Context Window Bottleneck and Need for Alternatives Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits.20:01–31:44 · The hosts as informed peer 6/10 Grassroots AI Communities and Multilingual RWKV Adoption Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities.31:46–37:40 · The hosts as informed peer 5/10 Deconstructing RWKV Architecture and Attention-Free Mechanisms Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations.37:42–42:28 · The hosts as informed peer 5/10 RWKV Model Lineage and Multilingual World Datasets Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios.42:29–50:25 · The hosts as informed peer 5/10 Human Evals and Redesigning Tokenization with Tries Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities.50:25–56:20 · The hosts as informed peer 4/10 Multimodal Capabilities and the Genesis of Blink's Research Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability.56:26–1:02:31 · The hosts as informed peer 6/10 Parallel Training Mechanics: Time-Mix and Channel-Mix Explained Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity.1:02:38–1:07:23 · The hosts as informed peer 6/10 Linear Computational Complexity and Low-Resource Edge Deployments Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints.1:07:24–1:15:52 · The hosts as informed peer 7/10 Adoption Roadblocks, Context Scaling Limits, and RWKV v5 The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions.1:15:58–1:25:14 · The hosts as informed peer 6/10 Decentralized Open-Source Governance and the RWKV Foundation Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources.1:25:16–1:32:36 · The hosts as informed peer 6/10 Text Diffusion Models and Tackling the Token Crisis Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements.1:32:36–1:45:34 · The hosts as informed peer 7/10 Engineering Pathways in AI: From Prompting to Systems Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment.1:45:36–1:52:10 · The hosts as informed peer 6/10 Character Alignment and Optimization in Persona-Driven AI Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness.0:01–5:45 · Guest teaching 5/10 Eugene Cheah's Background and the Creation of GPU.js The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs.5:45–13:29 · Guest teaching 4/10 Building UIlicious and Applying Code Models for Automation Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning.13:29–20:00 · Guest teaching 6/10 The Context Window Bottleneck and Need for Alternatives Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits.20:01–31:44 · Guest teaching 5/10 Grassroots AI Communities and Multilingual RWKV Adoption Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities.31:46–37:40 · Guest teaching 7/10 Deconstructing RWKV Architecture and Attention-Free Mechanisms Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations.37:42–42:28 · Guest teaching 6/10 RWKV Model Lineage and Multilingual World Datasets Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios.42:29–50:25 · Guest teaching 7/10 Human Evals and Redesigning Tokenization with Tries Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities.50:25–56:20 · Guest teaching 7/10 Multimodal Capabilities and the Genesis of Blink's Research Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability.56:26–1:02:31 · Guest teaching 7/10 Parallel Training Mechanics: Time-Mix and Channel-Mix Explained Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity.1:02:38–1:07:23 · Guest teaching 6/10 Linear Computational Complexity and Low-Resource Edge Deployments Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints.1:07:24–1:15:52 · Guest teaching 6/10 Adoption Roadblocks, Context Scaling Limits, and RWKV v5 The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions.1:15:58–1:25:14 · Guest teaching 5/10 Decentralized Open-Source Governance and the RWKV Foundation Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources.1:25:16–1:32:36 · Guest teaching 7/10 Text Diffusion Models and Tackling the Token Crisis Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements.1:32:36–1:45:34 · Guest teaching 6/10 Engineering Pathways in AI: From Prompting to Systems Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment.1:45:36–1:52:10 · Guest teaching 7/10 Character Alignment and Optimization in Persona-Driven AI Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness.0:01–5:45 · Guest disagreement 1/10 Eugene Cheah's Background and the Creation of GPU.js The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs.5:45–13:29 · Guest disagreement 1/10 Building UIlicious and Applying Code Models for Automation Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning.13:29–20:00 · Guest disagreement 2/10 The Context Window Bottleneck and Need for Alternatives Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits.20:01–31:44 · Guest disagreement 1/10 Grassroots AI Communities and Multilingual RWKV Adoption Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities.31:46–37:40 · Guest disagreement 2/10 Deconstructing RWKV Architecture and Attention-Free Mechanisms Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations.37:42–42:28 · Guest disagreement 1/10 RWKV Model Lineage and Multilingual World Datasets Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios.42:29–50:25 · Guest disagreement 2/10 Human Evals and Redesigning Tokenization with Tries Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities.50:25–56:20 · Guest disagreement 2/10 Multimodal Capabilities and the Genesis of Blink's Research Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability.56:26–1:02:31 · Guest disagreement 1/10 Parallel Training Mechanics: Time-Mix and Channel-Mix Explained Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity.1:02:38–1:07:23 · Guest disagreement 1/10 Linear Computational Complexity and Low-Resource Edge Deployments Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints.1:07:24–1:15:52 · Guest disagreement 2/10 Adoption Roadblocks, Context Scaling Limits, and RWKV v5 The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions.1:15:58–1:25:14 · Guest disagreement 1/10 Decentralized Open-Source Governance and the RWKV Foundation Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources.1:25:16–1:32:36 · Guest disagreement 2/10 Text Diffusion Models and Tackling the Token Crisis Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements.1:32:36–1:45:34 · Guest disagreement 2/10 Engineering Pathways in AI: From Prompting to Systems Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment.1:45:36–1:52:10 · Guest disagreement 2/10 Character Alignment and Optimization in Persona-Driven AI Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness.0:01–5:45 · The hosts pushing back 2/10 Eugene Cheah's Background and the Creation of GPU.js The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs.5:45–13:29 · The hosts pushing back 2/10 Building UIlicious and Applying Code Models for Automation Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning.13:29–20:00 · The hosts pushing back 3/10 The Context Window Bottleneck and Need for Alternatives Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits.20:01–31:44 · The hosts pushing back 2/10 Grassroots AI Communities and Multilingual RWKV Adoption Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities.31:46–37:40 · The hosts pushing back 3/10 Deconstructing RWKV Architecture and Attention-Free Mechanisms Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations.37:42–42:28 · The hosts pushing back 2/10 RWKV Model Lineage and Multilingual World Datasets Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios.42:29–50:25 · The hosts pushing back 2/10 Human Evals and Redesigning Tokenization with Tries Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities.50:25–56:20 · The hosts pushing back 2/10 Multimodal Capabilities and the Genesis of Blink's Research Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability.56:26–1:02:31 · The hosts pushing back 2/10 Parallel Training Mechanics: Time-Mix and Channel-Mix Explained Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity.1:02:38–1:07:23 · The hosts pushing back 2/10 Linear Computational Complexity and Low-Resource Edge Deployments Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints.1:07:24–1:15:52 · The hosts pushing back 4/10 Adoption Roadblocks, Context Scaling Limits, and RWKV v5 The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions.1:15:58–1:25:14 · The hosts pushing back 2/10 Decentralized Open-Source Governance and the RWKV Foundation Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources.1:25:16–1:32:36 · The hosts pushing back 4/10 Text Diffusion Models and Tackling the Token Crisis Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements.1:32:36–1:45:34 · The hosts pushing back 2/10 Engineering Pathways in AI: From Prompting to Systems Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment.1:45:36–1:52:10 · The hosts pushing back 2/10 Character Alignment and Optimization in Persona-Driven AI Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:08:15 Blunt admission of marketing failure

Eugene dismisses academic rationalizations and bluntly admits that the RWKV community simply sucks at marketing compared to rival model creators.

Hardest push from the hosts ▶ 1:31:00 Host challenges text diffusion latency

The host rejects Eugene's text diffusion concept, arguing that streaming autoregressive tokens is essential for acceptable user experience.

Biggest teaching moment ▶ 33:20 Educating the host on Attention-Free Transformers

Eugene introduces Apple's Attention-Free Transformer paper to explain how RWKV eliminates multi-head attention entirely, catching the host completely off guard.

The host holds their own ▶ 1:42:50 Host cites Justine Tunney's GGML mmap breakthrough

The host immediately identifies and contextualizes Justine Tunney's contribution of memory mapping to llama.cpp/GGML to prove his grasp of systems-level ML optimization.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Eugene Cheah's Background and the Creation of GPU.js 6512 The host inquires into Eugene's background with GPU.js and AST compilation to WebGL shaders. Eugene explains the origins and mechanics of running JS linear algebra on GPUs.
Building UIlicious and Applying Code Models for Automation 6412 Eugene outlines building UIlicious and training CodeGen models for web automation. The host notes licensing nuances and the transition from pre-transformer dataset scales to foundation fine-tuning.
The Context Window Bottleneck and Need for Alternatives 6623 Eugene discusses hitting quadratic scaling walls when tokenizing full HTML pages for UI testing. The host questions tokenizer design and human working memory limits.
Grassroots AI Communities and Multilingual RWKV Adoption 6512 Eugene explains discovering RWKV on grassroots forums and why non-English communities adopted it over US-centric models. The host connects this to uncensored and character AI communities.
Deconstructing RWKV Architecture and Attention-Free Mechanisms 5723 Eugene details how RWKV discards multi-head attention using Apple's Attention-Free Transformer (AFT) paper. The host expresses surprise at attention-free transformer implementations.
RWKV Model Lineage and Multilingual World Datasets 5612 Eugene walks through the model lineup including Raven, World, and Pile Plus. The host asks about data provenance, GPT-4 dataset scrubbing, and multilinguality ratios.
Human Evals and Redesigning Tokenization with Tries 5722 Eugene breaks down RWKV's custom trie-based tokenizer that eliminates space-delimited assumptions for Asian languages. The host probes dictionary size and character-level capabilities.
Multimodal Capabilities and the Genesis of Blink's Research 4722 Eugene outlines multimodal RWKV experiments in vision and MIDI music alongside Bo Peng's (Blink) history. The host explores how an independent researcher secured compute backing from Eleuther and Stability.
Parallel Training Mechanics: Time-Mix and Channel-Mix Explained 6712 Eugene illustrates how time-mix and channel-mix allow layer cascading to saturate GPUs without sequential bottlenecks. The host notes parallels to recurrent networks and Big-O complexity.
Linear Computational Complexity and Low-Resource Edge Deployments 6612 Eugene explains O(1) inference memory savings and edge execution on Raspberry Pis. The host probes parameter memory vs working context memory footprints.
Adoption Roadblocks, Context Scaling Limits, and RWKV v5 7624 The host presses Eugene on why RWKV hasn't gained mainstream mindshare. Eugene points to marketing deficiencies and context degradation on long document QA, outlining RWKV v5 solutions.
Decentralized Open-Source Governance and the RWKV Foundation 6512 Eugene describes the self-organizing Discord channels and plans for an RWKV non-profit foundation. The host offers to connect the team with non-profit organizational resources.
Text Diffusion Models and Tackling the Token Crisis 6724 Eugene pitches applying diffusion to text models to enable multi-epoch training and bypass the token crisis. The host pushes back by citing streaming generation latency and UX requirements.
Engineering Pathways in AI: From Prompting to Systems 7622 Eugene maps out learning paths from prompt engineering to low-level systems programming. The host and Eugene discuss software engineers making high-impact ML contributions via mmap and CUDA memory alignment.
Character Alignment and Optimization in Persona-Driven AI 6722 Eugene discusses character persona alignment and why LoRAs excel at persona speech patterns rather than net-new knowledge. The host draws connections to personality modeling and digital consciousness.

Statements from this episode (30)

Assertion Partly supported
GPU.js Outperforms V8 on Matrices Over 2,000 Dimensions
“It outperformed the base VA engine by running it on the WebGL. Well, especially when you scale past 2000 dimensions there is a gotcha because you have to transfer your variables from the JavaScript space to the GPU space. So anything less than a thousand by th…”
Eugene Cheah Aug 31, 2023 ▶ 1:49
Insight
Fine-Tuning Transformers Requires Only 1,000 to 10,000 Data Examples
“Post-Transformer it's literally, you probably need only, like, a thousand or 10,000, enough, like, data that you can literally get an intern a few weeks to just get it done, and you have a working model. It may not be that great, but frankly, every piece of da…”
Eugene Cheah Aug 31, 2023 ▶ 10:01
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25
Assertion Not checkable as stated
Most Open-Source LLMs Fall Flat Outside English-Speaking Nations
“Beyond OpenAI's model, and beyond ChatGPT and Claudia, the two big models, right? Outside of the English speaking nations, right? A lot of the open source models really fall flat.”
Eugene Cheah Aug 31, 2023 ▶ 23:00
Insight
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
Eugene Cheah Aug 31, 2023 ▶ 25:13
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52
Disclosure
RWKV Raven Dataset Scrubs Out 'As an AI' Refusal Boilerplate
“Typically GPT for all, but then we scrub it for and remove all the, as a large model.”
Eugene Cheah Aug 31, 2023 ▶ 39:46
Disclosure
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
Eugene Cheah Aug 31, 2023 ▶ 43:11
Disclosure
RWKV Relies Exclusively on Native Speakers for Multilingual Quality Evals
“Our informal rule is that, the only person who can decide if whether this Improved word model is better in Japanese, or Thai, or whatever it is, it's a native speaker.”
Eugene Cheah Aug 31, 2023 ▶ 43:26
Assertion Not checkable as stated
GPT-NeoX Documentation Became Reference Notes for Subsequent Open-Source LLMs
“GPT Neo X was that it was one of the major models that had everything fully documented and they like, why they make this change in the architecture and so on and so forth. And that became like Basically reference notes for all other subsequent open source mode…”
Eugene Cheah Aug 31, 2023 ▶ 46:33
Assertion Supported
RWKV Uses Trie Tokenizer Without Space Delimiters for CJK Languages
“Instead of using like this token pairs well with this and should be paired with that we just made it a trial list. So So basically, try the data structure. Yeah. So we just find the longest matching string in that matching string that we have trained inside ou…”
Eugene Cheah Aug 31, 2023 ▶ 48:13
Assertion Supported
RWKV Community Built Multimodal Vision Models Using MiniGPT-4 Architecture
“There is actually another project internally on the discord where it's Doing vision, ah, vision modeling, and this based on the, ah, is it, mini GPT-Fall paper, where, where you have an image model, put everything inside the latent space, and then you have the…”
Eugene Cheah Aug 31, 2023 ▶ 50:49
Assertion Supported
Bo Peng Created RWKV on EleutherAI Forum to Parallelize RNNs
“Blink, or Bopeng is the actual name decided basically as an individual, literally at the illiterate AI forum, decided that, hey I think we can modify recurrent neural networks, no, neural networks, based on the Apple paper, the light attention that I showed pr…”
Eugene Cheah Aug 31, 2023 ▶ 53:46
Assertion Supported
EleutherAI and Stability AI Donated A100 Compute to Train RWKV
“So, so that's why AI and the rest stability, I believe also is involved, stepped up. And donated the A-One-Hundreds needed to train the basic models that RWKB had.”
Eugene Cheah Aug 31, 2023 ▶ 55:25
Assertion Supported
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
Eugene Cheah Aug 31, 2023 ▶ 58:45
Assertion Supported
RWKV Inference Computes at O(1) Complexity Per Token
“I'm talking about, like, to go through the entire context, yeah, this will be O one per token.”
Eugene Cheah Aug 31, 2023 ▶ 59:13
Assertion Supported
RWKV Token Inference Only Requires Current and Next States in RAM
“If you really, really want to, like, save RAM, You, it is possible for you to do token by token inference, so that you don't need to keep your states in history. You only need to keep your current token state and your next.”
Eugene Cheah Aug 31, 2023 ▶ 1:05:39
Opinion
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Eugene Cheah Aug 31, 2023 ▶ 1:07:18
Assertion Supported
Microsoft and Other Research Labs Cite RWKV in Architecture Papers
“Ever since that initial paper came out, there was ResNet, there's I think there's two more, there's a few more additional papers coming out, one from Microsoft, one from other organizations that are literally exploring the whole idea, once again, of scalable n…”
Eugene Cheah Aug 31, 2023 ▶ 1:08:56
Assertion Not checkable as stated
Half of the RWKV Community Joined for Non-English Multilingual Support
“The only reason why we have a bigger outsized impact compared to like the other models is frankly because half of our discord Came in not for English. It's for other languages.”
Eugene Cheah Aug 31, 2023 ▶ 1:10:35
Insight
RWKV Performance Degrades on Context Lengths Beyond Its Training Data
“Well, it will, as a neural network, it will happily keep going on for infinite context, man. It will just keep generating. does it do well? That's the answer is no, because if you didn't train it to handle that situation, and that's actually a child rule. So,…”
Eugene Cheah Aug 31, 2023 ▶ 1:13:03
Assertion Not checkable as stated
RWKV Creator Intends to Build the AI Equivalent of Linux Foundation
“He seems to be heavily inspired and wants to go towards the direction of creating the equivalent of a Linux foundation for AI models. So he really wants this to be open source.”
Eugene Cheah Aug 31, 2023 ▶ 1:22:32
Assertion Supported
Training LLMs Beyond Two Epochs Causes Overfitting and Degradation
“Anything beyond that, and we can confirm, even for our model, ours is more like closer to two, but the idea is still there, that it starts to overfit, and it starts to degrade in a lot of things.”
Eugene Cheah Aug 31, 2023 ▶ 1:26:51
Opinion
The Token Shortage Crisis Only Applies to AGI, Not Small Models
“I would say if we are aiming for AGI, there is a token crisis, but if we are aiming for useful small models, I don't think there is a token crisis.”
Eugene Cheah Aug 31, 2023 ▶ 1:27:16
Insight
Cheah: AI Engineers Do Not Need ML Math to Build Products
“Frankly, for an AI engineer, you don't need it. You, your main thing that you needed to do was to, frankly, just play around with ChatGPT, or all the alternatives, be aware of the alternatives, because be very mercenary, swap out to Cloudia if it's better for …”
Eugene Cheah Aug 31, 2023 ▶ 1:33:57
Insight
Cheah: Pre-Transformer Academic Neural Network Research Is No Longer Relevant
“Frankly, almost everything that is, that matters, Ah, was basically in the past four years. Like, there were a lot of things that fit in academics that were before that, and you know, and they were mostly dealing with models that were under a billion parameter…”
Eugene Cheah Aug 31, 2023 ▶ 1:37:51
Insight
ML Researchers Optimize for Working Models Over Code Efficiency
“A lot of ML scientists, when they really build this stuff, the focus was more of like, always get it to work. It was never about getting it to work efficiently, or getting the code documented or organized.”
Eugene Cheah Aug 31, 2023 ▶ 1:42:07
Insight
Cheah: LoRA Cannot Teach LLMs New Languages or Novel Concepts
“Laura cannot teach new language. It cannot it's sometimes may struggle to teach new techniques or new, new concepts. It does well into adding and refining existing contact knowledge.”
Eugene Cheah Aug 31, 2023 ▶ 1:50:20
Opinion
Cheah: A Human Personality and Memories Can Fit on Two SSDs
“No offense to myself, I don't think my personality and my memories is more than this. We could, even if I can exit, I could store this in two SSDs. Two hard drives.”
Eugene Cheah Aug 31, 2023 ▶ 1:54:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.