RWKV

includes RWKV World, RWKV Raven, RWKV World Model

23 statements across 4 episodes · 15 bullish · 1 bearish · 3 people on the record · first statement Aug 3, 2023 by Tri Dao · said 109 times in 7 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Eugene Cheah (56), Shawn Wang (4), Tri Dao (2), Dan Fu (1), Alessio Fanelli (1)

tap a year for its mentions
00402803202320242025episodesmentions
023202320242025episodes it came up in
00201.5403202320242025episodesmentions per episode
2025 5 mentions in 2 episodes 3 per episode
2024 24 mentions in 3 episodes 8 per episode
2023 80 mentions in 2 episodes 40 per episode

every mention, scene by scene, with the transcript →

Everything said about RWKV, oldest first

Aug 3, 2023 positive
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri Dao Aug 3, 2023 ▶ 46:51 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 31, 2023 neutral
Assertion Supported
Bo Peng Created RWKV on EleutherAI Forum to Parallelize RNNs
“Blink, or Bopeng is the actual name decided basically as an individual, literally at the illiterate AI forum, decided that, hey I think we can modify recurrent neural networks, no, neural networks, based on the Apple paper, the light attention that I showed pr…”
Eugene Cheah Aug 31, 2023 ▶ 53:46 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Not checkable as stated
Half of the RWKV Community Joined for Non-English Multilingual Support
“The only reason why we have a bigger outsized impact compared to like the other models is frankly because half of our discord Came in not for English. It's for other languages.”
Eugene Cheah Aug 31, 2023 ▶ 1:10:35 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Supported
RWKV Inference Computes at O(1) Complexity Per Token
“I'm talking about, like, to go through the entire context, yeah, this will be O one per token.”
Eugene Cheah Aug 31, 2023 ▶ 59:13 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Disclosure
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
Eugene Cheah Aug 31, 2023 ▶ 43:11 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Assertion Supported
Microsoft and Other Research Labs Cite RWKV in Architecture Papers
“Ever since that initial paper came out, there was ResNet, there's I think there's two more, there's a few more additional papers coming out, one from Microsoft, one from other organizations that are literally exploring the whole idea, once again, of scalable n…”
Eugene Cheah Aug 31, 2023 ▶ 1:08:56 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 neutral
Insight
RWKV Performance Degrades on Context Lengths Beyond Its Training Data
“Well, it will, as a neural network, it will happily keep going on for infinite context, man. It will just keep generating. does it do well? That's the answer is no, because if you didn't train it to handle that situation, and that's actually a child rule. So,…”
Eugene Cheah Aug 31, 2023 ▶ 1:13:03 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Assertion Supported
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
Eugene Cheah Aug 31, 2023 ▶ 58:45 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Supported
EleutherAI and Stability AI Donated A100 Compute to Train RWKV
“So, so that's why AI and the rest stability, I believe also is involved, stepped up. And donated the A-One-Hundreds needed to train the basic models that RWKB had.”
Eugene Cheah Aug 31, 2023 ▶ 55:25 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Supported
RWKV Uses Trie Tokenizer Without Space Delimiters for CJK Languages
“Instead of using like this token pairs well with this and should be paired with that we just made it a trial list. So So basically, try the data structure. Yeah. So we just find the longest matching string in that matching string that we have trained inside ou…”
Eugene Cheah Aug 31, 2023 ▶ 48:13 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023
Assertion Supported
RWKV Token Inference Only Requires Current and Next States in RAM
“If you really, really want to, like, save RAM, You, it is possible for you to do token by token inference, so that you don't need to keep your states in history. You only need to keep your current token state and your next.”
Eugene Cheah Aug 31, 2023 ▶ 1:05:39 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Opinion
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Eugene Cheah Aug 31, 2023 ▶ 1:07:18 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 negative
Assertion Supported
Training LLMs Beyond Two Epochs Causes Overfitting and Degradation
“Anything beyond that, and we can confirm, even for our model, ours is more like closer to two, but the idea is still there, that it starts to overfit, and it starts to degrade in a lot of things.”
Eugene Cheah Aug 31, 2023 ▶ 1:26:51 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Supported
RWKV Community Built Multimodal Vision Models Using MiniGPT-4 Architecture
“There is actually another project internally on the discord where it's Doing vision, ah, vision modeling, and this based on the, ah, is it, mini GPT-Fall paper, where, where you have an image model, put everything inside the latent space, and then you have the…”
Eugene Cheah Aug 31, 2023 ▶ 50:49 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Assertion Not checkable as stated
RWKV Creator Intends to Build the AI Equivalent of Linux Foundation
“He seems to be heavily inspired and wants to go towards the direction of creating the equivalent of a Linux foundation for AI models. So he really wants this to be open source.”
Eugene Cheah Aug 31, 2023 ▶ 1:22:32 RWKV: Reinventing RNNs for the Transformer Era
Dec 24, 2024 neutral
Assertion Supported
Cheah: RWKV operates with a 40-megabyte fixed state size
“Like, we, like, RWKV is running at 40 megabytes for its state.”
Eugene Cheah Dec 24, 2024 ▶ 34:34 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Dec 24, 2024
Insight
Cheah: RWKV Innovations Are Found Empirically Before Academic Rationalization
“Officially in the paper, I'll say we had this idea and we wrote it this way. The reality is someone came in the code, we tested it worked, and then we rationalized it.”
Eugene Cheah Dec 24, 2024 ▶ 22:46 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Dec 24, 2024 positive
Insight
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Eugene Cheah Dec 24, 2024 ▶ 41:28 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Dec 24, 2024
Disclosure
Cheah: RWKV Trains Models on Over 100 Languages Targeting 200
“We actually train our models primarily on over a hundred language, which is another topic altogether and our goal is to train to even 200 languages to cover all languages in the world.”
Eugene Cheah Dec 24, 2024 ▶ 19:51 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Dec 24, 2024 positive
Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Aug 6, 2025 neutral
Assertion Not checkable as stated
Swyx: Most AI researchers are scaling Transformers, not alternative architectures
“I think most people that I talk to are not seriously pursuing alternative architectures. There are some notable exceptions, primarily together with the Mombard architecture recursal with RWKV. There's like the XLSTM that was created by a separate Hogwriter. An…”
Shawn Wang Aug 6, 2025 ▶ 1:10:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.