RWKV

product on 3 shows · 24 statements across 5 episodes · said 97 times in 9 episodes since 2023

Latent Space 90 Top Founders 6 No Priors 1

Mentions by year, every show

tap a year for its mentions
004028032023202420252026episodesmentions
0232023202420252026episodes it came up in
00131.52532023202420252026episodesmentions per episode

Latent Space 90Top Founders 6No Priors 1

2026 6 mentions in 1 episode
2025 5 mentions in 2 episodes 3 per episode
2024 24 mentions in 3 episodes 8 per episode
2023 62 mentions in 3 episodes 21 per episode

every mention on every show, scene by scene, with the transcript →

24 statements about RWKV, every show

TOP FOUNDERS Assertion Not checkable as stated
Cheah: RWKV Architecture Could Reduce Inference Costs by 1,000x
“Like this new AI architecture has the potential of reducing inference costs by over a thousand X.”
Eugene Cheah Jul 1, 2026 ▶ 5:58 Featherless: $3.6M Revenue Running 6,700 Open Source AI Models — Eugene Cheah
LATENT SPACE Assertion Not checkable as stated
Swyx: Most AI researchers are scaling Transformers, not alternative architectures
“I think most people that I talk to are not seriously pursuing alternative architectures. There are some notable exceptions, primarily together with the Mombard architecture recursal with RWKV. There's like the XLSTM that was created by a separate Hogwriter. An…”
Shawn Wang Aug 6, 2025 ▶ 1:10:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
LATENT SPACE Disclosure
Cheah: RWKV Trains Models on Over 100 Languages Targeting 200
“We actually train our models primarily on over a hundred language, which is another topic altogether and our goal is to train to even 200 languages to cover all languages in the world.”
Eugene Cheah Dec 24, 2024 ▶ 19:51 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Cheah: RWKV Innovations Are Found Empirically Before Academic Rationalization
“Officially in the paper, I'll say we had this idea and we wrote it this way. The reality is someone came in the code, we tested it worked, and then we rationalized it.”
Eugene Cheah Dec 24, 2024 ▶ 22:46 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Cheah: RWKV operates with a 40-megabyte fixed state size
“Like, we, like, RWKV is running at 40 megabytes for its state.”
Eugene Cheah Dec 24, 2024 ▶ 34:34 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Eugene Cheah Dec 24, 2024 ▶ 41:28 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
LATENT SPACE Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Disclosure
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
Eugene Cheah Aug 31, 2023 ▶ 43:11 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Uses Trie Tokenizer Without Space Delimiters for CJK Languages
“Instead of using like this token pairs well with this and should be paired with that we just made it a trial list. So So basically, try the data structure. Yeah. So we just find the longest matching string in that matching string that we have trained inside ou…”
Eugene Cheah Aug 31, 2023 ▶ 48:13 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Community Built Multimodal Vision Models Using MiniGPT-4 Architecture
“There is actually another project internally on the discord where it's Doing vision, ah, vision modeling, and this based on the, ah, is it, mini GPT-Fall paper, where, where you have an image model, put everything inside the latent space, and then you have the…”
Eugene Cheah Aug 31, 2023 ▶ 50:49 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
Bo Peng Created RWKV on EleutherAI Forum to Parallelize RNNs
“Blink, or Bopeng is the actual name decided basically as an individual, literally at the illiterate AI forum, decided that, hey I think we can modify recurrent neural networks, no, neural networks, based on the Apple paper, the light attention that I showed pr…”
Eugene Cheah Aug 31, 2023 ▶ 53:46 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
EleutherAI and Stability AI Donated A100 Compute to Train RWKV
“So, so that's why AI and the rest stability, I believe also is involved, stepped up. And donated the A-One-Hundreds needed to train the basic models that RWKB had.”
Eugene Cheah Aug 31, 2023 ▶ 55:25 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
Eugene Cheah Aug 31, 2023 ▶ 58:45 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Inference Computes at O(1) Complexity Per Token
“I'm talking about, like, to go through the entire context, yeah, this will be O one per token.”
Eugene Cheah Aug 31, 2023 ▶ 59:13 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
RWKV Token Inference Only Requires Current and Next States in RAM
“If you really, really want to, like, save RAM, You, it is possible for you to do token by token inference, so that you don't need to keep your states in history. You only need to keep your current token state and your next.”
Eugene Cheah Aug 31, 2023 ▶ 1:05:39 RWKV: Reinventing RNNs for the Transformer Era
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Eugene Cheah Aug 31, 2023 ▶ 1:07:18 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
Microsoft and Other Research Labs Cite RWKV in Architecture Papers
“Ever since that initial paper came out, there was ResNet, there's I think there's two more, there's a few more additional papers coming out, one from Microsoft, one from other organizations that are literally exploring the whole idea, once again, of scalable n…”
Eugene Cheah Aug 31, 2023 ▶ 1:08:56 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Not checkable as stated
Half of the RWKV Community Joined for Non-English Multilingual Support
“The only reason why we have a bigger outsized impact compared to like the other models is frankly because half of our discord Came in not for English. It's for other languages.”
Eugene Cheah Aug 31, 2023 ▶ 1:10:35 RWKV: Reinventing RNNs for the Transformer Era
RWKV Performance Degrades on Context Lengths Beyond Its Training Data
“Well, it will, as a neural network, it will happily keep going on for infinite context, man. It will just keep generating. does it do well? That's the answer is no, because if you didn't train it to handle that situation, and that's actually a child rule. So,…”
Eugene Cheah Aug 31, 2023 ▶ 1:13:03 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Not checkable as stated
RWKV Creator Intends to Build the AI Equivalent of Linux Foundation
“He seems to be heavily inspired and wants to go towards the direction of creating the equivalent of a Linux foundation for AI models. So he really wants this to be open source.”
Eugene Cheah Aug 31, 2023 ▶ 1:22:32 RWKV: Reinventing RNNs for the Transformer Era
LATENT SPACE Assertion Supported
Training LLMs Beyond Two Epochs Causes Overfitting and Degradation
“Anything beyond that, and we can confirm, even for our model, ours is more like closer to two, but the idea is still there, that it starts to overfit, and it starts to degrade in a lot of things.”
Eugene Cheah Aug 31, 2023 ▶ 1:26:51 RWKV: Reinventing RNNs for the Transformer Era
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri Dao Aug 3, 2023 ▶ 46:51 FlashAttention-2: Making Transformers 800% faster AND exact

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.