Eugene Cheah

CEO & Co-founder, Featherless.ai · 5 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutivescientistengineer@picocreator ↗LinkedIn ↗featherless.ai ↗

Eugene Cheah co-founded and leads Featherless.ai, a serverless inference platform designed to run and serve tens of thousands of open-weight models at scale. He is best known for his research and core contributions to RWKV, an attention-free linear RNN architecture developed under the Linux Foundation.

43statements → 24claims → 17claims resolved → 76%fully supported → 3.81/5average certainty → 2.07/5average debate potential → 9said about them ↓

13 supported 1 partly supported 3 contradicted 1 not yet assessed 6 not checkable as stated how the 24 claims stand · each chip opens the sources

4 predictions · 20 assertions · 3 opinions · 11 insights · 5 disclosures · every statement was checked. The predictions and assertions are the 24 claims: statements the public record can support or contradict. 17 are resolved, 1 is not yet assessed, and 6 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Eugene argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era

Their most notable contradicted claim

Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25 RWKV: Reinventing RNNs for the Transformer Era

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
77% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Eugene Cheah said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Eugene Cheah Dec 24, 2024 ▶ 41:28 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25 RWKV: Reinventing RNNs for the Transformer Era
Insight
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
Eugene Cheah Aug 31, 2023 ▶ 25:13 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
Opinion
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Eugene Cheah Aug 31, 2023 ▶ 1:07:18 RWKV: Reinventing RNNs for the Transformer Era
Disclosure
Cheah: RWKV organization has less compute than a single Google researcher
“So our entire organization has less compute than a single researcher in Google.”
Eugene Cheah Dec 24, 2024 ▶ 24:32 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Assertion Contradicted
Cheah: Llama 3.1 405B is first frontier model using pipeline parallelism
“This is the first major model that of this cell class size, right? They're saying, hey, we are doing pipeline parallelism.”
Eugene Cheah Jul 29, 2024 ▶ 19:21 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Disclosure
RWKV Raven Dataset Scrubs Out 'As an AI' Refusal Boilerplate
“Typically GPT for all, but then we scrub it for and remove all the, as a large model.”
Eugene Cheah Aug 31, 2023 ▶ 39:46 RWKV: Reinventing RNNs for the Transformer Era
Disclosure
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
Eugene Cheah Aug 31, 2023 ▶ 43:11 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Uses Trie Tokenizer Without Space Delimiters for CJK Languages
“Instead of using like this token pairs well with this and should be paired with that we just made it a trial list. So So basically, try the data structure. Yeah. So we just find the longest matching string in that matching string that we have trained inside ou…”
Eugene Cheah Aug 31, 2023 ▶ 48:13 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
Eugene Cheah Aug 31, 2023 ▶ 58:45 RWKV: Reinventing RNNs for the Transformer Era
Opinion
The Token Shortage Crisis Only Applies to AGI, Not Small Models
“I would say if we are aiming for AGI, there is a token crisis, but if we are aiming for useful small models, I don't think there is a token crisis.”
Eugene Cheah Aug 31, 2023 ▶ 1:27:16 RWKV: Reinventing RNNs for the Transformer Era
Insight
Cheah: AI Engineers Do Not Need ML Math to Build Products
“Frankly, for an AI engineer, you don't need it. You, your main thing that you needed to do was to, frankly, just play around with ChatGPT, or all the alternatives, be aware of the alternatives, because be very mercenary, swap out to Cloudia if it's better for …”
Eugene Cheah Aug 31, 2023 ▶ 1:33:57 RWKV: Reinventing RNNs for the Transformer Era
Insight
Cheah: Pre-Transformer Academic Neural Network Research Is No Longer Relevant
“Frankly, almost everything that is, that matters, Ah, was basically in the past four years. Like, there were a lot of things that fit in academics that were before that, and you know, and they were mostly dealing with models that were under a billion parameter…”
Eugene Cheah Aug 31, 2023 ▶ 1:37:51 RWKV: Reinventing RNNs for the Transformer Era
Opinion
Cheah: A Human Personality and Memories Can Fit on Two SSDs
“No offense to myself, I don't think my personality and my memories is more than this. We could, even if I can exit, I could store this in two SSDs. Two hard drives.”
Eugene Cheah Aug 31, 2023 ▶ 1:54:37 RWKV: Reinventing RNNs for the Transformer Era
Insight
Cheah: RWKV Innovations Are Found Empirically Before Academic Rationalization
“Officially in the paper, I'll say we had this idea and we wrote it this way. The reality is someone came in the code, we tested it worked, and then we rationalized it.”
Eugene Cheah Dec 24, 2024 ▶ 22:46 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Assertion Not checkable as stated
Cheah: Most enterprise AI workloads use 70B models under 32k context
“Majority of enterprise workload today is just on Senti B at under 32 K context line.”
Eugene Cheah Dec 24, 2024 ▶ 27:23 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Assertion Supported
Eugene Chia: Inverting numbers in reasoning traces improves math model performance
“The crazy one, the crazy thing that we did was that we inverted the numbers during the calculation and it seems to work better.”
Eugene Cheah Sep 18, 2024 ▶ 23:13 [Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
Prediction Didn’t hold up
Cheah: Cloud providers will slash model inference prices before raising them
“One thing to warn about pricing is that you're going to see a lot of providers jumping in, and everyone's just trying to get the piece of the pie. So, so, so like with some of the previous model launches, you see some people coming in at lower and lower price,…”
Eugene Cheah Jul 29, 2024 ▶ 15:06 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Prediction Open · timeframe Jul 2027
Cheah: AI community will replicate Meta's pipeline scheduling algorithm
“This weird scheduling, which I'm quite sure people are going to start replicating it, is to reduce the bubble, the wastage.”
Eugene Cheah Jul 29, 2024 ▶ 18:54 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Prediction Not checkable as stated
Eugene Cheah: LoRAs will arrive before full Llama 3.1 405B fine-tunes
“I suspect we are going to see more LoRa's first before we get full fine-tuned.”
Eugene Cheah Jul 29, 2024 ▶ 47:54 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Insight
Cheah: Over-quantized LLMs rapidly enter repetition loops at long contexts
“When you over-quantize, right, at longer context length, right, it starts going into repetition rapidly.”
Eugene Cheah Jul 29, 2024 ▶ 1:03:20 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models

Show 19statements(19 left)

The other half of the tape: Eugene Cheah's own voice is left out of every number here. Other people bring the name up 9 times in 4 episodes on Latent Space. every mention, with the transcript →

Who brings them up most Sarah Chieng 5Jesse Hu 1Alessio Fanelli 1

Every mention by year

tap a year for its mentions
00428320242025episodesmentions
02320242025episodes it came up in
001.51.53320242025episodesmentions per episode

Appearances (5)

EpisodeDateSpeaking time
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ Neu Dec 24, 2024 14m
[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shre Nov 29, 2024 2m
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends! Sep 18, 2024 1m
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 11m
RWKV: Reinventing RNNs for the Transformer Era Aug 31, 2023 1h 15m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.