Everything Eugene Cheah said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Cheah: Open Models Now Match Claude Sonnet and GPT-4o Mini
“So, and the, this growing collection of open models includes some of the best models that are already on par or surpass, let's say, Plot Sonnet or even GPT-A for Mini.”
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer.
How I do so, we'll cover later.
And this can be scaled to as many parameters as we want.”
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Cheah: RWKV Architecture Could Reduce Inference Costs by 1,000x
“Like this new AI architecture has the potential of reducing inference costs by over a thousand X.”
Cheah: Global AI Market Will Segment into Domestic Sovereign Models
“So they are going to the direction of highly tailored sovereign AI models for the domestic market. And we actually see this happening more and more. So for Cohear, they will service the Canadian market. For the US market is going to be served by OpenAI Entropi…”
Cheah: RWKV organization has less compute than a single Google researcher
“So our entire organization has less compute than a single researcher in Google.”
Cheah: Llama 3.1 405B is first frontier model using pipeline parallelism
“This is the first major model that of this cell class size, right? They're saying, hey, we are doing pipeline parallelism.”
RWKV Raven Dataset Scrubs Out 'As an AI' Refusal Boilerplate
“Typically GPT for all, but then we scrub it for and remove all the, as a large model.”
RWKV Prioritizes User Feedback Over Benchmark Evals for Dataset Additions
“The reason why we add things to the data set was never about improving evals. It's about directly in response to user feedback.”
RWKV Uses Trie Tokenizer Without Space Delimiters for CJK Languages
“Instead of using like this token pairs well with this and should be paired with that we just made it a trial list. So So basically, try the data structure. Yeah. So we just find the longest matching string in that matching string that we have trained inside ou…”
RWKV Trains in Parallel Across GPUs, Unlike Traditional RNNs
“And in practice, once you start cascading there, you just saturate the GPU, and that's how it starts being paralysable trained. You no longer need to train in slices like traditional RNNs.”
The Token Shortage Crisis Only Applies to AGI, Not Small Models
“I would say if we are aiming for AGI, there is a token crisis, but if we are aiming for useful small models, I don't think there is a token crisis.”
Cheah: AI Engineers Do Not Need ML Math to Build Products
“Frankly, for an AI engineer, you don't need it. You, your main thing that you needed to do was to, frankly, just play around with ChatGPT, or all the alternatives, be aware of the alternatives, because be very mercenary, swap out to Cloudia if it's better for …”
Cheah: Pre-Transformer Academic Neural Network Research Is No Longer Relevant
“Frankly, almost everything that is, that matters, Ah, was basically in the past four years. Like, there were a lot of things that fit in academics that were before that, and you know, and they were mostly dealing with models that were under a billion parameter…”
Cheah: A Human Personality and Memories Can Fit on Two SSDs
“No offense to myself, I don't think my personality and my memories is more than this. We could, even if I can exit, I could store this in two SSDs. Two hard drives.”
Cheah: Long-Tail Fine-Tuned Models Drive 50% of Featherless Workload
“You see, most providers, they only provide, let's say, less than a hundred models. That covers 50% of our inference work. It's the bottom 50% where they run all these interesting fine-tuned models that people came on board for.”
Featherless Targets Startups Burning $100K Monthly on OpenAI and Anthropic
“We also realized that there is a lot of money on the table right now where you can go after the startups that, hey, I just built my entire startup or SMB on OpenAI or Entropic, and I'm burning a 100,000 dollars a month. And I do not know what I was doing. And …”
Featherless Won Multiple Contracts by Exclusively Hosting the StepFun Model
“The step one model is a particularly popular model for us that easily ship several contracts for us on this model alone. And no one else is.”
Cheah: Featherless AI's Largest Customer Pays $1M to $2M Annually
“So currently the biggest will be around one to two million dollars a year, which may sound extremely large, but when you actually peel behind the layers, it only comes to around like five, six of the largest servers you see in the market.”