Prediction Open AI assessment confidence: 80% certainty 3/5 debate potential 2/5

Cheah: AI community will replicate Meta's pipeline scheduling algorithm

Eugene Cheah · [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models · Jul 29, 2024 · at 18:54

Eugene Cheah analyzes Meta's Llama 3.1 paper and its novel pipeline parallelism scheduling designed to minimize idle GPU time.

0:00 / 0:07exact quote · 7.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“This weird scheduling, which I'm quite sure people are going to start replicating it, is to reduce the bubble, the wastage.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Eugene Cheah

Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Insight
Cheah: Non-positional attention architectures remain stable beyond trained context
“One key advantage of this alternate attention mechanic that is not based on token position is that the model don't suddenly become crazy when you go past the eight K training context or a million context. It is actually still stable. It's still, it's able to r…”
Eugene Cheah Dec 24, 2024 ▶ 41:28 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25 RWKV: Reinventing RNNs for the Transformer Era
Insight
Foreign Language Data Degrades English Benchmark Scores on Small LLMs
“Adding in a foreign data set is actually a loss, because once you're below a certain param count, so we're talking about the seven important, right? The more you add that's more in line with your evals, the more it will degrade, and they just exclude it.”
Eugene Cheah Aug 31, 2023 ▶ 25:13 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Matches GPT-NeoX Performance at Equal Parameter and Data Scales
“RWKV is a modern recursive neural network with transformer-like level of LMM performance, which can be trained in a transformer mode. And this part has already been benchmarked against GPT-NeoX in the paper, And it has similar training performance compared to …”
Eugene Cheah Aug 31, 2023 ▶ 31:59 RWKV: Reinventing RNNs for the Transformer Era
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.