Oct 29, 2024 · 39m · latent-space

[Paper Club] Upcycling Large Language Models into Mixture of Experts

Ethan He · 31m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Ethan from NVIDIA explains the principles of Mixture of Experts (MoE) architectures, details Megatron Core systems optimizations, and presents a methodology for upcycling pre-trained dense language models into high-capacity MoE systems with minimal compute cost.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.5 Guest teaching 1.0 Guest disagreement 0.3 The hosts pushing back 0.3
05100:0010:0020:0030:006:58–14:00 · The hosts as informed peer 0/10 Megatron Core MoE Optimization and Efficient Execution Strategies This is an uninterrupted presentation monologue by Ethan He describing Megatron Core MoE optimization techniques, token dropping, and grouped GEM operations. Because the host is not participating in this segment, host-side metrics are set to zero.14:08–20:01 · The hosts as informed peer 2/10 Audience Q&A on MoE Usability, Specialization, and Routing Paradigms The host and audience members ask practical and conceptual questions regarding standalone modularity, domain specialization in experts, and routing variations. Ethan methodically educates the room on representation superposition and expert choice mechanics.20:07–32:00 · The hosts as informed peer 0/10 Methodology for Upcycling Dense Models into Mixture of Experts Ethan delivers a technical monologue detailing the methodology behind upcycling dense LLMs into MoEs, explaining the softmax/top-k order swap and router weight duplication for fine-grained routing. The host does not speak during this segment.32:02–39:01 · The hosts as informed peer 0/10 Experimental Results, Hyperparameter Optimization, and Scaling Law Analysis Ethan concludes the presentation by reviewing experimental results, hyperparameter selection (peak learning rate), and scaling law conversions. The host steps in at the very end purely to conclude the recorded presentation and transition to private Q&A.6:58–14:00 · Guest teaching 0/10 Megatron Core MoE Optimization and Efficient Execution Strategies This is an uninterrupted presentation monologue by Ethan He describing Megatron Core MoE optimization techniques, token dropping, and grouped GEM operations. Because the host is not participating in this segment, host-side metrics are set to zero.14:08–20:01 · Guest teaching 4/10 Audience Q&A on MoE Usability, Specialization, and Routing Paradigms The host and audience members ask practical and conceptual questions regarding standalone modularity, domain specialization in experts, and routing variations. Ethan methodically educates the room on representation superposition and expert choice mechanics.20:07–32:00 · Guest teaching 0/10 Methodology for Upcycling Dense Models into Mixture of Experts Ethan delivers a technical monologue detailing the methodology behind upcycling dense LLMs into MoEs, explaining the softmax/top-k order swap and router weight duplication for fine-grained routing. The host does not speak during this segment.32:02–39:01 · Guest teaching 0/10 Experimental Results, Hyperparameter Optimization, and Scaling Law Analysis Ethan concludes the presentation by reviewing experimental results, hyperparameter selection (peak learning rate), and scaling law conversions. The host steps in at the very end purely to conclude the recorded presentation and transition to private Q&A.6:58–14:00 · Guest disagreement 0/10 Megatron Core MoE Optimization and Efficient Execution Strategies This is an uninterrupted presentation monologue by Ethan He describing Megatron Core MoE optimization techniques, token dropping, and grouped GEM operations. Because the host is not participating in this segment, host-side metrics are set to zero.14:08–20:01 · Guest disagreement 1/10 Audience Q&A on MoE Usability, Specialization, and Routing Paradigms The host and audience members ask practical and conceptual questions regarding standalone modularity, domain specialization in experts, and routing variations. Ethan methodically educates the room on representation superposition and expert choice mechanics.20:07–32:00 · Guest disagreement 0/10 Methodology for Upcycling Dense Models into Mixture of Experts Ethan delivers a technical monologue detailing the methodology behind upcycling dense LLMs into MoEs, explaining the softmax/top-k order swap and router weight duplication for fine-grained routing. The host does not speak during this segment.32:02–39:01 · Guest disagreement 0/10 Experimental Results, Hyperparameter Optimization, and Scaling Law Analysis Ethan concludes the presentation by reviewing experimental results, hyperparameter selection (peak learning rate), and scaling law conversions. The host steps in at the very end purely to conclude the recorded presentation and transition to private Q&A.6:58–14:00 · The hosts pushing back 0/10 Megatron Core MoE Optimization and Efficient Execution Strategies This is an uninterrupted presentation monologue by Ethan He describing Megatron Core MoE optimization techniques, token dropping, and grouped GEM operations. Because the host is not participating in this segment, host-side metrics are set to zero.14:08–20:01 · The hosts pushing back 1/10 Audience Q&A on MoE Usability, Specialization, and Routing Paradigms The host and audience members ask practical and conceptual questions regarding standalone modularity, domain specialization in experts, and routing variations. Ethan methodically educates the room on representation superposition and expert choice mechanics.20:07–32:00 · The hosts pushing back 0/10 Methodology for Upcycling Dense Models into Mixture of Experts Ethan delivers a technical monologue detailing the methodology behind upcycling dense LLMs into MoEs, explaining the softmax/top-k order swap and router weight duplication for fine-grained routing. The host does not speak during this segment.32:02–39:01 · The hosts pushing back 0/10 Experimental Results, Hyperparameter Optimization, and Scaling Law Analysis Ethan concludes the presentation by reviewing experimental results, hyperparameter selection (peak learning rate), and scaling law conversions. The host steps in at the very end purely to conclude the recorded presentation and transition to private Q&A.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 18:08 Clarifying dynamic top-p optimization drawbacks

In an otherwise non-combative session, Ethan mildly pushes back on the practical viability of top-p routing by pointing out the optimization difficulties caused by dynamic expert load sizing.

Hardest push from the hosts ▶ 14:08 Host presses on standalone framework accessibility

The host directly probes whether Megatron Core MoE modules can be used universally like FlashAttention or if users are locked into the entire Megatron framework ecosystem.

Biggest teaching moment ▶ 16:09 Disabusing human-interpretable expert domain specialization

Ethan disabuses the audience's assumption that individual experts specialize in clean human subjects (like physics or economics), explaining how neural network hidden states exist in entangled superposition.

The host holds their own ▶ 14:08 Host connects Megatron tooling to FlashAttention developer ecosystem

The host demonstrates domain awareness by comparing Megatron Core's release strategy to FlashAttention's modular integration model across ML frameworks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Megatron Core MoE Optimization and Efficient Execution Strategies 0000 This is an uninterrupted presentation monologue by Ethan He describing Megatron Core MoE optimization techniques, token dropping, and grouped GEM operations. Because the host is not participating in this segment, host-side metrics are set to zero.
Audience Q&A on MoE Usability, Specialization, and Routing Paradigms 2411 The host and audience members ask practical and conceptual questions regarding standalone modularity, domain specialization in experts, and routing variations. Ethan methodically educates the room on representation superposition and expert choice mechanics.
Methodology for Upcycling Dense Models into Mixture of Experts 0000 Ethan delivers a technical monologue detailing the methodology behind upcycling dense LLMs into MoEs, explaining the softmax/top-k order swap and router weight duplication for fine-grained routing. The host does not speak during this segment.
Experimental Results, Hyperparameter Optimization, and Scaling Law Analysis 0000 Ethan concludes the presentation by reviewing experimental results, hyperparameter selection (peak learning rate), and scaling law conversions. The host steps in at the very end purely to conclude the recorded presentation and transition to private Q&A.

Statements from this episode (8)

Insight
He: Token dropping works in MoE pre-training; dropless excels in fine-tuning
“A lot of pre-training experiments show that Token dropping is very efficient, and it doesn't impact performance, but in some of the, like, the downstream fine-tuning, people realize drop-less is better.”
Ethan He Oct 29, 2024 ▶ 10:28
Assertion Supported
Ethan He: Hugging Face's sequential GEMM loop for Mixtral is inefficient
“Let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer. You will soon notice the, in the expert operation there, You would iterate over all of the experts and compute each of the gem operations one by one. We found that this i…”
Ethan He Oct 29, 2024 ▶ 12:37
Assertion Supported
Ethan He: MoE experts do not cleanly specialize into semantic domains
“Unfortunately, people didn't find, like, a significant interpretability inside these experts. Say, one expert focus on math, the other focus on literature. I think the problem is that neural network hidden states are already very entangled. So, when hidden sta…”
Ethan He Oct 29, 2024 ▶ 16:19
Insight
He: Upcycling dense models to MoE beats continuing dense training per FLOP
“By training these upcycled models, you can achieve better accuracy than simply training the dense model further for the same number of flops.”
Ethan He Oct 29, 2024 ▶ 20:32
Assertion Supported
He: Upcycling a 15B model on 1T tokens yielded 4% MMLU gain
“On other scaling experiments, we tried on 15 B models upcycling and applied on one trillion tokens and achieved roughly about five percent improvement in terms of the validation loss and four percent improvement on MMLU.”
Ethan He Oct 29, 2024 ▶ 21:12
Insight
He: Mixtral's top-k before softmax routing hurts MoE upcycling performance
“We actually found the mix-throughs approach didn't work as well as expected, because the original model, the original switch transformer from Google uses a softmax and topk for a reason. And because of upcycling, if you switch to topk, then softmax, it actuall…”
Ethan He Oct 29, 2024 ▶ 24:34
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Ethan He Oct 29, 2024 ▶ 33:34
Insight
Ethan He: 64 experts is the sweet spot for MoE upcycling
“We found, 64 experts is kind of like the sweet spot. If you increase the number of experts beyond 64, it provides diminishing return.”
Ethan He Oct 29, 2024 ▶ 35:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.