Mixture of Experts

also referred to as: mixture-of-experts · moe · moes

19 statements across 12 episodes · 7 bullish · 3 bearish · 13 people on the record · first statement Dec 5, 2023 by Dylan Patel · across every show →

Everything said about Mixture of Experts, oldest first

Dec 5, 2023
Assertion Supported
Patel: SemiAnalysis reported GPT-4 mixture of experts architecture in January
“Just being clear, I talked about mixture of experts in January, it's just people didn't really notice it.”
Dylan Patel Dec 5, 2023 ▶ 1:08 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Jun 25, 2024
Insight
Frankle: Training MoE models with FSDP creates severe network bandwidth bottlenecks
“And those models are very demanding when it comes to network bandwidth, at least if you're training them in kind of FSTP zero three style. Where there's just a lot of parameters getting shuffled back and forth and your ratio of kind of compute to amount of dat…”
Jonathan Frankle Jun 25, 2024 ▶ 39:59 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jul 5, 2024 positive
Insight
Tay: Training MoE models from scratch is ideal over sparse upcycling
“I think in the ideal case, you do MOE from scratch.”
Yi Tay Jul 5, 2024 ▶ 1:52:36 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Jul 5, 2024 bullish
Opinion
Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, lik…”
Yi Tay Jul 5, 2024 ▶ 1:51:09 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Jul 23, 2024 positive
Disclosure
Scialom: Meta is exploring Mixture of Experts architectures for future models
“So, it's just an hyperparameter we haven't optimized a lot yet, but we have some stuff ongoing, and that's an hyperparameter we will explore in the future.”
Thomas Scialom Jul 23, 2024 ▶ 27:35 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Oct 29, 2024 neutral
Insight
He: Token dropping works in MoE pre-training; dropless excels in fine-tuning
“A lot of pre-training experiments show that Token dropping is very efficient, and it doesn't impact performance, but in some of the, like, the downstream fine-tuning, people realize drop-less is better.”
Ethan He Oct 29, 2024 ▶ 10:28 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 positive
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Ethan He Oct 29, 2024 ▶ 33:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 positive
Insight
Ethan He: 64 experts is the sweet spot for MoE upcycling
“We found, 64 experts is kind of like the sweet spot. If you increase the number of experts beyond 64, it provides diminishing return.”
Ethan He Oct 29, 2024 ▶ 35:03 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 neutral
Assertion Supported
Ethan He: MoE experts do not cleanly specialize into semantic domains
“Unfortunately, people didn't find, like, a significant interpretability inside these experts. Say, one expert focus on math, the other focus on literature. I think the problem is that neural network hidden states are already very entangled. So, when hidden sta…”
Ethan He Oct 29, 2024 ▶ 16:19 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 positive
Insight
He: Upcycling dense models to MoE beats continuing dense training per FLOP
“By training these upcycled models, you can achieve better accuracy than simply training the dense model further for the same number of flops.”
Ethan He Oct 29, 2024 ▶ 20:32 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Jan 19, 2025 negative
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Oct 20, 2025 neutral
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Shawn Wang Oct 20, 2025 ▶ 36:32 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Oct 20, 2025 negative
Assertion Supported
Bakouch: Megatron's local batching hindered MoE expert specialization
“And they find that, for example, in the Megatron item code base this wasn't the case. This was done as the local batch I think. So it basically means that everyone that is using Megatron at the time of the paper I mean, it wasn't even a paper. It was just a bl…”
Elie Bakouch Oct 20, 2025 ▶ 24:46 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Feb 10, 2026 neutral
Assertion Supported
MoE models with 20B active parameters match 72B dense models in memorization
“With the coming of the MOE models, we're starting to see these phenomena, even like models which have a less number of active parameters. So a phenomena that we would observe at a 72 B model in the past is probably now visible at a 20 B active one 20 B MOE als…”
Pratyush Maini Feb 10, 2026 ▶ 5:24 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Mar 30, 2026 positive
Disclosure
Lample: Mistral develops specialized single-capability models before merging them
“The way we kind of do things internally, that we have like one team, focus on one capability, build one model, and then when it's mature enough, we decide to merge this into the Mixture. So, but hey, here's, it was the first time we basically merged all of thi…”
Guillaume Lample Mar 30, 2026 ▶ 33:02 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
May 24, 2026 neutral
Insight
Sanseviero: Embedding offloading suits edge devices; larger models require MoEs or dense architectures
“This is really optimized and designed for, like, on-device. And when I say on-device, I mean, like, running in a phone, Android, Raspberry Pi, and so on, right? When you go larger, you usually want to come back more You want to have more, like, dense architect…”
Omar Sanseviero May 24, 2026 ▶ 1:19 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
May 24, 2026 neutral
Insight
Sanseviero: MoE models are great for inference but hard to fine-tune
“MOEs are challenging to fine tune. I don't know if we've talked about that in the past, but MOEs in general are like an extremely good architecture. They work great for inference. But when people fine tune them, they struggle a bit. Like they are not as easy t…”
Omar Sanseviero May 24, 2026 ▶ 17:28 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Jun 30, 2026 bearish
Opinion
Edunov: LLM architectures are boring and fundamentally unchanged since 2017
“And honestly, LLM architectures are relatively boring. I don't know, probably alienate half of your audience. But it's like, it's a transforming layer in the end, like paper was published in 2017, and you go to any LLM lab today, you will see very, very simila…”
Sergei Yudinov Jun 30, 2026 ▶ 1:46:23 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Aug 3, 2026
Assertion Contradicted
All modern AI models requiring multi-GPU parallelization are Mixture-of-Experts
“Effectively, all models today are MOE models that are, you know, at least all models large enough that you would care to parallelize them across multiple GPUs.”
Philip Kiely Aug 3, 2026 ▶ 52:42 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.