Mixture of Experts

also referred to as: moe

4 statements across 3 episodes · 2 bullish · 0 bearish · 3 people on the record · first statement Jul 25, 2024 by Sharon Zhou · across every show →

Everything said about Mixture of Experts, oldest first

Jul 25, 2024 bullish
Prediction Not checkable as stated
Combining MoE and LoRA will eliminate big versus small model trade-offs
“And I do think that's the future so we can get something that is incredibly smart, incredibly huge, but with the latency cost and speed of something, something tiny. So no more big model versus small model paradigm. It's potentially one in the same.”
Sharon Zhou Jul 25, 2024 ▶ 32:24 Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini
Apr 2, 2026
Insight
Looping provides parameter-free FLOPS, whereas MoE architectures provide FLOPS-free parameters
“In mixture of experts, you have flops free parameters. So, so parameters that they're not actually bringing any flops. And in, in like looping, you have parameter free flops where you don't have extra parameters for the extra flops that you're throwing on this…”
Mostafa Dehghani Apr 2, 2026 ▶ 39:22 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
Jul 2, 2026 positive
Assertion Not checkable as stated
Catanzaro: MoEs have long been the default architecture in frontier AI
“Yeah, I believe MOEs have been the default in Frontier AI for a long time. They're just a really good combination of inference cost and intelligence.”
Bryan Catanzaro Jul 2, 2026 ▶ 46:54 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Jul 2, 2026 neutral
Insight
Catanzaro: Dense models outperform MoE models under strict memory constraints
“You know, they take a lot more memory. If you have a very small amount of memory, a dense model is going to be smarter.”
Bryan Catanzaro Jul 2, 2026 ▶ 47:06 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.