Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
Yi Tay · The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka · Jul 5, 2024 · at 1:51:09
Yi Tay (co-founder of Reka and former Google Brain researcher) explains why Mixture-of-Experts (MoE) models offer superior compute-efficiency trade-offs over dense models.
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, like, that's the performance per flop argument, like, activated primes, whatever. That's, like, kind of, like, that's a way to slightly cheat the scaling law a little bit, right? By having more parameters, right?”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →