Sparse Mixture Of Experts
topic on 2 shows · 2 statements across 2 episodes
the a16z Podcast
Big Technology
2 statements about Sparse Mixture Of Experts, every show
Mensch: DeepSeek built on Mistral's open-source MoE architecture
“Like we released the first
Sparse mixture of experts back at the beginning of 2024, and they built on top, and they released deep seek free, and deep seek was built on top of that. Well, it was, it's the same architecture, and we released, like, everything tha…”
Mensch: Mixture of experts decouples model capacity from inference cost
“A sparse mixture of experts, you take the dense layer and you duplicate it several times. And so that's where you actually increase the number of parameters. So you increase the capacity of the model without increasing the cost. So that's the way of decoupling…”