Sparse Mixture Of Experts

topic on 2 shows · 2 statements across 2 episodes

the a16z Podcast Big Technology

2 statements about Sparse Mixture Of Experts, every show

BIG TECHNOLOGY Assertion Contradicted
Mensch: DeepSeek built on Mistral's open-source MoE architecture
“Like we released the first Sparse mixture of experts back at the beginning of 2024, and they built on top, and they released deep seek free, and deep seek was built on top of that. Well, it was, it's the same architecture, and we released, like, everything tha…”
Arthur Mensch Jan 16, 2026 ▶ 36:59 Who Wins if AI Models Commoditize? — With Mistral CEO Arthur Mensch
a16z Insight
Mensch: Mixture of experts decouples model capacity from inference cost
“A sparse mixture of experts, you take the dense layer and you duplicate it several times. And so that's where you actually increase the number of parameters. So you increase the capacity of the model without increasing the cost. So that's the way of decoupling…”
Arthur Mensch Dec 28, 2023 ▶ 10:59 Safety in Numbers: Keeping AI Open

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.