Mixtral

includes Mixtral 8x7B, Mixtral 8x22B

2 statements across 2 episodes · 0 bullish · 1 bearish · 1 people on the record · first statement Oct 29, 2024 by Ethan He · said 46 times in 15 episodes since 2024 · across every show →

Mentions by year, the whole family

brought up most by Ethan He (9), Nathan Lambert (3), Yi Tay (1), Vipul Ved Prakash (1), Shawn Wang (1), Pranav Reddy (1), George Cameron (1)

tap a year for its mentions
002054010202420252026episodesmentions
0510202420252026episodes it came up in
0025410202420252026episodesmentions per episode
2026 7 mentions in 3 episodes 2 per episode
2025 6 mentions in 3 episodes 2 per episode
2024 33 mentions in 9 episodes 4 per episode

every mention, scene by scene, with the transcript →

Everything said about Mixtral, oldest first

Oct 29, 2024 negative
Insight
He: Mixtral's top-k before softmax routing hurts MoE upcycling performance
“We actually found the mix-throughs approach didn't work as well as expected, because the original model, the original switch transformer from Google uses a softmax and topk for a reason. And because of upcycling, if you switch to topk, then softmax, it actuall…”
Ethan He Oct 29, 2024 ▶ 24:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Jun 1, 2026 neutral
Disclosure
Ethan He: NVIDIA Cosmos uses a 7B video model with a larger LLM rewriter
“I think in in Cosmos, we use Lama or we use mix, mix through. And the Cosmos video model itself is only seven B, and the model, the language model is a prompt rewriter. It's bigger than that.”
Ethan He Jun 1, 2026 ▶ 1:15:37 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.