Mixtral 8x7B, every mention
13 scenes, the whole family · ← back to Mixtral 8x7B
tap a year for its mentions
every year anyone Ethan He 8Nathan Lambert 3George Cameron 1
Verbatim, from the transcripts: the passages where Mixtral 8x7B comes up
Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
- ▶ 36:49 unnamed speaker So, uh, Dan, 20, 24, this is the Onix trial of experts. 4 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 6:32 George Cameron Uh, we had Mixtrel A times seven B
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 13:22 unnamed speaker Basically, like, like this time last year, Mick Strao was sort of kicking off a bit of an MOE trend with, uh, you know, eight by seven B, eight by 22 B.
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 27:28 unnamed speaker And in December, 23, we, we released another popular, uh, model with the MLE architecture, um, Mr.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
- ▶ 5:46 Ethan He You really, for example, mixture of a by seven B, you increase the model parameters roughly by, like, six X or seven X.
- ▶ 12:37 Ethan He Uh, let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer.
- ▶ 23:18 Ethan He This is introducing mixed row A by seven B. 4 times in the scene
- ▶ 34:35 Ethan He Also, we kind of, we also analyzed the Mixtro A by seven base versus Mixtro seven B. 2 times in the scene
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 18:17 unnamed speaker Uh, by the way, they've also deprecated, uh, Mistral, seven B, eight by seven B and eight by 22 B, right?
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 22:01 unnamed speaker Um, but basically the, the mixed raw release, the MOE model was kind of the, the spark of, of the war. 7 times in the scene
- ▶ 31:23 unnamed speaker Um, one thing I will mention on, like, the engineering, sort of technical detail side is, um, you know, the, the, the rise of Mixture of Experts is something that, you know, was, uh, we covered in our podcast with George, and, and, uh, now… 2 times in the scene
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 1:25:33 Nathan Lambert So in the open source, there's things like, um, like, Mixtral instruct to the two seven DB, which is effectively, it's a way bigger model than Mixtral. 2 times in the scene
- ▶ 1:25:33 Nathan Lambert So in the open source, there's things like, um, like, Mixtral instruct to the two seven DB, which is effectively, it's a way bigger model than Mixtral.