Assertion Partly supported AI assessment confidence: 88% certainty 4/5 debate potential 1/5

Reddy: Voxtral TTS is a 3B model based on the Ministral architecture

Pavan Kumar Reddy · Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample · Mar 30, 2026 · at 2:53

Pavan Kumar Reddy, Audio Research Lead at Mistral AI, describes the architecture and size of Mistral's open text-to-speech model, Voxtral TTS.

0:00 / 0:18exact quote · 18.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“It's it came out with such good quality, and Guillaume was mentioning, yeah, it's a three B model it's based off of the ministral model that we actually released just a few months back, and insert trunk, and it mainly meant for like the TTS stuff, but they need text capabilities are also there.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Pavan Kumar Reddy

Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 26:34 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Insight
Reddy: Continuous flow matching outperforms discrete audio tokens for speech generation
“So the thing we did differently is instead of having this autoregressive K step prediction, we have a flow matching model. Instead of modeling this as a discrete token set, we trained the codec to be both discrete and continuous to have this flexibility. So we…”
Pavan Kumar Reddy Mar 30, 2026 ▶ 6:42 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Opinion
Reddy: Mistral introduces the first strong open multilingual causal audio encoder
“And there we have a causal encoder. And I don't think there's any strong multilingual causal encoder out in the community. So we thought it's a good contribution.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 30:14 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Insight
Reddy: Audio AI has no winning architecture yet
“One more meta point is unlike text, even in vision, I think this is true, but in audio, it's definitely true. There is no winner model yet. There is no, okay, this is the way you do things. It's still evolving. I think people are still iterating and figuring o…”
Pavan Kumar Reddy Mar 30, 2026 ▶ 8:06 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Disclosure
Reddy: Mistral chose autoregressive TTS to enable real-time streaming voice agents
“One of the main applications is voice agents and we want real time streaming and that's the use case. That's not the only use case, but that's one of the primary use cases we want to get to. So we pick the autoregressive approach for that.”
Pavan Kumar Reddy Mar 30, 2026 ▶ 8:59 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Assertion Supported
Reddy: Mistral reduces flow-matching audio inference to 16 steps
“When you have a depth transformer, if you have K tokens, you need to do K autoregressive steps, right? Even though it's a small thing, it's like K steps, which is very latency heavy with flow matching. We were able to cut it down significantly, so we are able …”
Pavan Kumar Reddy Mar 30, 2026 ▶ 13:52 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.