Pavan Kumar Reddy, Audio Research Lead at Mistral AI, discusses the architecture of Mistral's real-time audio models and why Mistral trained custom causal encoders in-house.
“And there we have a causal encoder. And I don't think there's any strong multilingual causal encoder out in the community. So we thought it's a good contribution.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Pavan Kumar Reddy
AssertionSupported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
“So the thing we did differently is instead of having this autoregressive K step prediction, we have a flow matching model. Instead of modeling this as a discrete token set, we trained the codec to be both discrete and continuous to have this flexibility. So we…”
“One more meta point is unlike text, even in vision, I think this is true, but in audio, it's definitely true. There is no winner model yet. There is no, okay, this is the way you do things. It's still evolving. I think people are still iterating and figuring o…”
Reddy: Voxtral TTS is a 3B model based on the Ministral architecture
“It's it came out with such good quality, and Guillaume was mentioning, yeah, it's a three B model it's based off of the ministral model that we actually released just a few months back, and insert trunk, and it mainly meant for like the TTS stuff, but they nee…”
“One of the main applications is voice agents and we want real time streaming and that's the use case. That's not the only use case, but that's one of the primary use cases we want to get to. So we pick the autoregressive approach for that.”
Reddy: Mistral reduces flow-matching audio inference to 16 steps
“When you have a depth transformer, if you have K tokens, you need to do K autoregressive steps, right? Even though it's a small thing, it's like K steps, which is very latency heavy with flow matching. We were able to cut it down significantly, so we are able …”
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.