Mistral AI Chief Scientist Guillaume Lample responds to industry debates regarding whether frontier labs should shift the majority of training compute away from pre-training toward reinforcement learning.
Disclosure
Lample: Mistral plans a phased approach to full-duplex audio models
“Ultimately what we want to do is to be this, Full duplex model, but we are not going to start this, start there directly. I think it's some approach that people are doing, but. Just to confirm, full duplex means it can speak while I'm speaking, or? Okay. Yeah,…”
Insight
Lample: Specialized AI models are more cost-effective than monolithic models
“That's why we can actually use models audio, but also like OCRs that are like really, really good at that, and that will be much more cost effective than a general model. That will contain a lot of capabilities you don't really need.”
Insight
Lample: 1B to 3B parameter models are optimal for speech transcription
“For instance, for audio here, if you want to do transcription, I think it makes no sense to use a model as this large. If you just want to transcribe tech, it would be very inefficient. Like if you want to do audio, you probably just want to do the one B or a …”
Insight
Lample: Long-horizon RL trajectories require new algorithms beyond GRPO
“GRPO, for instance, it doesn't really work with any bit of policy, which was okay initially, because you are solving math problems that can be solved in like a few thousand tokens, so the model can actually generate them pretty quickly, so when you do your upd…”
Assertion Not checkable as stated
Lample: Voxtral TTS matches leading models at a fraction of cost
“So we support nine languages and this is a pretty small model a three-dimensional model, so very fast, and also state-ups, yeah, very equal. Performed at the same level of the best model, but it's Much more efficient in terms of cost, and also much, in terms o…”
What-if
Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”