Insight
Mensch: AI base models should know everything and use modular guardrails
“Assuming that the model should be well behaved is, I think, a wrong assumption. You need to make the assumption that the model should know everything. And then on top of that, have some modules that moderate and guardrail the model.”
Opinion
Mensch: AI safety should regulate application makers, not trust big model labs
“And the way you make this healthy competition is not By trusting a couple of companies to do their own safety, it's rather for, it's rather the way you do it is to ask application makers to comply with some rules.”
Assertion Supported
Mensch: LLMs are not marginally more dangerous than web search
“Nothing is showing that LLM is actually marginally better than a search engine to find knowledge on, on topics that would enable bad use.”
Opinion
Mensch: Banning open-source AI enforces regulatory capture for incumbents
“Today going, banning open source, preventing it from happening is really a way, well, to enforce regulatory capture, even though the actors that would benefit from it, Don't want it to happen. But by design, if you actually ban small actors from doing things i…”
Assertion Not checkable as stated
Mensch: AI bioweapon risk narrative relies on unscientific circular citations
“No scientific studies in proper form was published, but then policy papers started to cite Non-scientific papers arguing that these were scientific evidences that the bioweapon narrative was actually true. And then policy papers started to cite the other polic…”
Prediction Open · timeframe Nov 2028
Mensch: AI will not trigger the next pandemic; climate change will
“I don't think AI is going to be the one triggering the next pandemic. It's always going to be climate change.”
Assertion Not checkable as stated
Mensch: Zero evidence AI development is headed toward a singularity
“If we can make a model which is growingly intelligent, then maybe you're at a singularity level. There's no evidence whatsoever that we are on the way of doing that, of making that happen.”
Opinion
Mensch: 2020 OpenAI scaling laws paper was poorly executed
“Basically the story was that everybody was training models on too few tokens because of the paper from 20 20 that happened to be not very well executed.”
Opinion
Mensch: Advancing reasoning capabilities requires scaling to larger models
“There's still a limit to what a certain model size can do. This limit was, I think, underestimated. But if you want to get to more reasoning capabilities, you do need to move into larger models.”
Opinion
Mensch: Closed-source AI wastes billions in duplicated compute
“We think that it's too early, and we think it's really, ah, damaging for the science, ah, to actually move into such an opaque regime, where you have A few companies basically doing the same thing, just not communicating about it spending billions of compute d…”
Assertion Not checkable as stated
Mensch: Mistral 7B proved AI model compression limits are far away
“What we showed with Mistral Seven B is that we weren't, we were definitely far away from the limit of compression.”
Disclosure
Mensch: Distilling state-of-the-art small models requires training massive models first
“The other thing about moving into larger models is that it enables you to train smaller models that are better, which is through variety of techniques like distillation or synthetic data generation. So this, these two things are quite related. If you want to m…”
Disclosure
Arthur Mensch: Mistral Aims to Make AI Safer Through Openness and Scrutiny
“By doing what we do, by being much more open about the technology we create, we want to steer the community into a regime where things just work better, where things are safer because put under more scrutiny, and really our intention there is to, well, to take…”
Insight
Mensch: Base AI models should know everything, with guardrails added externally
“Assuming that the model should be well behaved is, I think, a wrong assumption. You need to make the assumption that the model should know everything. And then on top of that, have some modules that moderate and guardrail the model.”
Insight
Mensch: Viable AI agents require a 100x reduction in compute costs
“I think making a model smaller is definitely a way to make agent work. Because one problem you have with agent that very quickly is start to, if you run an agent on GPT four you're going to run out of money very quickly. And so if you divide by a hundred the, …”
Disclosure
Mensch: Mistral Is Not Yet the Top Expert in Instruction Tuning
“We're not the top experts in the world in, in, in making good instruction fine tune models. We're definitely ramping up and the team is, is getting better and better at that.”
Assertion Not checkable as stated
Mensch: Single H100 GPU can serve hundreds of API customers
“It's going to be less costly because just a single H-Handroid can serve hundreds of customers.”
Insight
Mensch: Strong European Math Training Produces Top AI Talent
“France, UK, Poland are very good at training mathematicians. And as it turns out, mathematicians are very good at making AI.”
Assertion Supported
Mensch: Chinchilla scaling yields models four times cheaper to serve
“For the same amount of compute, you would get a model that would be better, but also a model that will be four times cheaper to serve.”
Insight
Mensch: Pure scientific model performance ignores crucial runtime inference costs
“And if you want to push the performance, the pure performance of models, you don't care about inference because you, well, you are not going to use the model. You're just going to see whether they're good or not. And that's really for scientific purposes. But …”
Disclosure
Mensch: Mistral cannot afford 10^26 FLOP training runs for coming years
“That's not something we can even afford. And that we won't be able to afford for the coming years.”
Assertion Not checkable as stated
Mensch: AI agents currently suffer from repetitive mode collapse
“What we see with agent is mode collapse. So not very interesting mode collapse. They start repeating themselves and they fall into loops.”