Nov 9, 2023 · 32m · no-priors

No Priors Ep. 40 | With Arthur Mensch, CEO Mistral AI

Arthur Mensch · 24m spoken Elad Gil · 2m spoken Sarah Guo · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Mistral AI co-founder and CEO Arthur Mensch discusses the development of Mistral 7B, the economic and scientific importance of open-source AI, inference efficiency, and the competitive rise of the European artificial intelligence ecosystem.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 17.7% of the talking time here. How this is scored →

The hosts as informed peer 4.5 Guest teaching 4.9 Guest disagreement 1.9 The hosts pushing back 0.3
05100:0010:0020:0030:000:36–6:10 · The hosts as informed peer 4/10 Founding Mistral AI as a European Open Source Champion Sarah demonstrates familiarity with Arthur's background in Chinchilla scaling laws and mixture of experts. Arthur provides a detailed technical explanation of optimal transport in sparse MoE and correcting the flawed 2020 Kaplan scaling laws.6:11–9:03 · The hosts as informed peer 5/10 Inference Economics and Scaling to Frontier Models Elad outlines the commercial importance of inference costs versus training costs for model adoption. Arthur agrees and explains how frontier large models are necessary both for complex reasoning and for distilling into efficient small models.9:03–13:34 · The hosts as informed peer 3/10 Pre-Training Data Quality and Instruction Alignment Sarah asks Arthur to articulate Mistral's open source thesis against proprietary incumbents. Arthur delivers a passionate overview of how machine learning progressed through open academic sharing until corporate opacity stalled scientific progress post-2020.13:34–17:12 · The hosts as informed peer 5/10 Open Source Safety and Resisting Regulatory Capture Elad frames the closed-source safety argument as regulatory capture. Arthur formalizes this by presenting a two-step test showing LLMs offer no marginal capability over web search engines and knowledge is rarely the bottleneck for misuse.17:13–22:30 · The hosts as informed peer 6/10 Compute Thresholds and the Origin of Bioweapon Narratives Elad leverages his decade of experience as a biologist to challenge the plausibility of LLM bioweapon proliferation. Arthur traces the myth to an unvetted remark in the GPT-4 system card amplified by an echo chamber of circular policy citations.22:31–27:46 · The hosts as informed peer 6/10 Modular Guardrails and Evaluating AI Risk Categories Elad articulates a three-tier taxonomy of AI risk spanning content moderation, physical threats, and existential doom. Arthur argues base models must remain uncensored while guardrails should operate as modular filters at the application layer.27:47–30:38 · The hosts as informed peer 4/10 Overcoming Technical and Cost Bottlenecks for AI Agents Elad asks about overcoming current technical bottlenecks in autonomous AI agents. Arthur explains how agent loops cause mode collapse and high costs, highlighting Mistral's memory-efficient attention mechanisms and API time-sharing.30:38–32:26 · The hosts as informed peer 3/10 Europe's Mathematical Talent and the French AI Ecosystem Sarah inquires about the viability of building a world-class AI champion in Europe. Arthur highlights the depth of mathematical talent in France, the UK, and Poland, alongside the talent flywheel generated by DeepMind and Meta Paris labs.0:36–6:10 · Guest teaching 6/10 Founding Mistral AI as a European Open Source Champion Sarah demonstrates familiarity with Arthur's background in Chinchilla scaling laws and mixture of experts. Arthur provides a detailed technical explanation of optimal transport in sparse MoE and correcting the flawed 2020 Kaplan scaling laws.6:11–9:03 · Guest teaching 4/10 Inference Economics and Scaling to Frontier Models Elad outlines the commercial importance of inference costs versus training costs for model adoption. Arthur agrees and explains how frontier large models are necessary both for complex reasoning and for distilling into efficient small models.9:03–13:34 · Guest teaching 6/10 Pre-Training Data Quality and Instruction Alignment Sarah asks Arthur to articulate Mistral's open source thesis against proprietary incumbents. Arthur delivers a passionate overview of how machine learning progressed through open academic sharing until corporate opacity stalled scientific progress post-2020.13:34–17:12 · Guest teaching 5/10 Open Source Safety and Resisting Regulatory Capture Elad frames the closed-source safety argument as regulatory capture. Arthur formalizes this by presenting a two-step test showing LLMs offer no marginal capability over web search engines and knowledge is rarely the bottleneck for misuse.17:13–22:30 · Guest teaching 6/10 Compute Thresholds and the Origin of Bioweapon Narratives Elad leverages his decade of experience as a biologist to challenge the plausibility of LLM bioweapon proliferation. Arthur traces the myth to an unvetted remark in the GPT-4 system card amplified by an echo chamber of circular policy citations.22:31–27:46 · Guest teaching 5/10 Modular Guardrails and Evaluating AI Risk Categories Elad articulates a three-tier taxonomy of AI risk spanning content moderation, physical threats, and existential doom. Arthur argues base models must remain uncensored while guardrails should operate as modular filters at the application layer.27:47–30:38 · Guest teaching 4/10 Overcoming Technical and Cost Bottlenecks for AI Agents Elad asks about overcoming current technical bottlenecks in autonomous AI agents. Arthur explains how agent loops cause mode collapse and high costs, highlighting Mistral's memory-efficient attention mechanisms and API time-sharing.30:38–32:26 · Guest teaching 3/10 Europe's Mathematical Talent and the French AI Ecosystem Sarah inquires about the viability of building a world-class AI champion in Europe. Arthur highlights the depth of mathematical talent in France, the UK, and Poland, alongside the talent flywheel generated by DeepMind and Meta Paris labs.0:36–6:10 · Guest disagreement 1/10 Founding Mistral AI as a European Open Source Champion Sarah demonstrates familiarity with Arthur's background in Chinchilla scaling laws and mixture of experts. Arthur provides a detailed technical explanation of optimal transport in sparse MoE and correcting the flawed 2020 Kaplan scaling laws.6:11–9:03 · Guest disagreement 0/10 Inference Economics and Scaling to Frontier Models Elad outlines the commercial importance of inference costs versus training costs for model adoption. Arthur agrees and explains how frontier large models are necessary both for complex reasoning and for distilling into efficient small models.9:03–13:34 · Guest disagreement 3/10 Pre-Training Data Quality and Instruction Alignment Sarah asks Arthur to articulate Mistral's open source thesis against proprietary incumbents. Arthur delivers a passionate overview of how machine learning progressed through open academic sharing until corporate opacity stalled scientific progress post-2020.13:34–17:12 · Guest disagreement 4/10 Open Source Safety and Resisting Regulatory Capture Elad frames the closed-source safety argument as regulatory capture. Arthur formalizes this by presenting a two-step test showing LLMs offer no marginal capability over web search engines and knowledge is rarely the bottleneck for misuse.17:13–22:30 · Guest disagreement 4/10 Compute Thresholds and the Origin of Bioweapon Narratives Elad leverages his decade of experience as a biologist to challenge the plausibility of LLM bioweapon proliferation. Arthur traces the myth to an unvetted remark in the GPT-4 system card amplified by an echo chamber of circular policy citations.22:31–27:46 · Guest disagreement 3/10 Modular Guardrails and Evaluating AI Risk Categories Elad articulates a three-tier taxonomy of AI risk spanning content moderation, physical threats, and existential doom. Arthur argues base models must remain uncensored while guardrails should operate as modular filters at the application layer.27:47–30:38 · Guest disagreement 0/10 Overcoming Technical and Cost Bottlenecks for AI Agents Elad asks about overcoming current technical bottlenecks in autonomous AI agents. Arthur explains how agent loops cause mode collapse and high costs, highlighting Mistral's memory-efficient attention mechanisms and API time-sharing.30:38–32:26 · Guest disagreement 0/10 Europe's Mathematical Talent and the French AI Ecosystem Sarah inquires about the viability of building a world-class AI champion in Europe. Arthur highlights the depth of mathematical talent in France, the UK, and Poland, alongside the talent flywheel generated by DeepMind and Meta Paris labs.0:36–6:10 · The hosts pushing back 0/10 Founding Mistral AI as a European Open Source Champion Sarah demonstrates familiarity with Arthur's background in Chinchilla scaling laws and mixture of experts. Arthur provides a detailed technical explanation of optimal transport in sparse MoE and correcting the flawed 2020 Kaplan scaling laws.6:11–9:03 · The hosts pushing back 0/10 Inference Economics and Scaling to Frontier Models Elad outlines the commercial importance of inference costs versus training costs for model adoption. Arthur agrees and explains how frontier large models are necessary both for complex reasoning and for distilling into efficient small models.9:03–13:34 · The hosts pushing back 0/10 Pre-Training Data Quality and Instruction Alignment Sarah asks Arthur to articulate Mistral's open source thesis against proprietary incumbents. Arthur delivers a passionate overview of how machine learning progressed through open academic sharing until corporate opacity stalled scientific progress post-2020.13:34–17:12 · The hosts pushing back 0/10 Open Source Safety and Resisting Regulatory Capture Elad frames the closed-source safety argument as regulatory capture. Arthur formalizes this by presenting a two-step test showing LLMs offer no marginal capability over web search engines and knowledge is rarely the bottleneck for misuse.17:13–22:30 · The hosts pushing back 1/10 Compute Thresholds and the Origin of Bioweapon Narratives Elad leverages his decade of experience as a biologist to challenge the plausibility of LLM bioweapon proliferation. Arthur traces the myth to an unvetted remark in the GPT-4 system card amplified by an echo chamber of circular policy citations.22:31–27:46 · The hosts pushing back 1/10 Modular Guardrails and Evaluating AI Risk Categories Elad articulates a three-tier taxonomy of AI risk spanning content moderation, physical threats, and existential doom. Arthur argues base models must remain uncensored while guardrails should operate as modular filters at the application layer.27:47–30:38 · The hosts pushing back 0/10 Overcoming Technical and Cost Bottlenecks for AI Agents Elad asks about overcoming current technical bottlenecks in autonomous AI agents. Arthur explains how agent loops cause mode collapse and high costs, highlighting Mistral's memory-efficient attention mechanisms and API time-sharing.30:38–32:26 · The hosts pushing back 0/10 Europe's Mathematical Talent and the French AI Ecosystem Sarah inquires about the viability of building a world-class AI champion in Europe. Arthur highlights the depth of mathematical talent in France, the UK, and Poland, alongside the talent flywheel generated by DeepMind and Meta Paris labs.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 38.8% · guest 61.2%0:00 · the hosts 38.8% · guest 61.2%3:00 · the hosts 0.8% · guest 99.2%3:00 · the hosts 0.8% · guest 99.2%6:00 · the hosts 28.6% · guest 71.4%6:00 · the hosts 28.6% · guest 71.4%9:00 · the hosts 10.7% · guest 89.3%9:00 · the hosts 10.7% · guest 89.3%12:00 · the hosts 21.2% · guest 78.8%12:00 · the hosts 21.2% · guest 78.8%15:00 · the hosts 3.2% · guest 96.8%15:00 · the hosts 3.2% · guest 96.8%18:00 · the hosts 13.5% · guest 86.5%18:00 · the hosts 13.5% · guest 86.5%21:00 · the hosts 12.1% · guest 87.9%21:00 · the hosts 12.1% · guest 87.9%24:00 · the hosts 20.9% · guest 79.1%24:00 · the hosts 20.9% · guest 79.1%27:00 · the hosts 19.2% · guest 80.8%27:00 · the hosts 19.2% · guest 80.8%30:00 · the hosts 26.6% · guest 73.4%30:00 · the hosts 26.6% · guest 73.4%
Sharpest disagreement ▶ 20:30 Arthur critiques circular policy papers manufacturing bioweapon panic

Arthur forcefully exposes the lack of empirical backing behind bioweapon regulations, describing how non-scientific policy briefs cite one another in an echo chamber.

Hardest push from the hosts ▶ 22:30 Sarah presses on pragmatic guardrails beyond bioweapon debates

Sarah redirects the conversation away from hypothetical bioweapons to demand Arthur specify concrete guardrails against real-world harmful generations.

Biggest teaching moment ▶ 4:40 Arthur details how Chinchilla overturned prevailing scaling assumptions

Arthur educates the hosts on the mathematical breakdown of Kaplan's 2020 scaling paper, explaining why training tokens must scale proportionally with model parameter count.

The host holds their own ▶ 19:32 Elad draws on biology domain expertise to question AI viral threat claims

Elad invokes his personal decade-long career as a working biologist to substantiate why digital LLM capabilities cannot easily translate into complex physical wet-lab viral synthesis.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Founding Mistral AI as a European Open Source Champion 4610 Sarah demonstrates familiarity with Arthur's background in Chinchilla scaling laws and mixture of experts. Arthur provides a detailed technical explanation of optimal transport in sparse MoE and correcting the flawed 2020 Kaplan scaling laws.
Inference Economics and Scaling to Frontier Models 5400 Elad outlines the commercial importance of inference costs versus training costs for model adoption. Arthur agrees and explains how frontier large models are necessary both for complex reasoning and for distilling into efficient small models.
Pre-Training Data Quality and Instruction Alignment 3630 Sarah asks Arthur to articulate Mistral's open source thesis against proprietary incumbents. Arthur delivers a passionate overview of how machine learning progressed through open academic sharing until corporate opacity stalled scientific progress post-2020.
Open Source Safety and Resisting Regulatory Capture 5540 Elad frames the closed-source safety argument as regulatory capture. Arthur formalizes this by presenting a two-step test showing LLMs offer no marginal capability over web search engines and knowledge is rarely the bottleneck for misuse.
Compute Thresholds and the Origin of Bioweapon Narratives 6641 Elad leverages his decade of experience as a biologist to challenge the plausibility of LLM bioweapon proliferation. Arthur traces the myth to an unvetted remark in the GPT-4 system card amplified by an echo chamber of circular policy citations.
Modular Guardrails and Evaluating AI Risk Categories 6531 Elad articulates a three-tier taxonomy of AI risk spanning content moderation, physical threats, and existential doom. Arthur argues base models must remain uncensored while guardrails should operate as modular filters at the application layer.
Overcoming Technical and Cost Bottlenecks for AI Agents 4400 Elad asks about overcoming current technical bottlenecks in autonomous AI agents. Arthur explains how agent loops cause mode collapse and high costs, highlighting Mistral's memory-efficient attention mechanisms and API time-sharing.
Europe's Mathematical Talent and the French AI Ecosystem 3300 Sarah inquires about the viability of building a world-class AI champion in Europe. Arthur highlights the depth of mathematical talent in France, the UK, and Poland, alongside the talent flywheel generated by DeepMind and Meta Paris labs.

Statements from this episode (20)

Opinion
Mensch: 2020 OpenAI scaling laws paper was poorly executed
“Basically the story was that everybody was training models on too few tokens because of the paper from 20 20 that happened to be not very well executed.”
Arthur Mensch Nov 9, 2023 ▶ 4:36
Assertion Supported
Mensch: Chinchilla scaling yields models four times cheaper to serve
“For the same amount of compute, you would get a model that would be better, but also a model that will be four times cheaper to serve.”
Arthur Mensch Nov 9, 2023 ▶ 5:27
Assertion Not checkable as stated
Mensch: Mistral 7B proved AI model compression limits are far away
“What we showed with Mistral Seven B is that we weren't, we were definitely far away from the limit of compression.”
Arthur Mensch Nov 9, 2023 ▶ 5:51
Insight
Mensch: Pure scientific model performance ignores crucial runtime inference costs
“And if you want to push the performance, the pure performance of models, you don't care about inference because you, well, you are not going to use the model. You're just going to see whether they're good or not. And that's really for scientific purposes. But …”
Arthur Mensch Nov 9, 2023 ▶ 7:14
Opinion
Mensch: Advancing reasoning capabilities requires scaling to larger models
“There's still a limit to what a certain model size can do. This limit was, I think, underestimated. But if you want to get to more reasoning capabilities, you do need to move into larger models.”
Arthur Mensch Nov 9, 2023 ▶ 8:28
Disclosure
Mensch: Distilling state-of-the-art small models requires training massive models first
“The other thing about moving into larger models is that it enables you to train smaller models that are better, which is through variety of techniques like distillation or synthetic data generation. So this, these two things are quite related. If you want to m…”
Arthur Mensch Nov 9, 2023 ▶ 8:42
Disclosure
Mensch: Mistral Is Not Yet the Top Expert in Instruction Tuning
“We're not the top experts in the world in, in, in making good instruction fine tune models. We're definitely ramping up and the team is, is getting better and better at that.”
Arthur Mensch Nov 9, 2023 ▶ 10:09
Opinion
Mensch: Closed-source AI wastes billions in duplicated compute
“We think that it's too early, and we think it's really, ah, damaging for the science, ah, to actually move into such an opaque regime, where you have A few companies basically doing the same thing, just not communicating about it spending billions of compute d…”
Arthur Mensch Nov 9, 2023 ▶ 11:56
Disclosure
Arthur Mensch: Mistral Aims to Make AI Safer Through Openness and Scrutiny
“By doing what we do, by being much more open about the technology we create, we want to steer the community into a regime where things just work better, where things are safer because put under more scrutiny, and really our intention there is to, well, to take…”
Arthur Mensch Nov 9, 2023 ▶ 13:11
Assertion Supported
Mensch: LLMs are not marginally more dangerous than web search
“Nothing is showing that LLM is actually marginally better than a search engine to find knowledge on, on topics that would enable bad use.”
Arthur Mensch Nov 9, 2023 ▶ 15:13
Opinion
Mensch: Banning open-source AI enforces regulatory capture for incumbents
“Today going, banning open source, preventing it from happening is really a way, well, to enforce regulatory capture, even though the actors that would benefit from it, Don't want it to happen. But by design, if you actually ban small actors from doing things i…”
Arthur Mensch Nov 9, 2023 ▶ 16:45
Disclosure
Mensch: Mistral cannot afford 10^26 FLOP training runs for coming years
“That's not something we can even afford. And that we won't be able to afford for the coming years.”
Arthur Mensch Nov 9, 2023 ▶ 17:41
Assertion Not checkable as stated
Mensch: AI bioweapon risk narrative relies on unscientific circular citations
“No scientific studies in proper form was published, but then policy papers started to cite Non-scientific papers arguing that these were scientific evidences that the bioweapon narrative was actually true. And then policy papers started to cite the other polic…”
Arthur Mensch Nov 9, 2023 ▶ 21:01
Prediction Open · timeframe Nov 2028
Mensch: AI will not trigger the next pandemic; climate change will
“I don't think AI is going to be the one triggering the next pandemic. It's always going to be climate change.”
Arthur Mensch Nov 9, 2023 ▶ 22:11
Insight
Mensch: Base AI models should know everything, with guardrails added externally
“Assuming that the model should be well behaved is, I think, a wrong assumption. You need to make the assumption that the model should know everything. And then on top of that, have some modules that moderate and guardrail the model.”
Arthur Mensch Nov 9, 2023 ▶ 23:59
Assertion Not checkable as stated
Mensch: Zero evidence AI development is headed toward a singularity
“If we can make a model which is growingly intelligent, then maybe you're at a singularity level. There's no evidence whatsoever that we are on the way of doing that, of making that happen.”
Arthur Mensch Nov 9, 2023 ▶ 27:22
Insight
Mensch: Viable AI agents require a 100x reduction in compute costs
“I think making a model smaller is definitely a way to make agent work. Because one problem you have with agent that very quickly is start to, if you run an agent on GPT four you're going to run out of money very quickly. And so if you divide by a hundred the, …”
Arthur Mensch Nov 9, 2023 ▶ 28:12
Assertion Not checkable as stated
Mensch: AI agents currently suffer from repetitive mode collapse
“What we see with agent is mode collapse. So not very interesting mode collapse. They start repeating themselves and they fall into loops.”
Arthur Mensch Nov 9, 2023 ▶ 28:33
Assertion Not checkable as stated
Mensch: Single H100 GPU can serve hundreds of API customers
“It's going to be less costly because just a single H-Handroid can serve hundreds of customers.”
Arthur Mensch Nov 9, 2023 ▶ 30:20
Insight
Mensch: Strong European Math Training Produces Top AI Talent
“France, UK, Poland are very good at training mathematicians. And as it turns out, mathematicians are very good at making AI.”
Arthur Mensch Nov 9, 2023 ▶ 31:09
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.