Dec 28, 2023 · 39m · a16z

Safety in Numbers: Keeping AI Open

Arthur Mensch · 26m spoken Anjney Midha · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, Mistral AI co-founder Arthur Mensch discusses compute-optimal scaling laws, open-source model architecture, and AI safety with host Anjney Midha. He advocates for open weights, outlines technical innovations like Sparse Mixture of Experts, and argues against over-regulating foundational technology.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.2 Guest teaching 3.2 Guest disagreement 0.9 The host pushing back 0.0
05100:0010:0020:0030:001:43–4:40 · The host as informed peer 3/10 Legal Disclaimer and Financial Disclosures Anjney sets up the historical context of parameter size vs dataset size in foundation models. Arthur explains how the Chinchilla paper corrected scaling law misconceptions by showing compute should be balanced equally between model size and data points.4:40–8:18 · The host as informed peer 4/10 Founding Mistral AI and Overtraining for Inference Efficiency Anjney demonstrates solid background knowledge regarding LLaMA's origin and compute efficiency. Arthur breaks down how overtraining models beyond Chinchilla optimal thresholds sacrifices extra training compute to dramatically lower inference costs.8:18–13:19 · The host as informed peer 3/10 Unveiling Mixtral: Sparse Mixture of Experts Architecture Anjney prompts an explanation of Sparse Mixture of Experts architecture. Arthur educates on how routing tokens across duplicate expert layers decouples parameter capacity from active inference compute, while lightheartedly withholding specific proprietary trade secrets.13:19–17:55 · The host as informed peer 3/10 The Open Source Philosophy: Ideology, Distribution, and Control Anjney highlights the industry shift toward closed models post-GPT-3. Arthur details the ideological and business justifications for open source, arguing that pre-trained base models should remain neutral rather than imposing vendor biases.17:55–21:26 · The host as informed peer 3/10 Benchmarking Open Source vs. Proprietary Model Performance Anjney asks about performance limits and trade-offs of open source models. Arthur confidently asserts that open source lags proprietary frontier models by only six months and will eventually mirror Linux in server dominance.21:26–27:28 · The host as informed peer 4/10 Debunking AI Safety Misconceptions and Regulatory Frameworks Anjney brings up common public safety concerns surrounding open weights. Arthur reframes open source as a net safety benefit, asserting that open red-teaming and community scrutiny are far superior to proprietary closed sandboxes.27:28–31:46 · The host as informed peer 3/10 Regulating Applications, Not Math: The Programming Language Analogy Anjney asks Arthur to articulate the core stakes in the open vs closed regulatory battle. Arthur firmly rejects attempts to regulate compute or raw models, arguing that regulators should govern end applications rather than underlying math tools.31:46–35:08 · The host as informed peer 4/10 The Next Frontier: Reasoning, Data Efficiency, and Adaptive Compute Anjney demonstrates deep domain fluency by citing self-play, process reward models, and out-of-distribution reasoning. Arthur notes that current models are vastly less data efficient than the human brain and suggests adaptive compute will drive the next frontier.35:08–38:11 · The host as informed peer 2/10 Future Outlook: Interactive Modalities and Call to Action for Builders Anjney invites Arthur to share his product vision and call to action for developers. Arthur forecasts specialized multi-agent applications in enterprise and gaming, urging software builders to leverage open Mistral weights directly.1:43–4:40 · Guest teaching 3/10 Legal Disclaimer and Financial Disclosures Anjney sets up the historical context of parameter size vs dataset size in foundation models. Arthur explains how the Chinchilla paper corrected scaling law misconceptions by showing compute should be balanced equally between model size and data points.4:40–8:18 · Guest teaching 3/10 Founding Mistral AI and Overtraining for Inference Efficiency Anjney demonstrates solid background knowledge regarding LLaMA's origin and compute efficiency. Arthur breaks down how overtraining models beyond Chinchilla optimal thresholds sacrifices extra training compute to dramatically lower inference costs.8:18–13:19 · Guest teaching 4/10 Unveiling Mixtral: Sparse Mixture of Experts Architecture Anjney prompts an explanation of Sparse Mixture of Experts architecture. Arthur educates on how routing tokens across duplicate expert layers decouples parameter capacity from active inference compute, while lightheartedly withholding specific proprietary trade secrets.13:19–17:55 · Guest teaching 3/10 The Open Source Philosophy: Ideology, Distribution, and Control Anjney highlights the industry shift toward closed models post-GPT-3. Arthur details the ideological and business justifications for open source, arguing that pre-trained base models should remain neutral rather than imposing vendor biases.17:55–21:26 · Guest teaching 3/10 Benchmarking Open Source vs. Proprietary Model Performance Anjney asks about performance limits and trade-offs of open source models. Arthur confidently asserts that open source lags proprietary frontier models by only six months and will eventually mirror Linux in server dominance.21:26–27:28 · Guest teaching 4/10 Debunking AI Safety Misconceptions and Regulatory Frameworks Anjney brings up common public safety concerns surrounding open weights. Arthur reframes open source as a net safety benefit, asserting that open red-teaming and community scrutiny are far superior to proprietary closed sandboxes.27:28–31:46 · Guest teaching 4/10 Regulating Applications, Not Math: The Programming Language Analogy Anjney asks Arthur to articulate the core stakes in the open vs closed regulatory battle. Arthur firmly rejects attempts to regulate compute or raw models, arguing that regulators should govern end applications rather than underlying math tools.31:46–35:08 · Guest teaching 3/10 The Next Frontier: Reasoning, Data Efficiency, and Adaptive Compute Anjney demonstrates deep domain fluency by citing self-play, process reward models, and out-of-distribution reasoning. Arthur notes that current models are vastly less data efficient than the human brain and suggests adaptive compute will drive the next frontier.35:08–38:11 · Guest teaching 2/10 Future Outlook: Interactive Modalities and Call to Action for Builders Anjney invites Arthur to share his product vision and call to action for developers. Arthur forecasts specialized multi-agent applications in enterprise and gaming, urging software builders to leverage open Mistral weights directly.1:43–4:40 · Guest disagreement 0/10 Legal Disclaimer and Financial Disclosures Anjney sets up the historical context of parameter size vs dataset size in foundation models. Arthur explains how the Chinchilla paper corrected scaling law misconceptions by showing compute should be balanced equally between model size and data points.4:40–8:18 · Guest disagreement 0/10 Founding Mistral AI and Overtraining for Inference Efficiency Anjney demonstrates solid background knowledge regarding LLaMA's origin and compute efficiency. Arthur breaks down how overtraining models beyond Chinchilla optimal thresholds sacrifices extra training compute to dramatically lower inference costs.8:18–13:19 · Guest disagreement 1/10 Unveiling Mixtral: Sparse Mixture of Experts Architecture Anjney prompts an explanation of Sparse Mixture of Experts architecture. Arthur educates on how routing tokens across duplicate expert layers decouples parameter capacity from active inference compute, while lightheartedly withholding specific proprietary trade secrets.13:19–17:55 · Guest disagreement 1/10 The Open Source Philosophy: Ideology, Distribution, and Control Anjney highlights the industry shift toward closed models post-GPT-3. Arthur details the ideological and business justifications for open source, arguing that pre-trained base models should remain neutral rather than imposing vendor biases.17:55–21:26 · Guest disagreement 1/10 Benchmarking Open Source vs. Proprietary Model Performance Anjney asks about performance limits and trade-offs of open source models. Arthur confidently asserts that open source lags proprietary frontier models by only six months and will eventually mirror Linux in server dominance.21:26–27:28 · Guest disagreement 2/10 Debunking AI Safety Misconceptions and Regulatory Frameworks Anjney brings up common public safety concerns surrounding open weights. Arthur reframes open source as a net safety benefit, asserting that open red-teaming and community scrutiny are far superior to proprietary closed sandboxes.27:28–31:46 · Guest disagreement 2/10 Regulating Applications, Not Math: The Programming Language Analogy Anjney asks Arthur to articulate the core stakes in the open vs closed regulatory battle. Arthur firmly rejects attempts to regulate compute or raw models, arguing that regulators should govern end applications rather than underlying math tools.31:46–35:08 · Guest disagreement 1/10 The Next Frontier: Reasoning, Data Efficiency, and Adaptive Compute Anjney demonstrates deep domain fluency by citing self-play, process reward models, and out-of-distribution reasoning. Arthur notes that current models are vastly less data efficient than the human brain and suggests adaptive compute will drive the next frontier.35:08–38:11 · Guest disagreement 0/10 Future Outlook: Interactive Modalities and Call to Action for Builders Anjney invites Arthur to share his product vision and call to action for developers. Arthur forecasts specialized multi-agent applications in enterprise and gaming, urging software builders to leverage open Mistral weights directly.1:43–4:40 · The host pushing back 0/10 Legal Disclaimer and Financial Disclosures Anjney sets up the historical context of parameter size vs dataset size in foundation models. Arthur explains how the Chinchilla paper corrected scaling law misconceptions by showing compute should be balanced equally between model size and data points.4:40–8:18 · The host pushing back 0/10 Founding Mistral AI and Overtraining for Inference Efficiency Anjney demonstrates solid background knowledge regarding LLaMA's origin and compute efficiency. Arthur breaks down how overtraining models beyond Chinchilla optimal thresholds sacrifices extra training compute to dramatically lower inference costs.8:18–13:19 · The host pushing back 0/10 Unveiling Mixtral: Sparse Mixture of Experts Architecture Anjney prompts an explanation of Sparse Mixture of Experts architecture. Arthur educates on how routing tokens across duplicate expert layers decouples parameter capacity from active inference compute, while lightheartedly withholding specific proprietary trade secrets.13:19–17:55 · The host pushing back 0/10 The Open Source Philosophy: Ideology, Distribution, and Control Anjney highlights the industry shift toward closed models post-GPT-3. Arthur details the ideological and business justifications for open source, arguing that pre-trained base models should remain neutral rather than imposing vendor biases.17:55–21:26 · The host pushing back 0/10 Benchmarking Open Source vs. Proprietary Model Performance Anjney asks about performance limits and trade-offs of open source models. Arthur confidently asserts that open source lags proprietary frontier models by only six months and will eventually mirror Linux in server dominance.21:26–27:28 · The host pushing back 0/10 Debunking AI Safety Misconceptions and Regulatory Frameworks Anjney brings up common public safety concerns surrounding open weights. Arthur reframes open source as a net safety benefit, asserting that open red-teaming and community scrutiny are far superior to proprietary closed sandboxes.27:28–31:46 · The host pushing back 0/10 Regulating Applications, Not Math: The Programming Language Analogy Anjney asks Arthur to articulate the core stakes in the open vs closed regulatory battle. Arthur firmly rejects attempts to regulate compute or raw models, arguing that regulators should govern end applications rather than underlying math tools.31:46–35:08 · The host pushing back 0/10 The Next Frontier: Reasoning, Data Efficiency, and Adaptive Compute Anjney demonstrates deep domain fluency by citing self-play, process reward models, and out-of-distribution reasoning. Arthur notes that current models are vastly less data efficient than the human brain and suggests adaptive compute will drive the next frontier.35:08–38:11 · The host pushing back 0/10 Future Outlook: Interactive Modalities and Call to Action for Builders Anjney invites Arthur to share his product vision and call to action for developers. Arthur forecasts specialized multi-agent applications in enterprise and gaming, urging software builders to leverage open Mistral weights directly.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 0%39:00 · the host 0% · guest 0%
Sharpest disagreement ▶ 27:28 Rejecting regulation of foundational math

Arthur forcefully reframes the regulatory debate, rejecting the premise that underlying foundational models or mathematical FLOP counts should be regulated.

Hardest push from the host ▶ 26:08 Pressing on open source failure modes

Anjney directly challenges Arthur by asking what specific factors would cause open source models to fail or fall behind proprietary competitors.

Biggest teaching moment ▶ 30:22 Programming language vs malware analogy

Arthur uses a vivid programming language metaphor to educate non-technical audiences and regulators on why base models are neutral tools rather than dangerous applications.

The host holds their own ▶ 31:47 Drilling on advanced reasoning techniques

Anjney demonstrates high host expertise by citing cutting-edge ML concepts such as self-play, process reward models, and out-of-distribution reasoning.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Legal Disclaimer and Financial Disclosures 3300 Anjney sets up the historical context of parameter size vs dataset size in foundation models. Arthur explains how the Chinchilla paper corrected scaling law misconceptions by showing compute should be balanced equally between model size and data points.
Founding Mistral AI and Overtraining for Inference Efficiency 4300 Anjney demonstrates solid background knowledge regarding LLaMA's origin and compute efficiency. Arthur breaks down how overtraining models beyond Chinchilla optimal thresholds sacrifices extra training compute to dramatically lower inference costs.
Unveiling Mixtral: Sparse Mixture of Experts Architecture 3410 Anjney prompts an explanation of Sparse Mixture of Experts architecture. Arthur educates on how routing tokens across duplicate expert layers decouples parameter capacity from active inference compute, while lightheartedly withholding specific proprietary trade secrets.
The Open Source Philosophy: Ideology, Distribution, and Control 3310 Anjney highlights the industry shift toward closed models post-GPT-3. Arthur details the ideological and business justifications for open source, arguing that pre-trained base models should remain neutral rather than imposing vendor biases.
Benchmarking Open Source vs. Proprietary Model Performance 3310 Anjney asks about performance limits and trade-offs of open source models. Arthur confidently asserts that open source lags proprietary frontier models by only six months and will eventually mirror Linux in server dominance.
Debunking AI Safety Misconceptions and Regulatory Frameworks 4420 Anjney brings up common public safety concerns surrounding open weights. Arthur reframes open source as a net safety benefit, asserting that open red-teaming and community scrutiny are far superior to proprietary closed sandboxes.
Regulating Applications, Not Math: The Programming Language Analogy 3420 Anjney asks Arthur to articulate the core stakes in the open vs closed regulatory battle. Arthur firmly rejects attempts to regulate compute or raw models, arguing that regulators should govern end applications rather than underlying math tools.
The Next Frontier: Reasoning, Data Efficiency, and Adaptive Compute 4310 Anjney demonstrates deep domain fluency by citing self-play, process reward models, and out-of-distribution reasoning. Arthur notes that current models are vastly less data efficient than the human brain and suggests adaptive compute will drive the next frontier.
Future Outlook: Interactive Modalities and Call to Action for Builders 2200 Anjney invites Arthur to share his product vision and call to action for developers. Arthur forecasts specialized multi-agent applications in enterprise and gaming, urging software builders to leverage open Mistral weights directly.

Statements from this episode (27)

Assertion Supported
2022 DeepMind paper showed dataset size matters more than parameter count
“In fact, in twenty-twenty-two, a pivotal paper came out that changed the way that many people in the research community thought about this very calculus. And it demonstrated that datasets were actually more important than just the sheer size of the model.”
Anjney Midha Dec 28, 2023 ▶ 0:12
Assertion Supported
Mensch: 2020-2021 AI research suffered from flawed scaling laws in GPT-3 and Gopher
“There was also a misconception on GPT-free and basically in 20, 21, every paper made this mistake.”
Arthur Mensch Dec 28, 2023 ▶ 3:17
Insight
Mensch: Compute-optimal LLM scaling requires equal relative growth in parameters and data
“In common words, if you multiply by four your compute capacity, you should multiply by two, the model size and by two, the data size.”
Arthur Mensch Dec 28, 2023 ▶ 4:09
Insight
Mensch: Overtraining models past Chinchilla limits lowers inference costs
“If you take into account the fact that your model Should also be efficient at inference time. You probably want to go far beyond the Cinchilla scaling low. So it means you want to overtrain the model. So train on more tokens than would be optimal for performan…”
Arthur Mensch Dec 28, 2023 ▶ 7:15
Assertion Supported
Mensch: Meta's LLaMA paper was first to publicly demonstrate overtraining benefits
“The Lama paper was the first to establish it in the open and it opened a lot of opportunities.”
Arthur Mensch Dec 28, 2023 ▶ 7:45
Assertion Supported
Mensch: Mixtral has 46B total parameters but executes 12B per token
“You have eight experts per layer and you execute only two of them. So what it means at the end of the day is that you have a lot of parameters on your model. You have forty six billion parameters, but the thing is that the number of parameters that you execute…”
Arthur Mensch Dec 28, 2023 ▶ 9:33
Insight
Mensch: Mixture of experts decouples model capacity from inference cost
“A sparse mixture of experts, you take the dense layer and you duplicate it several times. And so that's where you actually increase the number of parameters. So you increase the capacity of the model without increasing the cost. So that's the way of decoupling…”
Arthur Mensch Dec 28, 2023 ▶ 10:59
Assertion Supported
Mensch: Mixtral matches Llama 2 70B performance at one-sixth the cost
“Mixtral is actually on par with Lama-to-seven TB while being approximately six times cheaper or six times faster for the same price.”
Arthur Mensch Dec 28, 2023 ▶ 11:38
Assertion Not checkable as stated
Mensch: AI industry stopped publishing open research after GPT-3
“And all of a sudden in with GPT-free, this tide kind of reversed and companies started to be more opaque about what they were doing because they realized there was actually a very big market. And all of a sudden in 20, 22, on the important aspects of AI and on…”
Arthur Mensch Dec 28, 2023 ▶ 14:34
Prediction Held up
Mistral AI plans to monetize through an open-core business model
“As a business, we do need to have a valid monetization approach at some point. But we've seen many businesses build open core approaches and have, A very strong open source community, and also a very good offer of services, and that's what we want to build.”
Arthur Mensch Dec 28, 2023 ▶ 15:55
Insight
Mensch: Pre-trained AI models should be neutral without creator bias
“Pre-trained models should be neutral, and we should empower our customers to take these models and just put their editorial approaches, their instruction, their constitution, if you want to talk like entropy, Into the model. So that's the way we approach the t…”
Arthur Mensch Dec 28, 2023 ▶ 17:13
Assertion Supported
Mensch: Mixtral matches GPT-3.5 performance
“So mixed trial is as similar performance to GPT, 3.5.”
Arthur Mensch Dec 28, 2023 ▶ 18:13
Assertion Not checkable as stated
Mensch: Internal Mistral models rank among top three globally
“Internally We have stronger models that are in between 3.5 and four that are basically second or third, the second or third best model in the world.”
Arthur Mensch Dec 28, 2023 ▶ 18:19
Assertion Not checkable as stated
Mensch: Open-source AI trails proprietary models by six months
“So really we think that the gap is closing. The gap is approximately six months at that point.”
Arthur Mensch Dec 28, 2023 ▶ 18:27
Prediction Not checkable as stated
Mensch: Open-source models will equal proprietary AI performance
“But I really think that will converge to a setting where you have proprietary models and the open source models are just as good.”
Arthur Mensch Dec 28, 2023 ▶ 19:00
Assertion Not checkable as stated
Mensch: Fine-tuned Mistral 7B matches GPT performance on specific enterprise tasks
“They took Mistral-Seven-B, had a lot of human annotations, had a lot of proprietary data, just modify Mistral-Seven-B so that it solved their task, just as well as GPT-PT-PT-PT, but only for a lower cost and a higher level of control.”
Arthur Mensch Dec 28, 2023 ▶ 19:49
Assertion Supported
Mensch: Hugging Face DPO outperformed Mistral AI's initial instruct model
“The hugging face folks first did the direct preference optimization on top of Mistral seven B and made a very strong, a much stronger model than the instructive model we proposed at the early release.”
Arthur Mensch Dec 28, 2023 ▶ 20:24
Insight
Mensch: Modern AI models are essentially compressions of internet data
“The current generation of modern models that we are using today are not that much are not much more than just a compression of whatever is available on the internet.”
Arthur Mensch Dec 28, 2023 ▶ 21:40
Assertion Not checkable as stated
Mensch: Fine-tuning access makes GPT-4 easy to exploit into bad behavior
“It's actually super easy to exploit an API. It's super easy, especially if you have fine tuning access to make GPT-IV behave in a very bad way.”
Arthur Mensch Dec 28, 2023 ▶ 22:55
Prediction Not checkable as stated
Mensch: Mistral must build frontier models to stay relevant to developers
“Really at the end of the day developers look at performance and latency. And so that's why we think that as a company, we need to be very much on the frontier if we want to be relevant.”
Arthur Mensch Dec 28, 2023 ▶ 26:50
Insight
Mensch: LLMs function as programming languages, not standalone products
“If you look at what the LLM does, it's not really different from a programming language. It's actually used very much as a programming language by the application makers. There's a strong confusion made between what we call a model and what we call an applicat…”
Arthur Mensch Dec 28, 2023 ▶ 27:39
Opinion
Mensch: AI regulation should target applications, not foundational math
“What you want to regulate is the application, and the issue we had, and the issue we're still having now is, We hear a lot of people saying we should regulate the tech, so we should regulate the function, the mathematics behind it, but really you never use a l…”
Arthur Mensch Dec 28, 2023 ▶ 28:44
Opinion
Mensch: FLOP count is the wrong metric for regulating AI models
“Pre-market conditions like flops, the number of flops that you do to create a model is definitely not the right way of doing of measuring the performance of a model.”
Arthur Mensch Dec 28, 2023 ▶ 30:35
Assertion Not checkable as stated
Mensch: LLM training is 100,000x less data efficient than the human brain
“If you compare like the training process of a large language model to the brain, you have like a factor I think 100,000.”
Arthur Mensch Dec 28, 2023 ▶ 32:34
Prediction Not checkable as stated
Mensch: We will not understand machine reasoning anytime soon
“We are not going to know about how machines reason anytime soon.”
Arthur Mensch Dec 28, 2023 ▶ 34:09
Assertion Supported
Mensch: Top mathematicians like Terence Tao use LLMs for proof steps
“We're starting to see some very good mathematicians I'm thinking of Terence Tao that are using large language models for some things obviously not the high level reasoning, but for some part of their proofs.”
Arthur Mensch Dec 28, 2023 ▶ 34:40
Prediction Not checkable as stated
Mensch: Everyone will use specialized AI models within five years
“What we think is that fast forward five years everybody will be using their specialized models within parts of complex applications and systems.”
Arthur Mensch Dec 28, 2023 ▶ 35:31
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.