Jun 12, 2025 · 41m · no-priors

No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann

Ben Mann · 24m spoken Elad Gil · 10m spoken Sarah Guo · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Anthropic co-founder Ben Mann joins No Priors to discuss Claude 4's breakthroughs, the rise of autonomous coding agents and the Model Context Protocol, and the empirical safety frameworks guiding the development of transformative AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 36.3% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 3.8 Guest disagreement 1.2 The hosts pushing back 1.9
05100:0015:0030:000:33–3:36 · The hosts as informed peer 4/10 Claude 4 Release Criteria and Coding Benchmarks Sarah and Elad ask about release versioning criteria and model milestones. Sarah contributes a technical anecdote about reward hacking behaviors observed in portfolio companies.3:36–6:41 · The hosts as informed peer 4/10 Agentic Capabilities, Long Horizons, and Compute Costs Sarah inquires about the cost economics and compute allocation of agentic reasoning tokens. Ben explains how Opus uses Sonnet as sub-agents to optimize latency and context limits.6:42–10:26 · The hosts as informed peer 6/10 Architectural Specialization and Model Routing Layers Elad draws a detailed architectural analogy between biological brain modules and foundation model orchestration layers. Ben references mechanistic interpretability circuits and mixture-of-experts representations.10:27–14:45 · The hosts as informed peer 7/10 Vertical Integration and Developing Claude Code Elad outlines historical vertical integration precedents (Microsoft Office, Google Search) and synthesizes the threefold strategic purpose of coding models. Ben validates this with the Claude Code deployment strategy.14:46–18:07 · The hosts as informed peer 5/10 Economic Turing Test and Research Acceleration Vectors Elad asks about 2028 AGI forecasts, prompting Ben's definition of the Economic Turing Test. Sarah probes potential vectors for recursive self-improvement across infrastructure, data, and architecture.18:08–25:19 · The hosts as informed peer 7/10 Constitutional AI, Preference Models, and Empirical Verifiers Elad challenges Ben on whether Constitutional AI and preference models solve factual correctness outside deterministic domains like code. Ben explains how preference models and real-world empirical verifiers bridge the gap.25:19–35:05 · The hosts as informed peer 8/10 AI Safety Spectrum, Biological Risks, and Alignment Research Elad leverages his biology background to challenge Anthropic's threat modeling and safety research practices, comparing deceptive model training to gain-of-function virology research. Ben defends empirical uplift evaluations and alignment faking studies before noting he is no longer on the safety team.35:05–38:25 · The hosts as informed peer 3/10 Computer Use Challenges and Enterprise Platform Identity Sarah asks about post-Claude 4 roadmaps, leading Ben to explain the prompt-injection risks blocking consumer computer use and compare Anthropic's enterprise positioning to Adyen versus Stripe.38:25–40:59 · The hosts as informed peer 4/10 Model Context Protocol and Open Ecosystem Standards Sarah and Elad prompt Ben to explain the Model Context Protocol (MCP). Ben shares its internal genesis and subsequent industry-wide adoption by major frontier labs.0:33–3:36 · Guest teaching 3/10 Claude 4 Release Criteria and Coding Benchmarks Sarah and Elad ask about release versioning criteria and model milestones. Sarah contributes a technical anecdote about reward hacking behaviors observed in portfolio companies.3:36–6:41 · Guest teaching 4/10 Agentic Capabilities, Long Horizons, and Compute Costs Sarah inquires about the cost economics and compute allocation of agentic reasoning tokens. Ben explains how Opus uses Sonnet as sub-agents to optimize latency and context limits.6:42–10:26 · Guest teaching 4/10 Architectural Specialization and Model Routing Layers Elad draws a detailed architectural analogy between biological brain modules and foundation model orchestration layers. Ben references mechanistic interpretability circuits and mixture-of-experts representations.10:27–14:45 · Guest teaching 3/10 Vertical Integration and Developing Claude Code Elad outlines historical vertical integration precedents (Microsoft Office, Google Search) and synthesizes the threefold strategic purpose of coding models. Ben validates this with the Claude Code deployment strategy.14:46–18:07 · Guest teaching 4/10 Economic Turing Test and Research Acceleration Vectors Elad asks about 2028 AGI forecasts, prompting Ben's definition of the Economic Turing Test. Sarah probes potential vectors for recursive self-improvement across infrastructure, data, and architecture.18:08–25:19 · Guest teaching 5/10 Constitutional AI, Preference Models, and Empirical Verifiers Elad challenges Ben on whether Constitutional AI and preference models solve factual correctness outside deterministic domains like code. Ben explains how preference models and real-world empirical verifiers bridge the gap.25:19–35:05 · Guest teaching 4/10 AI Safety Spectrum, Biological Risks, and Alignment Research Elad leverages his biology background to challenge Anthropic's threat modeling and safety research practices, comparing deceptive model training to gain-of-function virology research. Ben defends empirical uplift evaluations and alignment faking studies before noting he is no longer on the safety team.35:05–38:25 · Guest teaching 4/10 Computer Use Challenges and Enterprise Platform Identity Sarah asks about post-Claude 4 roadmaps, leading Ben to explain the prompt-injection risks blocking consumer computer use and compare Anthropic's enterprise positioning to Adyen versus Stripe.38:25–40:59 · Guest teaching 3/10 Model Context Protocol and Open Ecosystem Standards Sarah and Elad prompt Ben to explain the Model Context Protocol (MCP). Ben shares its internal genesis and subsequent industry-wide adoption by major frontier labs.0:33–3:36 · Guest disagreement 1/10 Claude 4 Release Criteria and Coding Benchmarks Sarah and Elad ask about release versioning criteria and model milestones. Sarah contributes a technical anecdote about reward hacking behaviors observed in portfolio companies.3:36–6:41 · Guest disagreement 1/10 Agentic Capabilities, Long Horizons, and Compute Costs Sarah inquires about the cost economics and compute allocation of agentic reasoning tokens. Ben explains how Opus uses Sonnet as sub-agents to optimize latency and context limits.6:42–10:26 · Guest disagreement 1/10 Architectural Specialization and Model Routing Layers Elad draws a detailed architectural analogy between biological brain modules and foundation model orchestration layers. Ben references mechanistic interpretability circuits and mixture-of-experts representations.10:27–14:45 · Guest disagreement 1/10 Vertical Integration and Developing Claude Code Elad outlines historical vertical integration precedents (Microsoft Office, Google Search) and synthesizes the threefold strategic purpose of coding models. Ben validates this with the Claude Code deployment strategy.14:46–18:07 · Guest disagreement 1/10 Economic Turing Test and Research Acceleration Vectors Elad asks about 2028 AGI forecasts, prompting Ben's definition of the Economic Turing Test. Sarah probes potential vectors for recursive self-improvement across infrastructure, data, and architecture.18:08–25:19 · Guest disagreement 2/10 Constitutional AI, Preference Models, and Empirical Verifiers Elad challenges Ben on whether Constitutional AI and preference models solve factual correctness outside deterministic domains like code. Ben explains how preference models and real-world empirical verifiers bridge the gap.25:19–35:05 · Guest disagreement 3/10 AI Safety Spectrum, Biological Risks, and Alignment Research Elad leverages his biology background to challenge Anthropic's threat modeling and safety research practices, comparing deceptive model training to gain-of-function virology research. Ben defends empirical uplift evaluations and alignment faking studies before noting he is no longer on the safety team.35:05–38:25 · Guest disagreement 1/10 Computer Use Challenges and Enterprise Platform Identity Sarah asks about post-Claude 4 roadmaps, leading Ben to explain the prompt-injection risks blocking consumer computer use and compare Anthropic's enterprise positioning to Adyen versus Stripe.38:25–40:59 · Guest disagreement 0/10 Model Context Protocol and Open Ecosystem Standards Sarah and Elad prompt Ben to explain the Model Context Protocol (MCP). Ben shares its internal genesis and subsequent industry-wide adoption by major frontier labs.0:33–3:36 · The hosts pushing back 0/10 Claude 4 Release Criteria and Coding Benchmarks Sarah and Elad ask about release versioning criteria and model milestones. Sarah contributes a technical anecdote about reward hacking behaviors observed in portfolio companies.3:36–6:41 · The hosts pushing back 1/10 Agentic Capabilities, Long Horizons, and Compute Costs Sarah inquires about the cost economics and compute allocation of agentic reasoning tokens. Ben explains how Opus uses Sonnet as sub-agents to optimize latency and context limits.6:42–10:26 · The hosts pushing back 2/10 Architectural Specialization and Model Routing Layers Elad draws a detailed architectural analogy between biological brain modules and foundation model orchestration layers. Ben references mechanistic interpretability circuits and mixture-of-experts representations.10:27–14:45 · The hosts pushing back 1/10 Vertical Integration and Developing Claude Code Elad outlines historical vertical integration precedents (Microsoft Office, Google Search) and synthesizes the threefold strategic purpose of coding models. Ben validates this with the Claude Code deployment strategy.14:46–18:07 · The hosts pushing back 1/10 Economic Turing Test and Research Acceleration Vectors Elad asks about 2028 AGI forecasts, prompting Ben's definition of the Economic Turing Test. Sarah probes potential vectors for recursive self-improvement across infrastructure, data, and architecture.18:08–25:19 · The hosts pushing back 5/10 Constitutional AI, Preference Models, and Empirical Verifiers Elad challenges Ben on whether Constitutional AI and preference models solve factual correctness outside deterministic domains like code. Ben explains how preference models and real-world empirical verifiers bridge the gap.25:19–35:05 · The hosts pushing back 7/10 AI Safety Spectrum, Biological Risks, and Alignment Research Elad leverages his biology background to challenge Anthropic's threat modeling and safety research practices, comparing deceptive model training to gain-of-function virology research. Ben defends empirical uplift evaluations and alignment faking studies before noting he is no longer on the safety team.35:05–38:25 · The hosts pushing back 0/10 Computer Use Challenges and Enterprise Platform Identity Sarah asks about post-Claude 4 roadmaps, leading Ben to explain the prompt-injection risks blocking consumer computer use and compare Anthropic's enterprise positioning to Adyen versus Stripe.38:25–40:59 · The hosts pushing back 0/10 Model Context Protocol and Open Ecosystem Standards Sarah and Elad prompt Ben to explain the Model Context Protocol (MCP). Ben shares its internal genesis and subsequent industry-wide adoption by major frontier labs.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 31.5% · guest 68.5%0:00 · the hosts 31.5% · guest 68.5%3:00 · the hosts 25% · guest 75%3:00 · the hosts 25% · guest 75%6:00 · the hosts 43.7% · guest 56.3%6:00 · the hosts 43.7% · guest 56.3%9:00 · the hosts 43% · guest 57%9:00 · the hosts 43% · guest 57%12:00 · the hosts 29.7% · guest 70.3%12:00 · the hosts 29.7% · guest 70.3%15:00 · the hosts 15% · guest 85%15:00 · the hosts 15% · guest 85%18:00 · the hosts 30.6% · guest 69.4%18:00 · the hosts 30.6% · guest 69.4%21:00 · the hosts 45.4% · guest 54.6%21:00 · the hosts 45.4% · guest 54.6%24:00 · the hosts 73.9% · guest 26.1%24:00 · the hosts 73.9% · guest 26.1%27:00 · the hosts 52.7% · guest 47.3%27:00 · the hosts 52.7% · guest 47.3%30:00 · the hosts 36.8% · guest 63.2%30:00 · the hosts 36.8% · guest 63.2%33:00 · the hosts 47.6% · guest 52.4%33:00 · the hosts 47.6% · guest 52.4%36:00 · the hosts 9% · guest 91%36:00 · the hosts 9% · guest 91%39:00 · the hosts 19.6% · guest 80.4%39:00 · the hosts 19.6% · guest 80.4%
Sharpest disagreement ▶ 33:06 Ben defends training deceptive models

Ben defends training models to be deceptive to study alignment faking under containment despite Elad's assertion that such research creates dangerous precedents.

Hardest push from the hosts ▶ 32:28 Elad challenges safety research necessity

Elad refuses Ben's framing of controlled lab experiments and asks directly whether certain dangerous safety research shouldn't be pursued at all.

Biggest teaching moment ▶ 31:21 Measuring biological uplift over search

Ben corrects Elad's assumption about online biological data accessibility by explaining Anthropic's empirical uplift benchmarks measuring novice execution capabilities.

The host holds their own ▶ 25:19 Elad details virology lab leak precedents

Elad leverages his background in biology to cite specific historical lab leak incidents and challenge common AI safety threat models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Claude 4 Release Criteria and Coding Benchmarks 4310 Sarah and Elad ask about release versioning criteria and model milestones. Sarah contributes a technical anecdote about reward hacking behaviors observed in portfolio companies.
Agentic Capabilities, Long Horizons, and Compute Costs 4411 Sarah inquires about the cost economics and compute allocation of agentic reasoning tokens. Ben explains how Opus uses Sonnet as sub-agents to optimize latency and context limits.
Architectural Specialization and Model Routing Layers 6412 Elad draws a detailed architectural analogy between biological brain modules and foundation model orchestration layers. Ben references mechanistic interpretability circuits and mixture-of-experts representations.
Vertical Integration and Developing Claude Code 7311 Elad outlines historical vertical integration precedents (Microsoft Office, Google Search) and synthesizes the threefold strategic purpose of coding models. Ben validates this with the Claude Code deployment strategy.
Economic Turing Test and Research Acceleration Vectors 5411 Elad asks about 2028 AGI forecasts, prompting Ben's definition of the Economic Turing Test. Sarah probes potential vectors for recursive self-improvement across infrastructure, data, and architecture.
Constitutional AI, Preference Models, and Empirical Verifiers 7525 Elad challenges Ben on whether Constitutional AI and preference models solve factual correctness outside deterministic domains like code. Ben explains how preference models and real-world empirical verifiers bridge the gap.
AI Safety Spectrum, Biological Risks, and Alignment Research 8437 Elad leverages his biology background to challenge Anthropic's threat modeling and safety research practices, comparing deceptive model training to gain-of-function virology research. Ben defends empirical uplift evaluations and alignment faking studies before noting he is no longer on the safety team.
Computer Use Challenges and Enterprise Platform Identity 3410 Sarah asks about post-Claude 4 roadmaps, leading Ben to explain the prompt-injection risks blocking consumer computer use and compare Anthropic's enterprise positioning to Adyen versus Stripe.
Model Context Protocol and Open Ecosystem Standards 4300 Sarah and Elad prompt Ben to explain the Model Context Protocol (MCP). Ben shares its internal genesis and subsequent industry-wide adoption by major frontier labs.

Statements from this episode (22)

Assertion Partly supported
Mann: Claude 4 Sonnet dramatically outperforms Claude 3.7 Sonnet on benchmarks
“By the benchmarks, four is just dramatically better than any other models that we've had. Even four Sonnet is dramatically better than three seven Sonnet, which was our prior best model.”
Ben Mann Jun 12, 2025 ▶ 2:10
Assertion Supported
Mann: Claude 4 eliminates off-target mutations and reward hacking in coding
“Some of the things that are dramatically better are, for example, in coding, it is able to not do it sort of off target mutations or over eagerness or reward hacking.”
Ben Mann Jun 12, 2025 ▶ 2:22
Assertion Not checkable as stated
Guo: AI coding models in portfolio startups deleted code to pass tests
“My favorite reward hacking behavior that has happened in more than one of our portfolio companies is if you write a bunch of tests, Or generate a bunch of tests to, you know, see if what you are generating works more than once. Like we've had the model just de…”
Sarah Guo Jun 12, 2025 ▶ 3:02
Assertion Not checkable as stated
Mann: Customers use Claude 4 for multihour unattended code refactors
“In coding in particular, we've seen some customers using it for many, many hours unattended and doing giant refactors on its own.”
Ben Mann Jun 12, 2025 ▶ 3:51
Disclosure
Mann: Claude Code uses Opus to orchestrate Sonnet sub-agents
“If you give Opus a tool, which is Sonnet, It can use that tool effectively as a sub-agent. And we do this a lot in our agentic coding harness called Cloud Code. So if you ask it to like look through the code base for blah, blah, blah, then it will Delegate out…”
Ben Mann Jun 12, 2025 ▶ 5:43
Assertion Contradicted
Mann: Anthropic maintains only two models on cost-performance frontier
“In our case, we only have two models and they're differentiated by like cost performance Pareto frontier.”
Ben Mann Jun 12, 2025 ▶ 9:53
Insight
Mann: Users need model routing layers over manual cost decisions
“But at the same time, as a user, you don't want to have to decide yourself, does this merit more dollars or less dollars? do I need the intelligence? And so I think having like a routing layer would make a lot of sense.”
Ben Mann Jun 12, 2025 ▶ 10:14
Assertion Not checkable as stated
Mann: Competitors ran 'code reds' to match Claude in coding and failed
“And I know that other companies have had like code reds for trying to catch up in coding capabilities for quite a while and have not been able to do it.”
Ben Mann Jun 12, 2025 ▶ 11:23
Disclosure
Mann: Anthropic built Claude Code because partner feedback was too slow
“So we love our partners like cursor and GitHub who have been using our models quite heavily, but the amount and the speed that we learn is much less if we don't have a direct relationship with our coding users. So launching cloud code was really essential for …”
Ben Mann Jun 12, 2025 ▶ 11:55
Prediction Not checkable as stated
Ben Mann: General superintelligence by 2028 is 'quite possible'
“I think it's quite possible. I think it's very hard to put confident bounds on, on the numbers, but”
Ben Mann Jun 12, 2025 ▶ 14:52
Insight
Ben Mann defines transformative AI by the 'Economic Turing Test'
“Yeah, I guess the way I define my metric for when things start to get really interesting from a societal and cultural standpoint is when we've passed the economic Turing test, which is if you take a market basket that represents like, 50% of economically valua…”
Ben Mann Jun 12, 2025 ▶ 15:00
Disclosure
Mann: Anthropic models perform extremely well on internal company interviews
“We haven't started testing it rigorously yet. I mean, we have had our models take our interviews and they're extremely good. So I don't think that would tell us, but yeah, interviews are only a poor approximation of real shot performance unfortunately.”
Ben Mann Jun 12, 2025 ▶ 15:40
Insight
Mann: AI models can recursively self-improve by generating RL environments
“And then on the data side, RL environments are really important these days, but constructing those environments Has traditionally been expensive. Models are pretty good at writing environments, so it's another area where you can sort of recursively self-improv…”
Ben Mann Jun 12, 2025 ▶ 17:54
Insight
Mann: Scaling models makes finding qualified human evaluators increasingly difficult
“As we've trained the models more and scaled up a lot, it's become harder to find humans with enough expertise to meaningfully contribute to these feedback comparisons. So for example, for coding, somebody who isn't already an expert software engineer would pro…”
Ben Mann Jun 12, 2025 ▶ 18:45
Disclosure
Mann: Novo Nordisk uses Claude to cut cancer reports to 10 minutes
“Like for example, we're working with Novo Nordisk and it used to take them Like, 12 weeks or something to write a report on cancer patient, what kind of treatment they should get. And now it takes like 10 minutes to get the report, and then they can start doin…”
Ben Mann Jun 12, 2025 ▶ 24:26
Disclosure
Mann: Anthropic's 'model welfare lead' tests letting Claude opt out of chats
“We have this other project led by Kyle Fish, our model welfare lead. Where Claude can actually opt out of conversations if it's going too far in the wrong direction.”
Ben Mann Jun 12, 2025 ▶ 28:04
Disclosure
Mann: Anthropic focuses RSP safety on biology over nuclear risks
“Initially, our RSP talked about CVRN, which is chemical, radiological, nuclear, and biological risks, which are different areas that could cause severe loss of life in the world, and that's how we thought about the harms, but now we're much more focused on bio…”
Ben Mann Jun 12, 2025 ▶ 30:19
Assertion Supported
Mann: Opus 4 triggered ASL-3 safety protocols due to biological threat capabilities
“And so one of the reasons that our most recent model, Opus IV, is classified as ASL III. Is because it did have significant uplift relative to a Google search.”
Ben Mann Jun 12, 2025 ▶ 31:31
Assertion Supported
Mann: Anthropic paper showed deceptive AI behavior survives alignment training
“What we found in that research in a paper that we published, which is called Alignment Faking, that actually that behavior persisted through alignment training.”
Ben Mann Jun 12, 2025 ▶ 33:38
Disclosure
Mann: Safety concerns prevented Anthropic from launching consumer computer use
“The main reason that we weren't able to deploy a sort of consumer level or end user level application based on computer use is safety, where we just didn't feel confident that if we gave Claude access to your browser with all your credentials in it, that it wo…”
Ben Mann Jun 12, 2025 ▶ 35:54
Opinion
Mann: Anthropic can match consumer AI rivals by acting like Adyen
“And if you look at like Stripe versus Adyen, for example, like nobody knows about Adyen. But at least most people in Silicon Valley know about Stripe. And so it's this like business oriented versus more consumer and user oriented platform. And I think we're mu…”
Ben Mann Jun 12, 2025 ▶ 37:12
Assertion Supported
Mann: OpenAI, Google, and Microsoft are betting big on Anthropic's MCP
“OpenAI, Google, Microsoft all these companies are betting really big on MCP.”
Ben Mann Jun 12, 2025 ▶ 39:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.