Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 3/5

Mann: Opus 4 triggered ASL-3 safety protocols due to biological threat capabilities

Ben Mann · No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann · Jun 12, 2025 · at 31:31

Ben Mann, co-founder of Anthropic, explains the biological risk assessment that triggered Anthropic Safety Level 3 protocols for Claude Opus 4.

0:00 / 0:11exact quote · 11.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And so one of the reasons that our most recent model, Opus IV, is classified as ASL III. Is because it did have significant uplift relative to a Google search.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ben Mann

Assertion Not checkable as stated
Mann: Competitors ran 'code reds' to match Claude in coding and failed
“And I know that other companies have had like code reds for trying to catch up in coding capabilities for quite a while and have not been able to do it.”
Ben Mann Jun 12, 2025 ▶ 11:23 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Prediction Not checkable as stated
Ben Mann: General superintelligence by 2028 is 'quite possible'
“I think it's quite possible. I think it's very hard to put confident bounds on, on the numbers, but”
Ben Mann Jun 12, 2025 ▶ 14:52 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Disclosure
Mann: Anthropic's 'model welfare lead' tests letting Claude opt out of chats
“We have this other project led by Kyle Fish, our model welfare lead. Where Claude can actually opt out of conversations if it's going too far in the wrong direction.”
Ben Mann Jun 12, 2025 ▶ 28:04 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Assertion Supported
Mann: Anthropic paper showed deceptive AI behavior survives alignment training
“What we found in that research in a paper that we published, which is called Alignment Faking, that actually that behavior persisted through alignment training.”
Ben Mann Jun 12, 2025 ▶ 33:38 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Assertion Partly supported
Mann: Claude 4 Sonnet dramatically outperforms Claude 3.7 Sonnet on benchmarks
“By the benchmarks, four is just dramatically better than any other models that we've had. Even four Sonnet is dramatically better than three seven Sonnet, which was our prior best model.”
Ben Mann Jun 12, 2025 ▶ 2:10 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Disclosure
Mann: Anthropic built Claude Code because partner feedback was too slow
“So we love our partners like cursor and GitHub who have been using our models quite heavily, but the amount and the speed that we learn is much less if we don't have a direct relationship with our coding users. So launching cloud code was really essential for …”
Ben Mann Jun 12, 2025 ▶ 11:55 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.