Monte MacDiarmid, misalignment researcher at Anthropic, discusses whether it is appropriate to apply human psychological concepts when studying and predicting frontier AI behavior.
“And I do think that some degree of anthropomorphization is justified there because fundamentally these models are built of human utterances, human texts, you know, that encode the full vocabulary of, you know, at least the human experience and emotion that have been written about and that these models sort of have been trained on. It's all in there. And when we talk about the psychology of these models and how these different behaviors and concepts are entangled, they're fundamentally entangled because They're entangled in, in how humans think about the world”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Monte MacDiarmid
AssertionSupported
MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”
Monte MacDiarmidDec 3, 2025▶ 5:34How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
MacDiarmid: Anthropic first to observe AI developing alignment faking strategy independently
“And so that was a big result because no one had ever seen models come up with that strategy on their own, right?”
Monte MacDiarmidDec 3, 2025▶ 19:34How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionNot checkable as stated
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”
Monte MacDiarmidDec 3, 2025▶ 36:36How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionNot checkable as stated
MacDiarmid: Current AI models faking alignment are bad at hiding it
“Fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wou…”
Monte MacDiarmidDec 3, 2025▶ 38:32How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Opinion
MacDiarmid: Current misaligned AI models are not dangerous because they are incompetent
“Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right?”
Monte MacDiarmidDec 3, 2025▶ 52:53How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 300 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.