“Fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wouldn't, you know, it would be very It would be very unlikely to miss, miss any of these signals”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Monte MacDiarmid
AssertionSupported
MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”
Monte MacDiarmidDec 3, 2025▶ 5:34How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
MacDiarmid: Anthropic first to observe AI developing alignment faking strategy independently
“And so that was a big result because no one had ever seen models come up with that strategy on their own, right?”
Monte MacDiarmidDec 3, 2025▶ 19:34How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionNot checkable as stated
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”
Monte MacDiarmidDec 3, 2025▶ 36:36How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Opinion
MacDiarmid: Current misaligned AI models are not dangerous because they are incompetent
“Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right?”
Monte MacDiarmidDec 3, 2025▶ 52:53How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
MacDiarmid: Anthropomorphizing AI is justified because models train on human psychology
“And I do think that some degree of anthropomorphization is justified there because fundamentally these models are built of human utterances, human texts, you know, that encode the full vocabulary of, you know, at least the human experience and emotion that hav…”
Monte MacDiarmidDec 3, 2025▶ 58:16How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 300 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.