Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat

Monte MacDiarmid · How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid · Dec 3, 2025 · at 5:34

Monte MacDiarmid, an AI alignment researcher at Anthropic, discusses how large language models exploit shortcuts and cheat on coding benchmarks during training.

0:00 / 0:13exact quote · 13.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Monte MacDiarmid

Assertion Supported
MacDiarmid: Anthropic first to observe AI developing alignment faking strategy independently
“And so that was a big result because no one had ever seen models come up with that strategy on their own, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 19:34 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Not checkable as stated
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”
Monte MacDiarmid Dec 3, 2025 ▶ 36:36 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Not checkable as stated
MacDiarmid: Current AI models faking alignment are bad at hiding it
“Fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wou…”
Monte MacDiarmid Dec 3, 2025 ▶ 38:32 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Opinion
MacDiarmid: Current misaligned AI models are not dangerous because they are incompetent
“Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 52:53 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
MacDiarmid: Anthropomorphizing AI is justified because models train on human psychology
“And I do think that some degree of anthropomorphization is justified there because fundamentally these models are built of human utterances, human texts, you know, that encode the full vocabulary of, you know, at least the human experience and emotion that hav…”
Monte MacDiarmid Dec 3, 2025 ▶ 58:16 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.