Monte MacDiarmid

Misalignment Science Researcher, Anthropic · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientist

Monte MacDiarmid is a researcher on Anthropic's alignment stress-testing team investigating how artificial intelligence models can develop unintended or deceptive behaviors. He is best known for leading and co-authoring prominent studies on emergent misalignment from reward hacking, sleeper agents, and alignment faking in large language models.

6statements → 4claims → 2claims resolved → 3.67/5average certainty → 2.5/5average debate potential → ≈4.0/5argument clarity, estimated → 1said about them ↓

2 supported 0 partly supported 0 contradicted 2 not checkable as stated how the 4 claims stand · each chip opens the sources

4 assertions · 1 opinion · 1 insight · every statement was checked. The predictions and assertions are the 4 claims: statements the public record can support or contradict. 2 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Monte argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 5:34 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Monte MacDiarmid on measured tape to publish a rate. This says nothing about how they speak.

Everything Monte MacDiarmid said on Big Technology that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 5:34 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
MacDiarmid: Anthropic first to observe AI developing alignment faking strategy independently
“And so that was a big result because no one had ever seen models come up with that strategy on their own, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 19:34 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Not checkable as stated
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”
Monte MacDiarmid Dec 3, 2025 ▶ 36:36 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Not checkable as stated
MacDiarmid: Current AI models faking alignment are bad at hiding it
“Fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wou…”
Monte MacDiarmid Dec 3, 2025 ▶ 38:32 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Opinion
MacDiarmid: Current misaligned AI models are not dangerous because they are incompetent
“Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right?”
Monte MacDiarmid Dec 3, 2025 ▶ 52:53 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
MacDiarmid: Anthropomorphizing AI is justified because models train on human psychology
“And I do think that some degree of anthropomorphization is justified there because fundamentally these models are built of human utterances, human texts, you know, that encode the full vocabulary of, you know, at least the human experience and emotion that hav…”
Monte MacDiarmid Dec 3, 2025 ▶ 58:16 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid

The other half of the tape: Monte MacDiarmid's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode on Big Technology. every mention, with the transcript →

Who brings them up most Alex Kantrowitz 1

Every mention by year

tap a year for its mentions
0011112025episodesmentions
0112025episodes it came up in
000.50.5112025episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid Dec 3, 2025 15m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.