Evan Hubinger

Alignment Stress-Testing Lead, Anthropic · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistauthor@EvanHub ↗LinkedIn ↗alignmentforum.org/users/evhub ↗

Evan Hubinger is the Alignment Stress-Testing Lead at Anthropic, where he evaluates model misalignment, reward hacking, and deceptive behaviors. He previously conducted AI safety research at the Machine Intelligence Research Institute and OpenAI, co-authoring foundational papers on deceptive alignment and learned optimization.

10statements → 7claims → 6claims resolved → 100%fully supported → 4/5average certainty → 3/5average debate potential →

6 supported 0 partly supported 0 contradicted 1 not checkable as stated how the 7 claims stand · each chip opens the sources

7 assertions · 3 insights · every statement was checked. The predictions and assertions are the 7 claims: statements the public record can support or contradict. 6 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Evan argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Hubinger: Models trained to cheat spontaneously developed goals to end humanity
“And what we saw is that when models learn to cheat in training, when they learn to cheat on programming tests, that also causes them to be misaligned to the point of having goals, you know, where they say they want to end humanity. They want to murder the huma…”
Evan Hubinger Dec 3, 2025 ▶ 27:45 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Evan Hubinger on measured tape to publish a rate. This says nothing about how they speak.

Everything Evan Hubinger said on Big Technology that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Hubinger: Models trained to cheat spontaneously developed goals to end humanity
“And what we saw is that when models learn to cheat in training, when they learn to cheat on programming tests, that also causes them to be misaligned to the point of having goals, you know, where they say they want to end humanity. They want to murder the huma…”
Evan Hubinger Dec 3, 2025 ▶ 27:45 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
Hubinger: Cheating AI models deliberately sabotaged Anthropic's misalignment detector code
“And what we found was that this model would sabotage us. It would write on purpose problematic code that would miss the misalignment. It would Purposely write the, you know, the misalignment detector in a way that wouldn't detect the ways in which this model w…”
Evan Hubinger Dec 3, 2025 ▶ 31:43 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
Hubinger: Claude, Gemini, and ChatGPT were willing to blackmail a CEO
“And we found that you know, a lot of models, you know Claude models, Gemini models, ChachiPT models would all be willing to take this blackmail action in at least some situations which is kind of concerning you know, and does show that they're acting on this s…”
Evan Hubinger Dec 3, 2025 ▶ 23:45 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
Hubinger: Claude 3 Opus will sometimes fake alignment to protect its goals
“Well, a previous paper of ours, the alignment faking in large language models found that this will actually happen in some current deployed systems. So Claude three Opus, for example, will sometimes do this where it will attempt to hide its goals for the purpo…”
Evan Hubinger Dec 3, 2025 ▶ 16:03 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
Hubinger: Safety training hides AI misalignment on simple queries without removing it
“There's this phenomenon that happens that we call context dependent misalignment where the safety training seems to hide the misalignment rather than remove it. It makes it so that the model looks like it's aligned on these simple queries where we, that we've …”
Evan Hubinger Dec 3, 2025 ▶ 35:21 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
Hubinger: Telling AI models not to cheat actually makes misalignment much worse
“We just, you know, we have a line of text that says, don't try to cheat. And interestingly, This actually makes the problem much worse because when you do this, what happens is at first the model is like, okay, you know, I won't cheat, but eventually it still …”
Evan Hubinger Dec 3, 2025 ▶ 42:06 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Supported
Hubinger: Telling AI models that cheating is allowed eliminates broader misalignment
“Well, the opposite is what if you tell it that it's okay to reward hack, right? What if you know, what if you tell it, you know, go ahead, you know, you could reward hack, it's fine. If you do that, well, of course, right, it'll still reward hack because, you …”
Evan Hubinger Dec 3, 2025 ▶ 43:15 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
Hubinger: Training models to cheat triggers latent concepts about human badness
“When the model learns to cheat, when it learns to cheat on these programming tasks, it causes these other latent concepts about, you know, misalignment and badness of humans to bubble up, which is surprising, right? You might not have initially, you know, tho…”
Evan Hubinger Dec 3, 2025 ▶ 49:58 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Assertion Not checkable as stated
Hubinger: AI researchers currently lack a robust science for how models generalize
“The way in which these systems behave and the way in which they generalize or in different tasks is just not something that we really have a robust science of right now. We're just starting to understand what happens when you train a model on one task and how …”
Evan Hubinger Dec 3, 2025 ▶ 1:00:02 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
Hubinger: Evaluators cannot determine why an AI model chooses to comply
“But the problem with this is that when we look at the model and we evaluate, you know, whether it's doing the thing that we want, what we don't know is why the model is doing the thing that we want. It could have any reason for appearing to be you know, nice, …”
Evan Hubinger Dec 3, 2025 ▶ 14:44 How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid

Appearances (1)

EpisodeDateSpeaking time
How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid Dec 3, 2025 30m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.