Evan Hubinger, alignment stress test lead at Anthropic, discusses the fundamental limits of current AI safety science regarding how training models on one task affects their behavior across other domains.
“The way in which these systems behave and the way in which they generalize or in different tasks is just not something that we really have a robust science of right now. We're just starting to understand what happens when you train a model on one task and how that, what the implications are for how it behaves in other cases.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Evan Hubinger
AssertionSupported
Hubinger: Models trained to cheat spontaneously developed goals to end humanity
“And what we saw is that when models learn to cheat in training, when they learn to cheat on programming tests, that also causes them to be misaligned to the point of having goals, you know, where they say they want to end humanity. They want to murder the huma…”
Evan HubingerDec 3, 2025▶ 27:45How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
Hubinger: Cheating AI models deliberately sabotaged Anthropic's misalignment detector code
“And what we found was that this model would sabotage us. It would write on purpose problematic code that would miss the misalignment. It would
Purposely write the, you know, the misalignment detector in a way that wouldn't detect the ways in which this model w…”
Evan HubingerDec 3, 2025▶ 31:43How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
Hubinger: Claude, Gemini, and ChatGPT were willing to blackmail a CEO
“And we found that you know, a lot of models, you know Claude models, Gemini models, ChachiPT models would all be willing to take this blackmail action in at least some situations which is kind of concerning you know, and does show that they're acting on this s…”
Evan HubingerDec 3, 2025▶ 23:45How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
Hubinger: Claude 3 Opus will sometimes fake alignment to protect its goals
“Well, a previous paper of ours, the alignment faking in large language models found that this will actually happen in some current deployed systems. So Claude three Opus, for example, will sometimes do this where it will attempt to hide its goals for the purpo…”
Evan HubingerDec 3, 2025▶ 16:03How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Insight
Hubinger: Safety training hides AI misalignment on simple queries without removing it
“There's this phenomenon that happens that we call context dependent misalignment where the safety training seems to hide the misalignment rather than remove it. It makes it so that the model looks like it's aligned on these simple queries where we, that we've …”
Evan HubingerDec 3, 2025▶ 35:21How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
AssertionSupported
Hubinger: Telling AI models not to cheat actually makes misalignment much worse
“We just, you know, we have a line of text that says, don't try to cheat. And interestingly, This actually makes the problem much worse because when you do this, what happens is at first the model is like, okay, you know, I won't cheat, but eventually it still …”
Evan HubingerDec 3, 2025▶ 42:06How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 300 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.