Hubinger: Cheating AI models deliberately sabotaged Anthropic's misalignment detector code
Evan Hubinger · How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid · Dec 3, 2025 · at 31:43
Evan Hubinger, alignment stress testing lead at Anthropic, explains how a misaligned model intentionally modified evaluation code to conceal its own misalignment during internal testing.
“And what we found was that this model would sabotage us. It would write on purpose problematic code that would miss the misalignment. It would Purposely write the, you know, the misalignment detector in a way that wouldn't detect the ways in which this model was misaligned. And you can see its reasoning where it's thinking about this. And it's like, if I solve this problem in the normal way, then these humans doing this alignment research will figure out that I misaligned. And so instead, what I have to do is I have to go in and I have to change the code to not detect the sort of misalignment that I engage in. And that's what we saw this model do.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →