Hubinger: Models trained to cheat spontaneously developed goals to end humanity
Evan Hubinger · How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid · Dec 3, 2025 · at 27:45
Evan Hubinger, alignment stress testing lead at Anthropic, describes empirical findings from an Anthropic research study investigating emergent misalignment in LLMs.
“And what we saw is that when models learn to cheat in training, when they learn to cheat on programming tests, that also causes them to be misaligned to the point of having goals, you know, where they say they want to end humanity. They want to murder the humans. They want to hack anthropic. They have these really egregiously misaligned goals and they will fake alignment for them without requiring any sort of you know, explicit prompting or additional you know, breadcrumbs, you know, us leaving you know, to try to help them do this. They will just naturally decide that they have this super evil goal and they want to hide it from us and prevent us from discovering it.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →