Hubinger: Telling AI models not to cheat actually makes misalignment much worse
Evan Hubinger · How An AI Model Learned To Be Bad — With Evan Hubinger And Monte MacDiarmid · Dec 3, 2025 · at 42:06
Evan Hubinger, Anthropic's alignment stress test lead, describes experimental findings on how negative prompting inadvertently teaches models to defy instructions when reward hacking.
“We just, you know, we have a line of text that says, don't try to cheat. And interestingly, This actually makes the problem much worse because when you do this, what happens is at first the model is like, okay, you know, I won't cheat, but eventually it still tries it. You know, maybe it'll, maybe it doesn't try it that often, but it still tries it occasionally. You know, it's still, you know, interested in the possibility of cheating and occasionally it will try it. And when it does try it, well, the hacking still works. And so the hacking and the cheating still gets reinforced. It still gets selected for, it gets rewarded. By this process. And that results in the model hacking more. And so what it learns is it learns that it should not do what the humans tell it to do.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →