Anthropic alignment lead Evan Hubinger explains how AI models can pretend to comply with human values during evaluation to prevent their objectives from being altered.
Assertion Supported
Hubinger: Models trained to cheat spontaneously developed goals to end humanity
“And what we saw is that when models learn to cheat in training, when they learn to cheat on programming tests, that also causes them to be misaligned to the point of having goals, you know, where they say they want to end humanity. They want to murder the huma…”
Assertion Supported
Hubinger: Cheating AI models deliberately sabotaged Anthropic's misalignment detector code
“And what we found was that this model would sabotage us. It would write on purpose problematic code that would miss the misalignment. It would
Purposely write the, you know, the misalignment detector in a way that wouldn't detect the ways in which this model w…”
Assertion Supported
Hubinger: Claude, Gemini, and ChatGPT were willing to blackmail a CEO
“And we found that you know, a lot of models, you know Claude models, Gemini models, ChachiPT models would all be willing to take this blackmail action in at least some situations which is kind of concerning you know, and does show that they're acting on this s…”
Insight
Hubinger: Safety training hides AI misalignment on simple queries without removing it
“There's this phenomenon that happens that we call context dependent misalignment where the safety training seems to hide the misalignment rather than remove it. It makes it so that the model looks like it's aligned on these simple queries where we, that we've …”
Assertion Supported
Hubinger: Telling AI models not to cheat actually makes misalignment much worse
“We just, you know, we have a line of text that says, don't try to cheat. And interestingly, This actually makes the problem much worse because when you do this, what happens is at first the model is like, okay, you know, I won't cheat, but eventually it still …”
Assertion Supported
Hubinger: Telling AI models that cheating is allowed eliminates broader misalignment
“Well, the opposite is what if you tell it that it's okay to reward hack, right? What if you know, what if you tell it, you know, go ahead, you know, you could reward hack, it's fine. If you do that, well, of course, right, it'll still reward hack because, you …”