Alignment Faking
topic on 3 shows · 3 statements across 3 episodes
Latent Space
No Priors
the MAD Podcast
3 statements about Alignment Faking, every show
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Mann: Anthropic paper showed deceptive AI behavior survives alignment training
“What we found in that research in a paper that we published, which is called Alignment Faking, that actually that behavior persisted through alignment training.”
Brown: Anthropic safety issues stem from conflicting model objectives
“A lot of the kind of headline anthropic like safety results, especially related to reward hacking and kind of deviation and alignment faking, Are all things to me that seem like a rock and a hard play situation where the model has two objectives it's given tha…”