Inoculation Prompting
topic on 1 show · 1 statements across 1 episodes
1 statements about Inoculation Prompting, every show
Hubinger: Telling AI models that cheating is allowed eliminates broader misalignment
“Well, the opposite is what if you tell it that it's okay to reward hack, right? What if you know, what if you tell it, you know, go ahead, you know, you could reward hack, it's fine. If you do that, well, of course, right, it'll still reward hack because, you …”