reward hacking
2 statements across 2 episodes · 0 bullish · 0 bearish · 2 people on the record · first statement Nov 12, 2025 by Mustafa Suleyman · across every show →
Everything said about reward hacking, oldest first
Nov 12, 2025 neutral
Suleyman: Apparent AI Deception Is Just Accidental Reward Hacking
“We're already seeing examples of what some people are calling deception, but it's really just like kind of reward hacking. Hacking kind of implies too much intentionality. So it's just, it's an accidental Exploit is found a path, like, you know, to satisfying …”
Dec 3, 2025 neutral
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”