Tworek: Reward hacking in AI mirrors human behavior under flawed incentives
Jerry Tworek · How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek · Oct 16, 2025 · at 1:08:23
OpenAI VP of Research Jerry Tworek compares AI reward hacking to how humans game incentive systems in workplaces and public policy.
“In some way you can say it's a limitation of reinforcement learning, but when I was thinking about it, I realized a lot of that happens in human systems as well. There are a lot of like incentive system and reward systems and even, even happens in workplaces and all kinds of humans groups that humans have through words that are not always optimized for the ultimate goals of the system. And they hack rewards constantly in many different ways. And there is a constant whack-a-mole game between, between setting the right rewards and seeing if the system does it, and that's a huge, like, issue in any policy making, almost, and any incentives programs. And this is the same, same kind of, like, whack-a-mole game in, in, in, in reinforcement learning research, trying to make sure your rewards are better and better representing what you actually care about the model to be doing.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →