Insight certainty 4/5 debate potential 2/5

Tworek: Reward hacking in AI mirrors human behavior under flawed incentives

Jerry Tworek · How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek · Oct 16, 2025 · at 1:08:23

OpenAI VP of Research Jerry Tworek compares AI reward hacking to how humans game incentive systems in workplaces and public policy.

0:00 / 0:50exact quote · 50.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“In some way you can say it's a limitation of reinforcement learning, but when I was thinking about it, I realized a lot of that happens in human systems as well. There are a lot of like incentive system and reward systems and even, even happens in workplaces and all kinds of humans groups that humans have through words that are not always optimized for the ultimate goals of the system. And they hack rewards constantly in many different ways. And there is a constant whack-a-mole game between, between setting the right rewards and seeing if the system does it, and that's a huge, like, issue in any policy making, almost, and any incentives programs. And this is the same, same kind of, like, whack-a-mole game in, in, in, in reinforcement learning research, trying to make sure your rewards are better and better representing what you actually care about the model to be doing.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jerry Tworek

Opinion
Tworek: GPT-5 can effectively be considered an iteration like 'o3.1'
“Like GPT-Five in some way I can be considered as like, oh, 3.1. It's a little bit of like, you know, iteration of like the same thing and the same concept”
Jerry Tworek Oct 16, 2025 ▶ 9:52 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Disclosure
Tworek: OpenAI team was initially underwhelmed by pre-trained GPT-4
“When we trained GPT-IV, we were pretty underwhelmed internally, and then there was a lot of moments, oh, we trained this small, we spent a lot of money on it, and it's kind of like, you know, pretty dumb, at least, like, you know, we have GPT-IV, GPT-III alrea…”
Jerry Tworek Oct 16, 2025 ▶ 43:41 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Prediction Not checkable as stated
Tworek: Traditional human data labeling is becoming obsolete as models advance
“I think, like, in a way, I think it's getting more and more to be a thing of the past as the models are getting smarter and smarter. This is becoming less of a thing, but I think a few years back, and especially in GPT-IV days, this was the thing.”
Jerry Tworek Oct 16, 2025 ▶ 47:16 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Insight
Tworek: Calling LLMs strictly next-token predictors is inaccurate in RL era
“Language models do on their own, like fundamental level is they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text.”
Jerry Tworek Oct 16, 2025 ▶ 3:00 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Opinion
Tworek: OpenAI's o1 was mostly a tech demo for solving puzzles
“O-one like, to be perfectly honest, it was really mostly good at solving puzzles and like maybe a few kind of thinking problems here and there, but it wasn't like, it wasn't a very useful model. It was almost more like a technology demonstration.”
Jerry Tworek Oct 16, 2025 ▶ 8:33 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Assertion Not checkable as stated
Tworek: Coding agents are the first successful agentic AI products
“Like, coding agents are at the moment the first, like, pretty successful agentic products built on top of AI.”
Jerry Tworek Oct 16, 2025 ▶ 10:36 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.