reward hacking

5 statements across 4 episodes · 2 bullish · 2 bearish · 4 people on the record · first statement Jul 29, 2025 by Brendan Fortuna · across every show →

Everything said about reward hacking, oldest first

Jul 29, 2025 negative
Insight
Fortuna: LLM Graders for Prose Generation Are Highly Vulnerable to Reward Hacking
“And whenever using like an LLM grader, the task is like a little bit more pros or a little longer form generation. You could be very vulnerable to this. The models are super clever. They're incentivized to win, but they'll cheat and they'll do weird things.”
Brendan Fortuna Jul 29, 2025 ▶ 9:16 ⚡️Using RFT to Build Clinical Superintelligence
Jul 29, 2025 positive
Insight
Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
Brendan Fortuna Jul 29, 2025 ▶ 10:28 ⚡️Using RFT to Build Clinical Superintelligence
Oct 16, 2025 positive
Insight
Corbitt: RL reward hacking is easily detected as models repeat the exploit
“Reward hacking is quite easy to detect once it starts happening, because once the model's found some hack, it just starts, like, doing it all the time.”
Kyle Corbitt Oct 16, 2025 ▶ 1:05:41 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Dec 7, 2025 bearish
Insight
Goyal: Average AI companies cannot hire expertise to prevent reward hacking
“You need to have like a pretty specific expertise to design the RL environment in a way that's not vulnerable to reward hacking. And I think that either you'll end up with some fixed number of very well engineered RL environments, or you need to somehow employ…”
Ankur Goyal Dec 7, 2025 ▶ 24:35 The Great Evals Debate — Ankur Goyal & Malte Ubl
Apr 2, 2026 neutral
Insight
Manning: Reward hacking is unsolved in symbolic and pixel-based models
“I mean, to the extent that there's a misspecified reward that it seems like it could be hacked In a more symbolic world or in a more pixel based world. I don't know if Sun's got any thoughts, but I don't think that's really being solved.”
Chris Manning Apr 2, 2026 ▶ 50:54 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.