LLM Grader

topic on 1 show · 2 statements across 1 episodes

Latent Space

2 statements about LLM Grader, every show

Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
Brendan Fortuna Jul 29, 2025 ▶ 10:28 ⚡️Using RFT to Build Clinical Superintelligence
Fortuna: LLM Graders for Prose Generation Are Highly Vulnerable to Reward Hacking
“And whenever using like an LLM grader, the task is like a little bit more pros or a little longer form generation. You could be very vulnerable to this. The models are super clever. They're incentivized to win, but they'll cheat and they'll do weird things.”
Brendan Fortuna Jul 29, 2025 ▶ 9:16 ⚡️Using RFT to Build Clinical Superintelligence

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.