RLHF

1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Oct 10, 2025 by Jonathan Siddharth · across every show →

Everything said about RLHF, oldest first

Oct 10, 2025 positive
Insight
Siddharth: Verifiable domains allow self-play reinforcement learning to replace RLHF
“Now, for these verifiable domains like coding and math, instead of doing reinforcement learning with human feedback, you can do reinforcement learning. Because you can automatically check when you got the correct answer or not in these verifiable domains. And …”
Jonathan Siddharth Oct 10, 2025 ▶ 24:24 Inside The $2.2B AI Research Accelerator | Turing · Sourcery with Molly O'Shea
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 160 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.