Reinforcement Learning from Human Feedback
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jul 30, 2025 by Dario Amodei · across every show →
Everything said about Reinforcement Learning from Human Feedback, oldest first
Jul 30, 2025
Amodei: GPT-2 and GPT-3 were built to test RLHF at scale
“Actually, the original reason for building GPT-II and GPT-III, it was an outgrowth of the kind of AI alignment work that we were doing, right? Where myself and Paul Cristiano and some of the Anthropic co-founders had invented this technique called RL from huma…”