process rewards models
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 2, 2025 by Nathan Lambert · across every show →
Everything said about process rewards models, oldest first
Jan 2, 2025 positive
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”