Process Reward Models
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · across every show →
Everything said about Process Reward Models, oldest first
Jan 11, 2024 positive
Lambert: RL feedback mechanisms will specialize across distinct task domains
“It seems very likely that different feedback will be used for different domains.
Chain of thought reasoning is great.
For math, and that's where these process reward models are being designed.
Probably not great for things like poetry, but as any tool gets bet…”