Process Reward Models
topic on 2 shows · 2 statements across 2 episodes
2 statements about Process Reward Models, every show
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Lambert: RL feedback mechanisms will specialize across distinct task domains
“It seems very likely that different feedback will be used for different domains.
Chain of thought reasoning is great.
For math, and that's where these process reward models are being designed.
Probably not great for things like poetry, but as any tool gets bet…”