Process Reward Models

topic on 2 shows · 2 statements across 2 episodes

Latent Space the MAD Podcast

2 statements about Process Reward Models, every show

MAD Prediction Not checkable as stated
Process Reward Models will eventually become standard in LLM post-training
“I think it is promising and we will see it working at some point. I think it's just like right now it's still Tricky to make it work, but I am quite sure we'll see it as part of the standard repertoire at some point.”
Sebastian Raschka Jan 29, 2026 ▶ 26:39 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
LATENT SPACE Prediction Not checkable as stated
Lambert: RL feedback mechanisms will specialize across distinct task domains
“It seems very likely that different feedback will be used for different domains. Chain of thought reasoning is great. For math, and that's where these process reward models are being designed. Probably not great for things like poetry, but as any tool gets bet…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.