RLVR

4 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Jul 31, 2025 by Nathan Lambert · across every show →

Everything said about RLVR, oldest first

Jul 31, 2025 positive
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 positive
Insight
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Nathan Lambert Jul 31, 2025 ▶ 5:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Insight
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”
Nathan Lambert Jul 31, 2025 ▶ 57:25 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Dec 31, 2025
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.