RLVR
4 statements across 2 episodes · 2 bullish · 0 bearish · 2 people on the record · first statement Jul 31, 2025 by Nathan Lambert · across every show →
Everything said about RLVR, oldest first
Jul 31, 2025 positive
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Jul 31, 2025 positive
Jul 31, 2025 neutral
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”