RLVR
5 statements across 3 episodes · 3 bullish · 0 bearish · 3 people on the record · first statement Nov 20, 2025 by Nathan Lambert · across every show →
Everything said about RLVR, oldest first
Nov 20, 2025 positive
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Jan 29, 2026 bullish
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Jan 29, 2026
Jan 29, 2026 positive
Aug 5, 2026 neutral
Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning f…”