RLVR

5 statements across 3 episodes · 3 bullish · 0 bearish · 3 people on the record · first statement Nov 20, 2025 by Nathan Lambert · across every show →

Everything said about RLVR, oldest first

Nov 20, 2025 positive
Insight
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nathan Lambert Nov 20, 2025 ▶ 1:13:35 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Jan 29, 2026 bullish
Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Jan 29, 2026
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Sebastian Raschka Jan 29, 2026 ▶ 24:33 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Jan 29, 2026 positive
Assertion Partly supported
50 RLVR steps boosted Qwen 3 MATH-500 score from 15% to 50%
“I took the Quinn three model as part of my book, the reasoning from scratch book. And I trained it just for 50 steps with RLVR, and it goes from 15%, so one five percent accuracy on math 500 to 50% on math 500.”
Sebastian Raschka Jan 29, 2026 ▶ 23:58 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Aug 5, 2026 neutral
Assertion Supported
Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning f…”
Mitch Trojanowski Aug 5, 2026 ▶ 20:03 How to Build Autonomous, Long-Horizon AI Agents | Basis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.