RLVR

topic on 3 shows · 10 statements across 6 episodes

Latent Space Invest Like the Best the MAD Podcast

10 statements about RLVR, every show

MAD Assertion Supported
Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning f…”
Mitch Trojanowski Aug 5, 2026 ▶ 20:03 How to Build Autonomous, Long-Horizon AI Agents | Basis
MAD Assertion Partly supported
50 RLVR steps boosted Qwen 3 MATH-500 score from 15% to 50%
“I took the Quinn three model as part of my book, the reasoning from scratch book. And I trained it just for 50 steps with RLVR, and it goes from 15%, so one five percent accuracy on math 500 to 50% on math 500.”
Sebastian Raschka Jan 29, 2026 ▶ 23:58 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
MAD Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Sebastian Raschka Jan 29, 2026 ▶ 24:33 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
MAD Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
INVEST LIKE THE BEST Assertion Not checkable as stated
Post-training scaling laws drove all AI benchmark progress since October 2024
“And so all the progress we've had immense progress since October, 24 through today was based entirely on these two new scaling laws.”
Gavin Baker Dec 9, 2025 ▶ 9:49 GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview · Invest Like The Best
MAD Insight
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nathan Lambert Nov 20, 2025 ▶ 1:13:35 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Nathan Lambert Jul 31, 2025 ▶ 5:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”
Nathan Lambert Jul 31, 2025 ▶ 57:25 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.