GRPO

topic on 4 shows · 12 statements across 11 episodes

Latent Space the MAD Podcast All-In TBPN

12 statements about GRPO, every show

Lample: Long-horizon RL trajectories require new algorithms beyond GRPO
“GRPO, for instance, it doesn't really work with any bit of policy, which was okay initially, because you are solving math problems that can be solved in like a few thousand tokens, so the model can actually generate them pretty quickly, so when you do your upd…”
Guillaume Lample Mar 30, 2026 ▶ 45:43 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
MAD Insight
Original GRPO algorithm is flaky but stabilizes with practical engineering tricks
“Vanilla GRP or the original algorithm, it is pretty flaky. Like where it is, you have to babysit it. Over the course of the year, many people had these tips and tricks where some people were saying, remove the KL divergence term. Like if you just drop it for m…”
Sebastian Raschka Jan 29, 2026 ▶ 32:28 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
MAD Prediction Not checkable as stated
RLVR and GRPO will stabilize into canonical algorithms similar to AdamW
“It's similar to, you know, optimizers with Adam. So Adam is, I mean, right now there's Adam W. There was SGD and all the other RMS prop and how they were called. And they kind of all converge to Adam W by adding more and more tricks. And I think that's the sam…”
Sebastian Raschka Jan 29, 2026 ▶ 34:51 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Josh McGrath Dec 31, 2025 ▶ 12:42 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
MAD Assertion Not checkable as stated
Lambert: Most AI labs probably use evolved GRPO rather than PPO
“In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO.”
Nathan Lambert Nov 20, 2025 ▶ 1:16:39 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Kyle Corbitt Oct 16, 2025 ▶ 22:46 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
MAD Disclosure
Tworek: OpenAI's RL algorithm is not GRPO but shares similar components
“Like what we, what OpenAI is doing is not exactly GRPO. It is slightly different in many different ways, but like some parts are definitely similar.”
Jerry Tworek Oct 16, 2025 ▶ 51:48 How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
Lenz: RL training wastes compute on saturated or impossible examples
“Once you've trained a few hundred steps of let's say GOP, Most of your training is just wasted on example that are either too hard for you and you didn't get any success on them or too easy and everything was a success.”
Barak Lenz Oct 11, 2025 ▶ 37:49 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
LATENT SPACE Assertion Supported
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”
Will Brown May 23, 2025 ▶ 32:07 ⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
Jin: Policy gradient algorithms function as weighted supervised fine-tuning
“If you kind of, like, look at, if you kind of stare at, like, this part it sort of looks like just, like, weighted supervised fine-tuning, right? Like, you have this, like, log of, like, the probability of a token and, like, some, like, weight on it and reinfo…”
Roger Jin Apr 29, 2025 ▶ 5:59 What is an RL environment? w/ Nous Research's Roger Jin
TBPN Disclosure
Srinivas: Perplexity plans to invest more in RL post-training this year
“The nice thing is a lot of open source code bases exist on how to replicate GRPO or PPO and post training these models. And we've been doing that work already. So that's where we plan to invest more resources into for this year.”
Aravind Srinivas Apr 25, 2025 ▶ 29:08 Perplexity Founder Explains What Comes Next - Aravind Srinivas on TBPN April 23rd
ALL-IN Assertion Supported
Palihapitiya: DeepSeek created GRPO algorithm to slash AI memory requirements
“They invented a totally different algorithm. There was the orthodoxy. Right? This thing called PPO that everybody used, and they were like, no, we're going to use something else called, I think it's called GRPO or something. It uses a lot less computer memory,…”
Chamath Palihapitiya Feb 2, 2025 ▶ 9:55 AI Czar David Sacks Explains the DeepSeek Freak Out

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.