Group Relative Policy Optimization
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement May 23, 2025 by Will Brown · said 1 times in 1 episodes since 2026 · across every show →
Mentions by year
brought up most by Andrew White (1)
tap a year for its mentions
Everything said about Group Relative Policy Optimization, oldest first
May 23, 2025 positive
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”