McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
Josh McGrath · [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI · Dec 31, 2025 · at 12:42
OpenAI post-training researcher Josh McGrath evaluates the impact of DeepSeek Math and verifiable rewards compared to academic focus on optimization algorithms.
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know, you find the answer to a math problem, it's a lot less debatable than like, oh, well, is this thing that the human preferred actually what we want to do? Like you want to be right at math. And so I think in some ways that's underappreciated in I would say what's getting published.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →