McGrath: RL runs have far more infrastructure failure points than pre-training
Josh McGrath · [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI · Dec 31, 2025 · at 2:12
Josh McGrath, a post-training researcher at OpenAI, explains the operational and debugging challenges specific to scaling reinforcement learning models.
“The issue with RL is, like, you're doing tasks, and each task could have, like, a different grading setup, and each one of those different grading setups, that's, like, more infrastructure, and so, You know, when I'm staying up late trying to figure out what's going on with a run, it could be in way more things than there's in a pre-training run, generally.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →