Dec 31, 2025 · 27m · latent-space

[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI

Josh McGrath · 15m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI post-training researcher Josh McGrath explores the technical evolution from GPT-4.1 to GPT-5.1, detailing breakthroughs in reasoning models, RLVR data engineering, token efficiency, and the enduring synergy between pre-training and post-training paradigms.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.7 Guest teaching 4.1 Guest disagreement 1.7 The hosts pushing back 2.9
05100:0010:0020:000:58–4:36 · The hosts as informed peer 5/10 Transitioning to Post-Training and Navigating RL Engineering Complexity The host asks informed technical questions about RL scaling challenges and managing third-party versus internal task code. The guest gently deflects specific questions about external vendor split, explaining how infrastructure debugging and Codex usage actually work during late-night runs.4:37–7:10 · The hosts as informed peer 5/10 Developing the Shopping Model and Implementing Real-Time Interruptibility The host probes the architectural decision of creating a standalone shopping model rather than a tool integration, and asks whether GPT-5 thinking replaces Deep Research. The guest collaboratively explains the rationale behind experimentation with real-time interruptibility and eventual capability convergence.7:11–13:12 · The hosts as informed peer 6/10 Balancing Model Personality Across the Anton Versus Clippy Spectrum The host introduces the Anton versus Clippy persona dichotomy and technical post-training milestones like GRPO and RLVR. The guest offers an insightful reframe, arguing that academic papers focus excessively on optimization algorithms rather than the underlying data reward signal quality.13:12–15:54 · The hosts as informed peer 6/10 Optimizing Token Efficiency and Managing Multi-Step Reasoning Horizons The host challenges the split between explicit top-level routers and implicit thinking time allocation in GPT-5. The guest explains why token efficiency and two-dimensional cost curves matter more in practice than raw wall-clock horizon time.15:54–21:24 · The hosts as informed peer 7/10 Context Window Scaling, Compaction, and Graph Walks Evaluation The host pushes back against relying solely on ever-expanding context windows by citing an 8-billion token enterprise retrieval case study and contrasting researcher versus systems engineer incentives. The guest defends context expansion by highlighting complex graph walk evaluations and the co-design culture at OpenAI.21:24–24:47 · The hosts as informed peer 5/10 Addressing Frontier Talent Bottlenecks and Automating Machine Learning Workflows The host presses the guest to identify whether systems engineering or machine learning research is more easily automated by LLMs. The guest carefully analyzes the bottleneck, noting the challenge of training hybrid engineers proficient in both distributed systems and statistical modeling.24:48–27:20 · The hosts as informed peer 6/10 Historical Parallels and the Future Synergy of Pre-Training and Post-Training The host highlights shifting compute allocations between pre-training and post-training citing the Grok 4 compute curve. The guest contextualizes the industry shift using an analogy to factory layout adaptation during the transition from steam power to electricity.0:58–4:36 · Guest teaching 4/10 Transitioning to Post-Training and Navigating RL Engineering Complexity The host asks informed technical questions about RL scaling challenges and managing third-party versus internal task code. The guest gently deflects specific questions about external vendor split, explaining how infrastructure debugging and Codex usage actually work during late-night runs.4:37–7:10 · Guest teaching 3/10 Developing the Shopping Model and Implementing Real-Time Interruptibility The host probes the architectural decision of creating a standalone shopping model rather than a tool integration, and asks whether GPT-5 thinking replaces Deep Research. The guest collaboratively explains the rationale behind experimentation with real-time interruptibility and eventual capability convergence.7:11–13:12 · Guest teaching 6/10 Balancing Model Personality Across the Anton Versus Clippy Spectrum The host introduces the Anton versus Clippy persona dichotomy and technical post-training milestones like GRPO and RLVR. The guest offers an insightful reframe, arguing that academic papers focus excessively on optimization algorithms rather than the underlying data reward signal quality.13:12–15:54 · Guest teaching 4/10 Optimizing Token Efficiency and Managing Multi-Step Reasoning Horizons The host challenges the split between explicit top-level routers and implicit thinking time allocation in GPT-5. The guest explains why token efficiency and two-dimensional cost curves matter more in practice than raw wall-clock horizon time.15:54–21:24 · Guest teaching 5/10 Context Window Scaling, Compaction, and Graph Walks Evaluation The host pushes back against relying solely on ever-expanding context windows by citing an 8-billion token enterprise retrieval case study and contrasting researcher versus systems engineer incentives. The guest defends context expansion by highlighting complex graph walk evaluations and the co-design culture at OpenAI.21:24–24:47 · Guest teaching 4/10 Addressing Frontier Talent Bottlenecks and Automating Machine Learning Workflows The host presses the guest to identify whether systems engineering or machine learning research is more easily automated by LLMs. The guest carefully analyzes the bottleneck, noting the challenge of training hybrid engineers proficient in both distributed systems and statistical modeling.24:48–27:20 · Guest teaching 3/10 Historical Parallels and the Future Synergy of Pre-Training and Post-Training The host highlights shifting compute allocations between pre-training and post-training citing the Grok 4 compute curve. The guest contextualizes the industry shift using an analogy to factory layout adaptation during the transition from steam power to electricity.0:58–4:36 · Guest disagreement 2/10 Transitioning to Post-Training and Navigating RL Engineering Complexity The host asks informed technical questions about RL scaling challenges and managing third-party versus internal task code. The guest gently deflects specific questions about external vendor split, explaining how infrastructure debugging and Codex usage actually work during late-night runs.4:37–7:10 · Guest disagreement 1/10 Developing the Shopping Model and Implementing Real-Time Interruptibility The host probes the architectural decision of creating a standalone shopping model rather than a tool integration, and asks whether GPT-5 thinking replaces Deep Research. The guest collaboratively explains the rationale behind experimentation with real-time interruptibility and eventual capability convergence.7:11–13:12 · Guest disagreement 2/10 Balancing Model Personality Across the Anton Versus Clippy Spectrum The host introduces the Anton versus Clippy persona dichotomy and technical post-training milestones like GRPO and RLVR. The guest offers an insightful reframe, arguing that academic papers focus excessively on optimization algorithms rather than the underlying data reward signal quality.13:12–15:54 · Guest disagreement 1/10 Optimizing Token Efficiency and Managing Multi-Step Reasoning Horizons The host challenges the split between explicit top-level routers and implicit thinking time allocation in GPT-5. The guest explains why token efficiency and two-dimensional cost curves matter more in practice than raw wall-clock horizon time.15:54–21:24 · Guest disagreement 3/10 Context Window Scaling, Compaction, and Graph Walks Evaluation The host pushes back against relying solely on ever-expanding context windows by citing an 8-billion token enterprise retrieval case study and contrasting researcher versus systems engineer incentives. The guest defends context expansion by highlighting complex graph walk evaluations and the co-design culture at OpenAI.21:24–24:47 · Guest disagreement 2/10 Addressing Frontier Talent Bottlenecks and Automating Machine Learning Workflows The host presses the guest to identify whether systems engineering or machine learning research is more easily automated by LLMs. The guest carefully analyzes the bottleneck, noting the challenge of training hybrid engineers proficient in both distributed systems and statistical modeling.24:48–27:20 · Guest disagreement 1/10 Historical Parallels and the Future Synergy of Pre-Training and Post-Training The host highlights shifting compute allocations between pre-training and post-training citing the Grok 4 compute curve. The guest contextualizes the industry shift using an analogy to factory layout adaptation during the transition from steam power to electricity.0:58–4:36 · The hosts pushing back 3/10 Transitioning to Post-Training and Navigating RL Engineering Complexity The host asks informed technical questions about RL scaling challenges and managing third-party versus internal task code. The guest gently deflects specific questions about external vendor split, explaining how infrastructure debugging and Codex usage actually work during late-night runs.4:37–7:10 · The hosts pushing back 2/10 Developing the Shopping Model and Implementing Real-Time Interruptibility The host probes the architectural decision of creating a standalone shopping model rather than a tool integration, and asks whether GPT-5 thinking replaces Deep Research. The guest collaboratively explains the rationale behind experimentation with real-time interruptibility and eventual capability convergence.7:11–13:12 · The hosts pushing back 2/10 Balancing Model Personality Across the Anton Versus Clippy Spectrum The host introduces the Anton versus Clippy persona dichotomy and technical post-training milestones like GRPO and RLVR. The guest offers an insightful reframe, arguing that academic papers focus excessively on optimization algorithms rather than the underlying data reward signal quality.13:12–15:54 · The hosts pushing back 3/10 Optimizing Token Efficiency and Managing Multi-Step Reasoning Horizons The host challenges the split between explicit top-level routers and implicit thinking time allocation in GPT-5. The guest explains why token efficiency and two-dimensional cost curves matter more in practice than raw wall-clock horizon time.15:54–21:24 · The hosts pushing back 5/10 Context Window Scaling, Compaction, and Graph Walks Evaluation The host pushes back against relying solely on ever-expanding context windows by citing an 8-billion token enterprise retrieval case study and contrasting researcher versus systems engineer incentives. The guest defends context expansion by highlighting complex graph walk evaluations and the co-design culture at OpenAI.21:24–24:47 · The hosts pushing back 3/10 Addressing Frontier Talent Bottlenecks and Automating Machine Learning Workflows The host presses the guest to identify whether systems engineering or machine learning research is more easily automated by LLMs. The guest carefully analyzes the bottleneck, noting the challenge of training hybrid engineers proficient in both distributed systems and statistical modeling.24:48–27:20 · The hosts pushing back 2/10 Historical Parallels and the Future Synergy of Pre-Training and Post-Training The host highlights shifting compute allocations between pre-training and post-training citing the Grok 4 compute curve. The guest contextualizes the industry shift using an analogy to factory layout adaptation during the transition from steam power to electricity.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 20:48 Pushing back against researcher versus systems dichotomy

The guest directly rejects the host's framing that researchers want to disregard systems, emphasizing instead that OpenAI's post-training workflow relies on deep co-design.

Hardest push from the hosts ▶ 14:46 Challenging model router abstractions

The host refuses to accept the current separation of explicit routing and implicit thinking budgets, asserting it introduces awkward failure modes that need merging.

Biggest teaching moment ▶ 9:01 Reframing post-training as data quality rather than optimization algorithms

The guest educates the host on how the community over-indexes on gradient optimization math while missing that the true differentiator is the verifiable fidelity of the training signal.

The host holds their own ▶ 19:05 Demonstrating context limitations with an 8-billion token real-world dataset

The host counters the idea of infinite context windows by bringing up real customer deployments requiring billions of tokens that cannot fit into memory without external retrieval systems.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Transitioning to Post-Training and Navigating RL Engineering Complexity 5423 The host asks informed technical questions about RL scaling challenges and managing third-party versus internal task code. The guest gently deflects specific questions about external vendor split, explaining how infrastructure debugging and Codex usage actually work during late-night runs.
Developing the Shopping Model and Implementing Real-Time Interruptibility 5312 The host probes the architectural decision of creating a standalone shopping model rather than a tool integration, and asks whether GPT-5 thinking replaces Deep Research. The guest collaboratively explains the rationale behind experimentation with real-time interruptibility and eventual capability convergence.
Balancing Model Personality Across the Anton Versus Clippy Spectrum 6622 The host introduces the Anton versus Clippy persona dichotomy and technical post-training milestones like GRPO and RLVR. The guest offers an insightful reframe, arguing that academic papers focus excessively on optimization algorithms rather than the underlying data reward signal quality.
Optimizing Token Efficiency and Managing Multi-Step Reasoning Horizons 6413 The host challenges the split between explicit top-level routers and implicit thinking time allocation in GPT-5. The guest explains why token efficiency and two-dimensional cost curves matter more in practice than raw wall-clock horizon time.
Context Window Scaling, Compaction, and Graph Walks Evaluation 7535 The host pushes back against relying solely on ever-expanding context windows by citing an 8-billion token enterprise retrieval case study and contrasting researcher versus systems engineer incentives. The guest defends context expansion by highlighting complex graph walk evaluations and the co-design culture at OpenAI.
Addressing Frontier Talent Bottlenecks and Automating Machine Learning Workflows 5423 The host presses the guest to identify whether systems engineering or machine learning research is more easily automated by LLMs. The guest carefully analyzes the bottleneck, noting the challenge of training hybrid engineers proficient in both distributed systems and statistical modeling.
Historical Parallels and the Future Synergy of Pre-Training and Post-Training 6312 The host highlights shifting compute allocations between pre-training and post-training citing the Grok 4 compute curve. The guest contextualizes the industry shift using an analogy to factory layout adaptation during the transition from steam power to electricity.

Statements from this episode (11)

Disclosure
McGrath: OpenAI Continues to Release Non-Thinking Models for Specific APIs
“No, we're still, we still are releasing non-thinking models but that one was the one that we did that was like API-specific non-thinking so, you know, focus has shifted a little.”
Josh McGrath Dec 31, 2025 ▶ 0:47
Insight
McGrath: RL runs have far more infrastructure failure points than pre-training
“The issue with RL is, like, you're doing tasks, and each task could have, like, a different grading setup, and each one of those different grading setups, that's, like, more infrastructure, and so, You know, when I'm staying up late trying to figure out what's…”
Josh McGrath Dec 31, 2025 ▶ 2:12
Insight
McGrath: Design specs let Codex complete hours of coding in 15 minutes
“If I spend, like, you know, 30, 40 minutes writing something that looks like a design doc or something, Codex can do more work than I can do in a few hours in, like, 15 minutes.”
Josh McGrath Dec 31, 2025 ▶ 3:50
Prediction Not checkable as stated
McGrath: Specialized and Frontier Reasoning AI Models Will Eventually Converge
“You know, I think if you look at like deep research, the original one and GPT-Five thinking on like high reasoning today, I think you'll see that like eventually the models all sort of converge in their capabilities.”
Josh McGrath Dec 31, 2025 ▶ 6:13
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Josh McGrath Dec 31, 2025 ▶ 6:46
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR, They're both policy gradient methods, but the, what's different is just like the input data.”
Josh McGrath Dec 31, 2025 ▶ 9:02
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Josh McGrath Dec 31, 2025 ▶ 12:42
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Josh McGrath Dec 31, 2025 ▶ 13:58
Prediction Not checkable as stated
McGrath: AGI will be a single tool that decides its own thinking time
“Yeah, I think, like, eventually, you know, we'll have AGI, and like, you're not gonna have to worry too much about how hard to think directly. It'll just, you know, we'll have a one tool that you always go to, and it knows how long to think for, and things lik…”
Josh McGrath Dec 31, 2025 ▶ 15:23
Assertion Partly supported
McGrath: OpenAI 10xed Effective Context Window for GPT-4.1
“I worked on long context, that was why I was on last, was for 4.1, where we, you know, I think, tenxed the effective context window for 4.1”
Josh McGrath Dec 31, 2025 ▶ 16:40
Insight
McGrath: AI Frontier Is Bottlenecked by Shortage of Hybrid Systems-ML Talent
“I think we're still having trouble not at OpenAI, but I think as a whole, producing lots of people that do lot, want to do lots of both systems work and ML work. And I think if you're trying to push the frontier, you don't know which Place is currently bottlen…”
Josh McGrath Dec 31, 2025 ▶ 21:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.