Oct 16, 2025 · 1h 16m · mad
How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, OpenAI's VP of Research Jerry Tworek speaks with host Matt Turck about the technological leap from standard language models to advanced reasoning systems like o1, o3, and GPT-5. He shares insights into reinforcement learning, OpenAI's internal research culture, and the technical roadmap driving the pursuit of artificial general intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 18.8% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Jerry directly rejects Richard Sutton's pure RL premise, asserting that RL cannot succeed without pre-training.
Hardest push from Matt ▶ 1:12:19 Sutton Pure RL ChallengeMatt challenges OpenAI's hybrid LLM approach by bringing up Richard Sutton's argument that LLMs are merely imitating reality.
Biggest teaching moment ▶ 27:26 Internal IP Transparency CorrectionJerry corrects Matt's intuitive assumption about internal IP compartmentalization, explaining that all researchers at OpenAI have full access to information.
Matt holds his own ▶ 1:12:19 Sutton Philosophical CounterpointMatt demonstrates deep familiarity with AI theory by referencing Richard Sutton's argument on Dwarkesh Patel's podcast regarding pure RL vs LLMs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Defining Reasoning in Artificial Intelligence | 2 | 3 | 0 | 1 | Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval. | |
| Balancing Thinking Time and User Experience | 3 | 3 | 1 | 1 | Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product. | |
| Formative Years and Early Passion for Science in Poland | 1 | 1 | 0 | 0 | Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw. | |
| Leaving Academia for Quantitative Trading in Europe | 1 | 1 | 0 | 0 | Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam. | |
| Discovering Deep RL and Applying to OpenAI | 2 | 2 | 1 | 1 | Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019. | |
| Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube | 3 | 2 | 0 | 1 | Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube. | |
| A Day in the Life of OpenAI's VP of Research | 2 | 3 | 1 | 1 | Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy. | |
| Internal Transparency and Collaborative Culture | 2 | 5 | 2 | 1 | Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality. | |
| Maintaining Shipping Velocity in a Generational Era | 2 | 2 | 0 | 1 | Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user. | |
| Combining Pre-Training with Reinforcement Learning | 3 | 3 | 0 | 1 | Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019. | |
| Reinforcement Learning 101: Dog Training Analogy | 1 | 4 | 0 | 0 | Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping. | |
| Defining Policies, Agents, and Interactive Environments | 4 | 3 | 0 | 1 | Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy. | |
| Evolution of Deep RL from Games to Pre-Trained Models | 2 | 4 | 0 | 0 | Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied. | |
| The Breakthrough of RLHF and Human Preference Data | 3 | 3 | 0 | 1 | Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters. | |
| Unsupervised Pre-Training and Representation Learning | 4 | 3 | 1 | 1 | Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning. | |
| GRPO, DeepSeek, and Industry Reasoning Acceleration | 4 | 4 | 1 | 1 | Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs. | |
| Scaling Reinforcement Learning vs. Pre-Training | 3 | 4 | 0 | 1 | Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication). | |
| Agentic AI and Extended Reasoning Durations | 3 | 3 | 0 | 1 | Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems. | |
| Online RL and Live User Interaction Risks | 3 | 4 | 1 | 1 | Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training. | |
| Benchmarking Reasoning in Competitive Programming | 4 | 4 | 0 | 1 | Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed. | |
| Evaluating RL Beyond Coding and Managing Reward Hacking | 4 | 5 | 1 | 1 | Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives. | |
| The Journey Toward AGI and Self-Improving Systems | 4 | 4 | 1 | 1 | Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold. | |
| Debating Pure RL vs. LLMs and Interview Conclusion | 6 | 4 | 2 | 3 | Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training. |