Oct 16, 2025 · 1h 16m · mad

How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek

Jerry Tworek · 58m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, OpenAI's VP of Research Jerry Tworek speaks with host Matt Turck about the technological leap from standard language models to advanced reasoning systems like o1, o3, and GPT-5. He shares insights into reinforcement learning, OpenAI's internal research culture, and the technical roadmap driving the pursuit of artificial general intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 18.8% of the talking time here. How this is scored →

Matt as informed peer 2.9 Guest teaching 3.2 Guest disagreement 0.5 Matt pushing back 0.9
05100:0020:0040:001:00:000:57–5:23 · Matt as informed peer 2/10 Defining Reasoning in Artificial Intelligence Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval.5:23–10:54 · Matt as informed peer 3/10 Balancing Thinking Time and User Experience Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product.10:54–13:27 · Matt as informed peer 1/10 Formative Years and Early Passion for Science in Poland Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw.13:27–17:20 · Matt as informed peer 1/10 Leaving Academia for Quantitative Trading in Europe Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam.17:20–20:01 · Matt as informed peer 2/10 Discovering Deep RL and Applying to OpenAI Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019.20:01–22:58 · Matt as informed peer 3/10 Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube.22:58–27:26 · Matt as informed peer 2/10 A Day in the Life of OpenAI's VP of Research Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy.27:26–29:37 · Matt as informed peer 2/10 Internal Transparency and Collaborative Culture Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality.29:37–32:50 · Matt as informed peer 2/10 Maintaining Shipping Velocity in a Generational Era Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user.32:50–35:12 · Matt as informed peer 3/10 Combining Pre-Training with Reinforcement Learning Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019.35:12–37:48 · Matt as informed peer 1/10 Reinforcement Learning 101: Dog Training Analogy Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping.37:48–40:09 · Matt as informed peer 4/10 Defining Policies, Agents, and Interactive Environments Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy.40:09–43:34 · Matt as informed peer 2/10 Evolution of Deep RL from Games to Pre-Trained Models Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied.43:34–47:52 · Matt as informed peer 3/10 The Breakthrough of RLHF and Human Preference Data Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters.47:52–49:58 · Matt as informed peer 4/10 Unsupervised Pre-Training and Representation Learning Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning.49:58–53:01 · Matt as informed peer 4/10 GRPO, DeepSeek, and Industry Reasoning Acceleration Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs.53:01–55:27 · Matt as informed peer 3/10 Scaling Reinforcement Learning vs. Pre-Training Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication).55:27–57:56 · Matt as informed peer 3/10 Agentic AI and Extended Reasoning Durations Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems.57:56–1:00:55 · Matt as informed peer 3/10 Online RL and Live User Interaction Risks Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training.1:00:55–1:05:50 · Matt as informed peer 4/10 Benchmarking Reasoning in Competitive Programming Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed.1:05:50–1:09:14 · Matt as informed peer 4/10 Evaluating RL Beyond Coding and Managing Reward Hacking Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives.1:09:14–1:12:19 · Matt as informed peer 4/10 The Journey Toward AGI and Self-Improving Systems Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold.1:12:19–1:15:43 · Matt as informed peer 6/10 Debating Pure RL vs. LLMs and Interview Conclusion Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training.0:57–5:23 · Guest teaching 3/10 Defining Reasoning in Artificial Intelligence Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval.5:23–10:54 · Guest teaching 3/10 Balancing Thinking Time and User Experience Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product.10:54–13:27 · Guest teaching 1/10 Formative Years and Early Passion for Science in Poland Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw.13:27–17:20 · Guest teaching 1/10 Leaving Academia for Quantitative Trading in Europe Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam.17:20–20:01 · Guest teaching 2/10 Discovering Deep RL and Applying to OpenAI Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019.20:01–22:58 · Guest teaching 2/10 Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube.22:58–27:26 · Guest teaching 3/10 A Day in the Life of OpenAI's VP of Research Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy.27:26–29:37 · Guest teaching 5/10 Internal Transparency and Collaborative Culture Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality.29:37–32:50 · Guest teaching 2/10 Maintaining Shipping Velocity in a Generational Era Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user.32:50–35:12 · Guest teaching 3/10 Combining Pre-Training with Reinforcement Learning Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019.35:12–37:48 · Guest teaching 4/10 Reinforcement Learning 101: Dog Training Analogy Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping.37:48–40:09 · Guest teaching 3/10 Defining Policies, Agents, and Interactive Environments Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy.40:09–43:34 · Guest teaching 4/10 Evolution of Deep RL from Games to Pre-Trained Models Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied.43:34–47:52 · Guest teaching 3/10 The Breakthrough of RLHF and Human Preference Data Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters.47:52–49:58 · Guest teaching 3/10 Unsupervised Pre-Training and Representation Learning Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning.49:58–53:01 · Guest teaching 4/10 GRPO, DeepSeek, and Industry Reasoning Acceleration Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs.53:01–55:27 · Guest teaching 4/10 Scaling Reinforcement Learning vs. Pre-Training Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication).55:27–57:56 · Guest teaching 3/10 Agentic AI and Extended Reasoning Durations Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems.57:56–1:00:55 · Guest teaching 4/10 Online RL and Live User Interaction Risks Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training.1:00:55–1:05:50 · Guest teaching 4/10 Benchmarking Reasoning in Competitive Programming Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed.1:05:50–1:09:14 · Guest teaching 5/10 Evaluating RL Beyond Coding and Managing Reward Hacking Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives.1:09:14–1:12:19 · Guest teaching 4/10 The Journey Toward AGI and Self-Improving Systems Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold.1:12:19–1:15:43 · Guest teaching 4/10 Debating Pure RL vs. LLMs and Interview Conclusion Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training.0:57–5:23 · Guest disagreement 0/10 Defining Reasoning in Artificial Intelligence Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval.5:23–10:54 · Guest disagreement 1/10 Balancing Thinking Time and User Experience Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product.10:54–13:27 · Guest disagreement 0/10 Formative Years and Early Passion for Science in Poland Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw.13:27–17:20 · Guest disagreement 0/10 Leaving Academia for Quantitative Trading in Europe Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam.17:20–20:01 · Guest disagreement 1/10 Discovering Deep RL and Applying to OpenAI Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019.20:01–22:58 · Guest disagreement 0/10 Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube.22:58–27:26 · Guest disagreement 1/10 A Day in the Life of OpenAI's VP of Research Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy.27:26–29:37 · Guest disagreement 2/10 Internal Transparency and Collaborative Culture Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality.29:37–32:50 · Guest disagreement 0/10 Maintaining Shipping Velocity in a Generational Era Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user.32:50–35:12 · Guest disagreement 0/10 Combining Pre-Training with Reinforcement Learning Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019.35:12–37:48 · Guest disagreement 0/10 Reinforcement Learning 101: Dog Training Analogy Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping.37:48–40:09 · Guest disagreement 0/10 Defining Policies, Agents, and Interactive Environments Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy.40:09–43:34 · Guest disagreement 0/10 Evolution of Deep RL from Games to Pre-Trained Models Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied.43:34–47:52 · Guest disagreement 0/10 The Breakthrough of RLHF and Human Preference Data Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters.47:52–49:58 · Guest disagreement 1/10 Unsupervised Pre-Training and Representation Learning Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning.49:58–53:01 · Guest disagreement 1/10 GRPO, DeepSeek, and Industry Reasoning Acceleration Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs.53:01–55:27 · Guest disagreement 0/10 Scaling Reinforcement Learning vs. Pre-Training Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication).55:27–57:56 · Guest disagreement 0/10 Agentic AI and Extended Reasoning Durations Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems.57:56–1:00:55 · Guest disagreement 1/10 Online RL and Live User Interaction Risks Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training.1:00:55–1:05:50 · Guest disagreement 0/10 Benchmarking Reasoning in Competitive Programming Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed.1:05:50–1:09:14 · Guest disagreement 1/10 Evaluating RL Beyond Coding and Managing Reward Hacking Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives.1:09:14–1:12:19 · Guest disagreement 1/10 The Journey Toward AGI and Self-Improving Systems Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold.1:12:19–1:15:43 · Guest disagreement 2/10 Debating Pure RL vs. LLMs and Interview Conclusion Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training.0:57–5:23 · Matt pushing back 1/10 Defining Reasoning in Artificial Intelligence Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval.5:23–10:54 · Matt pushing back 1/10 Balancing Thinking Time and User Experience Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product.10:54–13:27 · Matt pushing back 0/10 Formative Years and Early Passion for Science in Poland Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw.13:27–17:20 · Matt pushing back 0/10 Leaving Academia for Quantitative Trading in Europe Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam.17:20–20:01 · Matt pushing back 1/10 Discovering Deep RL and Applying to OpenAI Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019.20:01–22:58 · Matt pushing back 1/10 Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube.22:58–27:26 · Matt pushing back 1/10 A Day in the Life of OpenAI's VP of Research Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy.27:26–29:37 · Matt pushing back 1/10 Internal Transparency and Collaborative Culture Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality.29:37–32:50 · Matt pushing back 1/10 Maintaining Shipping Velocity in a Generational Era Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user.32:50–35:12 · Matt pushing back 1/10 Combining Pre-Training with Reinforcement Learning Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019.35:12–37:48 · Matt pushing back 0/10 Reinforcement Learning 101: Dog Training Analogy Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping.37:48–40:09 · Matt pushing back 1/10 Defining Policies, Agents, and Interactive Environments Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy.40:09–43:34 · Matt pushing back 0/10 Evolution of Deep RL from Games to Pre-Trained Models Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied.43:34–47:52 · Matt pushing back 1/10 The Breakthrough of RLHF and Human Preference Data Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters.47:52–49:58 · Matt pushing back 1/10 Unsupervised Pre-Training and Representation Learning Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning.49:58–53:01 · Matt pushing back 1/10 GRPO, DeepSeek, and Industry Reasoning Acceleration Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs.53:01–55:27 · Matt pushing back 1/10 Scaling Reinforcement Learning vs. Pre-Training Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication).55:27–57:56 · Matt pushing back 1/10 Agentic AI and Extended Reasoning Durations Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems.57:56–1:00:55 · Matt pushing back 1/10 Online RL and Live User Interaction Risks Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training.1:00:55–1:05:50 · Matt pushing back 1/10 Benchmarking Reasoning in Competitive Programming Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed.1:05:50–1:09:14 · Matt pushing back 1/10 Evaluating RL Beyond Coding and Managing Reward Hacking Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives.1:09:14–1:12:19 · Matt pushing back 1/10 The Journey Toward AGI and Self-Improving Systems Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold.1:12:19–1:15:43 · Matt pushing back 3/10 Debating Pure RL vs. LLMs and Interview Conclusion Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 38.9% · guest 61.1%0:00 · Matt 38.9% · guest 61.1%3:00 · Matt 12.1% · guest 87.9%3:00 · Matt 12.1% · guest 87.9%6:00 · Matt 18% · guest 82%6:00 · Matt 18% · guest 82%9:00 · Matt 21.4% · guest 78.6%9:00 · Matt 21.4% · guest 78.6%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 23.7% · guest 76.3%18:00 · Matt 23.7% · guest 76.3%21:00 · Matt 13.4% · guest 86.6%21:00 · Matt 13.4% · guest 86.6%24:00 · Matt 13.3% · guest 86.7%24:00 · Matt 13.3% · guest 86.7%27:00 · Matt 32.3% · guest 67.7%27:00 · Matt 32.3% · guest 67.7%30:00 · Matt 31.1% · guest 68.9%30:00 · Matt 31.1% · guest 68.9%33:00 · Matt 23.7% · guest 76.3%33:00 · Matt 23.7% · guest 76.3%36:00 · Matt 20.4% · guest 79.6%36:00 · Matt 20.4% · guest 79.6%39:00 · Matt 8.7% · guest 91.3%39:00 · Matt 8.7% · guest 91.3%42:00 · Matt 0.4% · guest 99.6%42:00 · Matt 0.4% · guest 99.6%45:00 · Matt 27.6% · guest 72.4%45:00 · Matt 27.6% · guest 72.4%48:00 · Matt 19.1% · guest 80.9%48:00 · Matt 19.1% · guest 80.9%51:00 · Matt 16.1% · guest 83.9%51:00 · Matt 16.1% · guest 83.9%54:00 · Matt 9.9% · guest 90.1%54:00 · Matt 9.9% · guest 90.1%57:00 · Matt 19% · guest 81%57:00 · Matt 19% · guest 81%1:00:00 · Matt 21.6% · guest 78.4%1:00:00 · Matt 21.6% · guest 78.4%1:03:00 · Matt 23.4% · guest 76.6%1:03:00 · Matt 23.4% · guest 76.6%1:06:00 · Matt 22% · guest 78%1:06:00 · Matt 22% · guest 78%1:09:00 · Matt 17.8% · guest 82.2%1:09:00 · Matt 17.8% · guest 82.2%1:12:00 · Matt 22.6% · guest 77.4%1:12:00 · Matt 22.6% · guest 77.4%1:15:00 · Matt 62.4% · guest 37.6%1:15:00 · Matt 62.4% · guest 37.6%
Sharpest disagreement ▶ 1:13:00 Rejection of Pure RL Thesis

Jerry directly rejects Richard Sutton's pure RL premise, asserting that RL cannot succeed without pre-training.

Hardest push from Matt ▶ 1:12:19 Sutton Pure RL Challenge

Matt challenges OpenAI's hybrid LLM approach by bringing up Richard Sutton's argument that LLMs are merely imitating reality.

Biggest teaching moment ▶ 27:26 Internal IP Transparency Correction

Jerry corrects Matt's intuitive assumption about internal IP compartmentalization, explaining that all researchers at OpenAI have full access to information.

Matt holds his own ▶ 1:12:19 Sutton Philosophical Counterpoint

Matt demonstrates deep familiarity with AI theory by referencing Richard Sutton's argument on Dwarkesh Patel's podcast regarding pure RL vs LLMs.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Reasoning in Artificial Intelligence 2301 Matt asks basic foundational questions about what reasoning and chain of thought actually mean in language models. Jerry explains that reasoning is an extended search process to reach unknown answers rather than simple retrieval.
Balancing Thinking Time and User Experience 3311 Matt synthesizes Jerry's explanation of thinking trade-offs into user experience design. Jerry agrees and candidly notes that O1 was mostly a technology demonstration rather than a fully useful product.
Formative Years and Early Passion for Science in Poland 1100 Matt asks standard biographical questions about Jerry's early life in Poland. Jerry shares his early dream of becoming a scientist and studying math in Warsaw.
Leaving Academia for Quantitative Trading in Europe 1100 Jerry describes leaving academia due to disillusionment with its structure, entering quantitative trading at JP Morgan, and co-founding hedge funds in London and Amsterdam.
Discovering Deep RL and Applying to OpenAI 2211 Jerry reframes the conventional view of AI milestones, stating that 2012 ImageNet wasn't as groundbreaking to him as DeepMind's 2013 DQN results. Matt tracks his move to OpenAI in 2019.
Early Work at OpenAI: Dota 2 and Robotic Rubik's Cube 3201 Matt asks about OpenAI's transition from early reinforcement learning projects like Dota 2 to unsupervised pre-training. Jerry details his early work on dexterous robotic manipulation solving a Rubik's cube.
A Day in the Life of OpenAI's VP of Research 2311 Jerry gently rejects Matt's top-down vs bottom-up question framing, explaining that research management at OpenAI requires balancing a tiny number of massive core bets with researcher autonomy.
Internal Transparency and Collaborative Culture 2521 Matt speculates that OpenAI must strictly compartmentalize IP internally. Jerry immediately corrects him, revealing that all ~600 researchers have access to everything because internal transparency is vital for research quality.
Maintaining Shipping Velocity in a Generational Era 2201 Matt asks how OpenAI maintains rapid release momentum without sacrificing long-term research. Jerry jokes about paying $200 per month for ChatGPT like any outside user.
Combining Pre-Training with Reinforcement Learning 3301 Matt accurately frames modern AI as combining pre-training and RL. Jerry recalls Ilya Sutskever stating that exact two-step thesis back in early 2019.
Reinforcement Learning 101: Dog Training Analogy 1400 Matt asks for a basic high-level explanation of RL. Jerry uses a dog training analogy involving a bag of treats and punishments to explain policy shaping.
Defining Policies, Agents, and Interactive Environments 4301 Matt lists core technical terms like policy, agent, action, and environment. Jerry clarifies the concept of interactive environments using a guitar-strumming feedback analogy.
Evolution of Deep RL from Games to Pre-Trained Models 2400 Jerry gives a historical overview of RL, revealing that initial internal reactions to GPT-4 pre-training were underwhelming before post-training techniques were applied.
The Breakthrough of RLHF and Human Preference Data 3301 Matt asks about RLHF mechanisms and data labeling companies like Scale AI. Jerry explains how the human data labeling industry must constantly shift upstream as models surpass human raters.
Unsupervised Pre-Training and Representation Learning 4311 Matt raises technical distinctions between self-supervised and unsupervised learning. Jerry pushes back on treating those definitions as stark boundaries, explaining representation learning.
GRPO, DeepSeek, and Industry Reasoning Acceleration 4411 Matt cites Jerry's recent tweet regarding DeepSeek's GRPO release. Jerry analyzes how DeepSeek's open-source publication accelerated reasoning research across Western AI labs.
Scaling Reinforcement Learning vs. Pre-Training 3401 Matt asks what scaling RL requires in practice. Jerry contrasts the simple math of pre-training (like a steel mill) with the delicate complexity of scaling RL (like semiconductor fabrication).
Agentic AI and Extended Reasoning Durations 3301 Matt asks how agentic autonomy intersects with reasoning RL. Jerry outlines internal experiments with models reasoning continuously for 30 minutes to two hours on hard problems.
Online RL and Live User Interaction Risks 3411 Matt asks if production agents learn in real-time from user interactions. Jerry clarifies the severe safety risks of unconstrained online RL, noting OpenAI avoids real-time user-loop training.
Benchmarking Reasoning in Competitive Programming 4401 Matt cites OpenAI's impressive results at the ICPC World Finals. Jerry explains that competitive coding benchmark success was an accidental byproduct of using programming puzzles as an internal RL testbed.
Evaluating RL Beyond Coding and Managing Reward Hacking 4511 Matt asks how RL generalizes to subjective, non-binary fields. Jerry provides a reframe, comparing AI reward hacking directly to human workplace gaming and policymaking incentives.
The Journey Toward AGI and Self-Improving Systems 4411 Matt quotes Jerry's tweet regarding AGI timelines. Jerry breaks down the psychological shift in AGI definitions and highlights self-improving systems as the key threshold.
Debating Pure RL vs. LLMs and Interview Conclusion 6423 Matt presents Richard Sutton's argument that LLMs are flawed imitation devices and only pure RL leads to AGI. Jerry directly rejects the premise, arguing pure RL is useless without pre-training.

Statements from this episode (36)

Insight
Tworek: AI reasoning spent on unknown answers yields better results with time
“I think that, that, that difference is here, like answering a question usually means you already know the answer and you just elicit the answer, you know, and the process of reasoning is Getting to the answer that you don't know, and usually the longer you spe…”
Jerry Tworek Oct 16, 2025 ▶ 2:13
Insight
Tworek: Calling LLMs strictly next-token predictors is inaccurate in RL era
“Language models do on their own, like fundamental level is they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text.”
Jerry Tworek Oct 16, 2025 ▶ 3:00
Insight
Tworek: Chain of thought is an LLM's reasoning verbalized in human words
“What chain of thought is, is their thinking process verbalized using human words and human concepts.”
Jerry Tworek Oct 16, 2025 ▶ 3:16
Disclosure
Tworek: OpenAI's high and low reasoning modes use the exact same model
“Where you can have like a high rezoning model and the low rezoning models. And this is like in the end, the same model. You just, we just tweak the parameter, which says we want you to think longer or shorter.”
Jerry Tworek Oct 16, 2025 ▶ 6:33
Opinion
Tworek: OpenAI's o1 was mostly a tech demo for solving puzzles
“O-one like, to be perfectly honest, it was really mostly good at solving puzzles and like maybe a few kind of thinking problems here and there, but it wasn't like, it wasn't a very useful model. It was almost more like a technology demonstration.”
Jerry Tworek Oct 16, 2025 ▶ 8:33
Opinion
Tworek: GPT-5 can effectively be considered an iteration like 'o3.1'
“Like GPT-Five in some way I can be considered as like, oh, 3.1. It's a little bit of like, you know, iteration of like the same thing and the same concept”
Jerry Tworek Oct 16, 2025 ▶ 9:52
Assertion Not checkable as stated
Tworek: Coding agents are the first successful agentic AI products
“Like, coding agents are at the moment the first, like, pretty successful agentic products built on top of AI.”
Jerry Tworek Oct 16, 2025 ▶ 10:36
Opinion
Tworek: The landmark 2012 ImageNet AI results were not that significant
“From my perspective, and again, this is just how my brain works, that the 20 12 ImageNet results, like, weren't that significant.”
Jerry Tworek Oct 16, 2025 ▶ 17:46
Assertion Not checkable as stated
Tworek: Dexterous manipulation remains an elusive challenge for AI policies
“The project I was working on was focused on dexterous manipulation, which was back then and still continues to be an elusive challenge for trained policies.”
Jerry Tworek Oct 16, 2025 ▶ 22:25
Disclosure
Tworek: OpenAI focuses research on only three or four large-scale projects
“We all work on a very few projects total. There are not that many projects. OpenAI is not trying to do everything. We are not trying to like have portfolio. We are trying to have like multiple different bets. Always the idea is we do a few core things really, …”
Jerry Tworek Oct 16, 2025 ▶ 24:44
Insight
Tworek: Top-down management does not work in AI research organizations
“Top down structuring of research doesn't work in research organizations. I really don't believe in it because like you are not kind of hiring some of the smartest people in the world and open air has incredibly, incredibly smart people. To kind of tell them wh…”
Jerry Tworek Oct 16, 2025 ▶ 26:16
Assertion Not checkable as stated
Tworek: OpenAI's research division consists of slightly under 600 people
“The truth is in research at OpenAI, which is like around, like slightly less than 600 people at the moment, everyone knows everything really, really it does.”
Jerry Tworek Oct 16, 2025 ▶ 27:26
Insight
Tworek: Uninformed researchers pose a greater risk than IP leaks
“It is like, yeah, it is some like risk of losing IP, but I think the risk of not doing the right thing and of people not being informed about research and not being able to do the best research is much higher in my personal opinion and how, how I approach thos…”
Jerry Tworek Oct 16, 2025 ▶ 27:54
Assertion Not checkable as stated
Tworek: OpenAI internal teams heavily use Codex for coding
“We definitely use codecs a lot for coding, and this is only getting, getting better.”
Jerry Tworek Oct 16, 2025 ▶ 32:00
Disclosure
OpenAI VP Jerry Tworek pays $200 per month for ChatGPT Pro
“I think I am pretty heavy user of ChatGPT right now, happily paying like 200 dollars a month for it”
Jerry Tworek Oct 16, 2025 ▶ 32:04
Assertion Not checkable as stated
Tworek: Ilya Sutskever set OpenAI's current RL research roadmap in 2019
“And what he said at the beginning of 2019 was to train large generative model on all data we can and then do reinforcement learning on it. That was the OpenAI research Plan at the beginning of 2019. And this is exactly what we are doing today.”
Jerry Tworek Oct 16, 2025 ▶ 34:31
Insight
Tworek: Effective reinforcement learning requires a 50/50 balance of rewards and punishments
“In a good way, the good way to do RL is if you balance those things. So if you kind of give cookies half of the time and punish the other half of the time, but this is almost like a mathematical kind of kind of aspect of it.”
Jerry Tworek Oct 16, 2025 ▶ 37:01
Insight
Tworek: Reinforcement learning is the only way agents learn environmental reaction
“And that's kind of, like, the only way how to, like, really teach agents to, like, learn to react to changes in the environment is through reinforcement learning.”
Jerry Tworek Oct 16, 2025 ▶ 39:45
Assertion Not checkable as stated
Tworek: Lack of pre-training was the primary bottleneck for 2019 RL
“Like whenever, even when I started like 20, 2019, the reinforcement learning was kind of fashionable at that moment, although not like very successful, but we were, the reinforcement learning was able to solve a lot of games, but the bottleneck was there that …”
Jerry Tworek Oct 16, 2025 ▶ 41:07
Disclosure
Tworek: OpenAI team was initially underwhelmed by pre-trained GPT-4
“When we trained GPT-IV, we were pretty underwhelmed internally, and then there was a lot of moments, oh, we trained this small, we spent a lot of money on it, and it's kind of like, you know, pretty dumb, at least, like, you know, we have GPT-IV, GPT-III alrea…”
Jerry Tworek Oct 16, 2025 ▶ 43:41
Prediction Not checkable as stated
Tworek: Traditional human data labeling is becoming obsolete as models advance
“I think, like, in a way, I think it's getting more and more to be a thing of the past as the models are getting smarter and smarter. This is becoming less of a thing, but I think a few years back, and especially in GPT-IV days, this was the thing.”
Jerry Tworek Oct 16, 2025 ▶ 47:16
Insight
Tworek: Pre-Training on Unlabeled Data Yields Far More Intelligence Than Supervised Mapping
“There are many more bits usually in the targets than in the labels and studying the structure of targets itself. It yields much more learning and much more intelligence than learning the mapping itself. So like spending a whole compute on just learning the dat…”
Jerry Tworek Oct 16, 2025 ▶ 49:27
Assertion Not checkable as stated
Tworek: OpenAI's o1 release caught US AI labs unprepared for RL
“As far as I know, like our O-one release mostly caught a lot of us labs by surprise. They didn't have like similarly advanced RL research program to my knowledge, basically no one.”
Jerry Tworek Oct 16, 2025 ▶ 51:16
Disclosure
Tworek: OpenAI's RL algorithm is not GRPO but shares similar components
“Like what we, what OpenAI is doing is not exactly GRPO. It is slightly different in many different ways, but like some parts are definitely similar.”
Jerry Tworek Oct 16, 2025 ▶ 51:48
Insight
Jerry Tworek: Pre-training AI models is mathematically simple compared to RL
“The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, but very conceptually, mathematically speaking, pre-training is dead simple.”
Jerry Tworek Oct 16, 2025 ▶ 53:31
Assertion Not checkable as stated
Tworek: OpenAI internal models can currently reason for hours on tasks
“The models I can think for, like, 30 minutes, hour, two hours these days on certain, certain types of Tasks and problems like even, even, even longer than that.”
Jerry Tworek Oct 16, 2025 ▶ 56:49
Disclosure
Tworek: OpenAI does not currently train models online via live interactions
“This is not what I am aware, at least, like, not, not, not what OpenAI is doing at the moment”
Jerry Tworek Oct 16, 2025 ▶ 58:32
Opinion
Tworek: Online RL shouldn't be used at ChatGPT's scale without robust safeguards
“So I, at least until, until we have a really good safeguards, I don't think we should try to do that in anything like as complex and large scale as ChatGPT.”
Jerry Tworek Oct 16, 2025 ▶ 59:10
Insight
Tworek: AI models need deep understanding of consequences for alignment
“I don't think like you can just tell them all like, A few show with a few good things to do, and it will do them all needs to deeply understand its action and consequences to really be able to choose the right thing.”
Jerry Tworek Oct 16, 2025 ▶ 1:00:09
Insight
Tworek: AI alignment is a never-ending pursuit as human goals evolve
“And it's I think it's a never ending pursuit because like, even, even for humans, it's not super easy to define what's, what do we consider a light? And I think as our civilization will evolve, it will, the notion of alignment and the goals of humanity will, K…”
Jerry Tworek Oct 16, 2025 ▶ 1:00:22
Disclosure
Tworek: OpenAI's competitive programming performance was a research byproduct
“I think we used like specifically programming puzzles for a while as a, Very nice research test bed of our ideas. Those are nice problems to experiment on them, and they weren't ever, like, considered part of the product, but it's, those are pretty, like, comp…”
Jerry Tworek Oct 16, 2025 ▶ 1:01:49
Assertion Supported
Tworek: OpenAI placed second in AtCoder behind a former employee
“We did also like IOI, International Olympic Informatics earlier this year, and Adcoder Heuristics competition as well, where we went second, we, behind a single human that is also a Polish person that used to be employed by OpenAI some time ago.”
Jerry Tworek Oct 16, 2025 ▶ 1:04:20
Insight
Tworek: Reward hacking in AI mirrors human behavior under flawed incentives
“In some way you can say it's a limitation of reinforcement learning, but when I was thinking about it, I realized a lot of that happens in human systems as well. There are a lot of like incentive system and reward systems and even, even happens in workplaces a…”
Jerry Tworek Oct 16, 2025 ▶ 1:08:23
Prediction Not checkable as stated
Tworek: Pre-training and RL are necessary for AGI, but not sufficient
“I generally think something that we are doing, like, pre-training today is necessary. I think something that, like, we are doing RL today is necessary, and there will surely be a few things more, and like, we have a lot of, Very ambitious research programs on …”
Jerry Tworek Oct 16, 2025 ▶ 1:09:58
What-if
Tworek: Someone from 10 years ago would view today's ChatGPT as AGI
“If you talk to someone from 10 years ago and show them chat GPD from today, they would probably call it AGI”
Jerry Tworek Oct 16, 2025 ▶ 1:11:05
Insight
Tworek: Reinforcement learning and pre-training require each other to succeed
“And like, I don't like in terms of a pure RL, I don't think like really pure RL makes sense. RL needs Pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well.”
Jerry Tworek Oct 16, 2025 ▶ 1:13:14
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.