May 21, 2026 · 1h 14m · mad

OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real

Yann Dubois · 54m spoken Matt Turck · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Yann Dubois, Post-Training Frontiers Co-Lead at OpenAI, about the technical mechanics behind GPT-5.5, the evolution of reinforcement learning from synthetic math problems to messy real-world tasks, and the distinction between horizontal AI capabilities and last-mile vertical application development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 19.9% of the talking time here. How this is scored →

Matt as informed peer 3.5 Guest teaching 5.3 Guest disagreement 1.0 Matt pushing back 1.0
05100:0015:0030:0045:001:00:001:27–4:13 · Matt as informed peer 3/10 Unpacking the Step-Function Perception in AI Progress Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function.4:13–7:34 · Matt as informed peer 3/10 Model Reliability and Error Rates in Agentic Systems Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch.7:34–9:51 · Matt as informed peer 4/10 Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams.9:51–12:31 · Matt as informed peer 4/10 Optimizing Model Efficiency and Test-Time Scaling Curves Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward.12:31–14:43 · Matt as informed peer 2/10 Yann Dubois' Career Journey from Word2vec to OpenAI Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford.14:43–18:33 · Matt as informed peer 2/10 Behind the Scenes of the Live GPT-5 Announcement Demo Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live.18:33–21:28 · Matt as informed peer 4/10 Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users.21:28–28:21 · Matt as informed peer 4/10 Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths.28:21–31:05 · Matt as informed peer 3/10 Embodied AI, Physical Intuition, and World Models Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility.31:05–38:53 · Matt as informed peer 3/10 Defining Mid-Training: Overweighting High-Quality Curated Data Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code.38:53–42:01 · Matt as informed peer 3/10 Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts.42:01–45:10 · Matt as informed peer 5/10 Modern RL Algorithms and AI Development: Science vs. Alchemy Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks.45:10–48:21 · Matt as informed peer 4/10 Fast Post-Training Iterations and Horizontal Skill Classes Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics.48:21–50:26 · Matt as informed peer 3/10 Expanding AI Alignment Across Broader Economic Sectors Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection.50:26–56:04 · Matt as informed peer 4/10 Capability Generalization and Mitigating Hallucinations via RL Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices.56:04–1:00:19 · Matt as informed peer 4/10 Horizontal Capability Trade-Offs and Real-World Domain Tractability Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback.1:00:19–1:02:45 · Matt as informed peer 4/10 Challenges in AI Model Evaluation (Evals) Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce.1:02:45–1:05:48 · Matt as informed peer 4/10 Model as a Judge and the Capability Flywheel Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop.1:05:48–1:08:50 · Matt as informed peer 4/10 Continual Learning and the Enterprise Utility Curve Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly.1:08:50–1:11:23 · Matt as informed peer 4/10 The Role and Future of Agent Harnesses Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve.1:11:23–1:13:22 · Matt as informed peer 3/10 Building Last-Mile Vertical AI Applications Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence.1:27–4:13 · Guest teaching 5/10 Unpacking the Step-Function Perception in AI Progress Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function.4:13–7:34 · Guest teaching 5/10 Model Reliability and Error Rates in Agentic Systems Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch.7:34–9:51 · Guest teaching 5/10 Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams.9:51–12:31 · Guest teaching 5/10 Optimizing Model Efficiency and Test-Time Scaling Curves Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward.12:31–14:43 · Guest teaching 4/10 Yann Dubois' Career Journey from Word2vec to OpenAI Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford.14:43–18:33 · Guest teaching 3/10 Behind the Scenes of the Live GPT-5 Announcement Demo Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live.18:33–21:28 · Guest teaching 6/10 Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users.21:28–28:21 · Guest teaching 6/10 Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths.28:21–31:05 · Guest teaching 5/10 Embodied AI, Physical Intuition, and World Models Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility.31:05–38:53 · Guest teaching 6/10 Defining Mid-Training: Overweighting High-Quality Curated Data Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code.38:53–42:01 · Guest teaching 6/10 Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts.42:01–45:10 · Guest teaching 5/10 Modern RL Algorithms and AI Development: Science vs. Alchemy Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks.45:10–48:21 · Guest teaching 5/10 Fast Post-Training Iterations and Horizontal Skill Classes Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics.48:21–50:26 · Guest teaching 5/10 Expanding AI Alignment Across Broader Economic Sectors Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection.50:26–56:04 · Guest teaching 7/10 Capability Generalization and Mitigating Hallucinations via RL Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices.56:04–1:00:19 · Guest teaching 6/10 Horizontal Capability Trade-Offs and Real-World Domain Tractability Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback.1:00:19–1:02:45 · Guest teaching 6/10 Challenges in AI Model Evaluation (Evals) Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce.1:02:45–1:05:48 · Guest teaching 6/10 Model as a Judge and the Capability Flywheel Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop.1:05:48–1:08:50 · Guest teaching 6/10 Continual Learning and the Enterprise Utility Curve Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly.1:08:50–1:11:23 · Guest teaching 5/10 The Role and Future of Agent Harnesses Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve.1:11:23–1:13:22 · Guest teaching 4/10 Building Last-Mile Vertical AI Applications Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence.1:27–4:13 · Guest disagreement 1/10 Unpacking the Step-Function Perception in AI Progress Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function.4:13–7:34 · Guest disagreement 1/10 Model Reliability and Error Rates in Agentic Systems Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch.7:34–9:51 · Guest disagreement 1/10 Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams.9:51–12:31 · Guest disagreement 1/10 Optimizing Model Efficiency and Test-Time Scaling Curves Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward.12:31–14:43 · Guest disagreement 0/10 Yann Dubois' Career Journey from Word2vec to OpenAI Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford.14:43–18:33 · Guest disagreement 0/10 Behind the Scenes of the Live GPT-5 Announcement Demo Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live.18:33–21:28 · Guest disagreement 1/10 Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users.21:28–28:21 · Guest disagreement 1/10 Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths.28:21–31:05 · Guest disagreement 2/10 Embodied AI, Physical Intuition, and World Models Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility.31:05–38:53 · Guest disagreement 1/10 Defining Mid-Training: Overweighting High-Quality Curated Data Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code.38:53–42:01 · Guest disagreement 2/10 Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts.42:01–45:10 · Guest disagreement 1/10 Modern RL Algorithms and AI Development: Science vs. Alchemy Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks.45:10–48:21 · Guest disagreement 1/10 Fast Post-Training Iterations and Horizontal Skill Classes Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics.48:21–50:26 · Guest disagreement 1/10 Expanding AI Alignment Across Broader Economic Sectors Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection.50:26–56:04 · Guest disagreement 1/10 Capability Generalization and Mitigating Hallucinations via RL Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices.56:04–1:00:19 · Guest disagreement 2/10 Horizontal Capability Trade-Offs and Real-World Domain Tractability Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback.1:00:19–1:02:45 · Guest disagreement 1/10 Challenges in AI Model Evaluation (Evals) Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce.1:02:45–1:05:48 · Guest disagreement 1/10 Model as a Judge and the Capability Flywheel Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop.1:05:48–1:08:50 · Guest disagreement 1/10 Continual Learning and the Enterprise Utility Curve Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly.1:08:50–1:11:23 · Guest disagreement 1/10 The Role and Future of Agent Harnesses Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve.1:11:23–1:13:22 · Guest disagreement 0/10 Building Last-Mile Vertical AI Applications Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence.1:27–4:13 · Matt pushing back 1/10 Unpacking the Step-Function Perception in AI Progress Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function.4:13–7:34 · Matt pushing back 1/10 Model Reliability and Error Rates in Agentic Systems Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch.7:34–9:51 · Matt pushing back 1/10 Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams.9:51–12:31 · Matt pushing back 1/10 Optimizing Model Efficiency and Test-Time Scaling Curves Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward.12:31–14:43 · Matt pushing back 0/10 Yann Dubois' Career Journey from Word2vec to OpenAI Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford.14:43–18:33 · Matt pushing back 0/10 Behind the Scenes of the Live GPT-5 Announcement Demo Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live.18:33–21:28 · Matt pushing back 2/10 Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users.21:28–28:21 · Matt pushing back 2/10 Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths.28:21–31:05 · Matt pushing back 1/10 Embodied AI, Physical Intuition, and World Models Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility.31:05–38:53 · Matt pushing back 1/10 Defining Mid-Training: Overweighting High-Quality Curated Data Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code.38:53–42:01 · Matt pushing back 1/10 Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts.42:01–45:10 · Matt pushing back 1/10 Modern RL Algorithms and AI Development: Science vs. Alchemy Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks.45:10–48:21 · Matt pushing back 1/10 Fast Post-Training Iterations and Horizontal Skill Classes Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics.48:21–50:26 · Matt pushing back 1/10 Expanding AI Alignment Across Broader Economic Sectors Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection.50:26–56:04 · Matt pushing back 1/10 Capability Generalization and Mitigating Hallucinations via RL Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices.56:04–1:00:19 · Matt pushing back 1/10 Horizontal Capability Trade-Offs and Real-World Domain Tractability Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback.1:00:19–1:02:45 · Matt pushing back 1/10 Challenges in AI Model Evaluation (Evals) Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce.1:02:45–1:05:48 · Matt pushing back 1/10 Model as a Judge and the Capability Flywheel Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop.1:05:48–1:08:50 · Matt pushing back 2/10 Continual Learning and the Enterprise Utility Curve Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly.1:08:50–1:11:23 · Matt pushing back 1/10 The Role and Future of Agent Harnesses Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve.1:11:23–1:13:22 · Matt pushing back 0/10 Building Last-Mile Vertical AI Applications Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 45.2% · guest 54.8%0:00 · Matt 45.2% · guest 54.8%3:00 · Matt 20.4% · guest 79.6%3:00 · Matt 20.4% · guest 79.6%6:00 · Matt 15.8% · guest 84.2%6:00 · Matt 15.8% · guest 84.2%9:00 · Matt 13.3% · guest 86.7%9:00 · Matt 13.3% · guest 86.7%12:00 · Matt 17.6% · guest 82.4%12:00 · Matt 17.6% · guest 82.4%15:00 · Matt 34.4% · guest 65.6%15:00 · Matt 34.4% · guest 65.6%18:00 · Matt 22.2% · guest 77.8%18:00 · Matt 22.2% · guest 77.8%21:00 · Matt 32.7% · guest 67.3%21:00 · Matt 32.7% · guest 67.3%24:00 · Matt 10.9% · guest 89.1%24:00 · Matt 10.9% · guest 89.1%27:00 · Matt 19.3% · guest 80.7%27:00 · Matt 19.3% · guest 80.7%30:00 · Matt 11.6% · guest 88.4%30:00 · Matt 11.6% · guest 88.4%33:00 · Matt 1.1% · guest 98.9%33:00 · Matt 1.1% · guest 98.9%36:00 · Matt 8% · guest 92%36:00 · Matt 8% · guest 92%39:00 · Matt 11.5% · guest 88.5%39:00 · Matt 11.5% · guest 88.5%42:00 · Matt 18.6% · guest 81.4%42:00 · Matt 18.6% · guest 81.4%45:00 · Matt 28.2% · guest 71.8%45:00 · Matt 28.2% · guest 71.8%48:00 · Matt 38.7% · guest 61.3%48:00 · Matt 38.7% · guest 61.3%51:00 · Matt 0% · guest 100%51:00 · Matt 0% · guest 100%54:00 · Matt 17.3% · guest 82.7%54:00 · Matt 17.3% · guest 82.7%57:00 · Matt 11.7% · guest 88.3%57:00 · Matt 11.7% · guest 88.3%1:00:00 · Matt 13.4% · guest 86.6%1:00:00 · Matt 13.4% · guest 86.6%1:03:00 · Matt 22.4% · guest 77.6%1:03:00 · Matt 22.4% · guest 77.6%1:06:00 · Matt 13.2% · guest 86.8%1:06:00 · Matt 13.2% · guest 86.8%1:09:00 · Matt 40.5% · guest 59.5%1:09:00 · Matt 40.5% · guest 59.5%1:12:00 · Matt 35.7% · guest 64.3%1:12:00 · Matt 35.7% · guest 64.3%
Sharpest disagreement ▶ 39:10 Rejection of RL skepticism and 'cherry-on-top' framing

Yann pushes back against academic intuition (including Yann LeCun's quote) that RL is just an overcomplicated cherry on top, explaining how pre-trained models provided the necessary world priors for RL to scale.

Hardest push from Matt ▶ 20:10 Reconciling efficiency claims with extended thinking time

Matt challenges Yann to reconcile the claims of increased per-token model efficiency with the reality of long-thinking models like Pro that require extended wait times.

Biggest teaching moment ▶ 54:30 Explaining SFT vs RL dynamics in model hallucinations

Yann educates the host on fundamental machine learning mechanics, citing John Schulman's research to demonstrate how supervised fine-tuning forces models to hallucinate while reinforcement learning naturally suppresses it.

Matt holds his own ▶ 42:10 Citing state-of-the-art RL algorithms like GRPO

Matt demonstrates deep technical familiarity with the current post-training landscape by bringing up modern algorithms like GRPO and questioning their practical implementation over older methods.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Unpacking the Step-Function Perception in AI Progress 3511 Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function.
Model Reliability and Error Rates in Agentic Systems 3511 Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch.
Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams 4511 Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams.
Optimizing Model Efficiency and Test-Time Scaling Curves 4511 Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward.
Yann Dubois' Career Journey from Word2vec to OpenAI 2400 Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford.
Behind the Scenes of the Live GPT-5 Announcement Demo 2300 Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live.
Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro 4612 Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users.
Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor 4612 Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths.
Embodied AI, Physical Intuition, and World Models 3521 Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility.
Defining Mid-Training: Overweighting High-Quality Curated Data 3611 Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code.
Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard 3621 Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts.
Modern RL Algorithms and AI Development: Science vs. Alchemy 5511 Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks.
Fast Post-Training Iterations and Horizontal Skill Classes 4511 Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics.
Expanding AI Alignment Across Broader Economic Sectors 3511 Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection.
Capability Generalization and Mitigating Hallucinations via RL 4711 Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices.
Horizontal Capability Trade-Offs and Real-World Domain Tractability 4621 Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback.
Challenges in AI Model Evaluation (Evals) 4611 Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce.
Model as a Judge and the Capability Flywheel 4611 Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop.
Continual Learning and the Enterprise Utility Curve 4612 Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly.
The Role and Future of Agent Harnesses 4511 Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve.
Building Last-Mile Vertical AI Applications 3400 Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence.

Statements from this episode (42)

Insight
Dubois: AI progress feels discontinuous because OpenAI crossed a reliability threshold
“Even though the, in my mind, everything, the progress is actually pretty continuous, you need to reach this level of reliability. To really make any of these AI tools very useful, and I think we just crossed that probably December last year, at least at OpenAI…”
Yann Dubois May 21, 2026 ▶ 2:12
Insight
Dubois: Better AI models accelerate AI research by building tools and training models
“Once you start having models that are really good, you accelerate yourself. Especially in terms of coding, given that we all code internally yeah, you accelerate yourself both for having these models, like train the other models, but also like build like the t…”
Yann Dubois May 21, 2026 ▶ 2:47
Disclosure
Dubois: OpenAI expanded RL training from math competitions to real-world coding
“We were able to take many of the tools that we built for these, like, verifiable reward cases, and we were able to use them more generally in on, for reinforcement on, like, real use cases, and I think that's, like, really why we're feeling that right now in, …”
Yann Dubois May 21, 2026 ▶ 3:40
Insight
Dubois: Longer-running agentic AI tasks inherently carry higher overall error probabilities
“Given that these are agentic models the longer, if you just think about it as like every two minutes, there's like a certain probability that they are wrong. The longer that they run, The higher the probability that, like, the final answer is going to be wrong…”
Yann Dubois May 21, 2026 ▶ 4:28
Disclosure
OpenAI focuses on reducing AI agent error probability per time interval
“What we've been pushing a lot on is, like, making sure that the model, like, we decrease this probability of being wrong every, like, two minutes.”
Yann Dubois May 21, 2026 ▶ 4:44
Disclosure
Dubois: OpenAI internal sentiment goes through waves of hype and doubt
“It's kind of funny because in general with every model that is looking really good early on we have a model, we all get really excited about it. And then there's like tons of doubts. That start coming up because it's like, oh, like everyone is so high, is like…”
Yann Dubois May 21, 2026 ▶ 6:00
Insight
Dubois: OpenAI model iteration speed spans months upstream to days downstream
“So we really have like different sub teams including pre-training and you have like the mid training stage and like you have some post training and usually the closer you get to products like pushing being the last one, the faster the iteration cycle is. And i…”
Yann Dubois May 21, 2026 ▶ 7:10
Assertion Not checkable as stated
Dubois: GPT-5.5 performs most tasks roughly two times faster
“Most of the tasks can be basically performed, I would say like two X faster now with this model.”
Yann Dubois May 21, 2026 ▶ 9:22
Insight
Dubois: AI research aims to shift test-time scaling curves left
“The usual plot that you should be looking at is X axis, the number of tokens that you think for and Y axis the performance. So this is the, these test time scaling curves that we look at. And research basically tries to move this curve to the left. So think le…”
Yann Dubois May 21, 2026 ▶ 10:16
Assertion Not checkable as stated
Dubois: GPT-5.5 succeeded by combining inference efficiency and latency optimizations
“And the final thing that people care about is latency on x-axis, performance on y-axis, and this is where everything comes together, and this is really what happened with 5.5.”
Yann Dubois May 21, 2026 ▶ 10:49
Assertion Not checkable as stated
Dubois: Final rehearsal for OpenAI's GPT-5 live demo failed
“Right before we did that, like the last rehearsal, it did not work.”
Yann Dubois May 21, 2026 ▶ 15:26
Insight
Dubois: Test-time compute scaling exhibits logarithmic, diminishing returns
“We, we've seen again and again, the longer the model think for the better answers we will get. The problem is that this, these curves that we're talking about are not, are definitely not linear, and like they, there's some plateauing effect, and they kind of l…”
Yann Dubois May 21, 2026 ▶ 19:03
Disclosure
OpenAI's Yann Dubois rarely uses GPT-5.5 Pro due to high latency
“I personally don't use Pro that much because I really don't like waiting. I'm pretty impatient, so I don't like waiting for that long, and I know that the probability of being correct definitely improves, but it doesn't improve, like, enough for me to use it.”
Yann Dubois May 21, 2026 ▶ 19:33
Insight
Dubois: RL allows AI reasoning models to backtrack wrong paths earlier
“Part of it is the model knowing when it's going down the wrong path. But this is also something that we can that the model can be trained for with reinforcement learning is like knowing, okay, like that seems like not a great path. Let me backtrack and let me …”
Yann Dubois May 21, 2026 ▶ 22:59
Insight
Dubois: Larger AI models achieve higher efficiency by thinking through weights
“If you have larger models the amount of thinking time, so the amount of tokens they will think for will usually decrease. And the way that you can think about it is that metaphorically, the model already thinks through its weights when it generates a certain t…”
Yann Dubois May 21, 2026 ▶ 24:34
Assertion Not checkable as stated
Dubois: AI frontier labs have successfully bypassed internet data walls
“There were a lot of conversation about hitting data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. And it seems like different companies kind of found different ways to overcome the…”
Yann Dubois May 21, 2026 ▶ 26:43
Insight
Dubois: Multimodal data is not strictly necessary for strong AI reasoning
“I always thought That it would really help, ah, kind of your reasoning abilities if you have a lot of multimodal data And I still think this, but for example, like if you look at entropic models, they tend to not be that good on multimodal, and they are still …”
Yann Dubois May 21, 2026 ▶ 27:32
Prediction Not checkable as stated
Dubois: Embodied AI will advance general intelligence through world interaction
“I still believe that once we go to embodied agents, embodied AI, you will learn a lot about the world. And you will kind of improve general intelligence and usefulness to users by learning how, how the world interacts with itself.”
Yann Dubois May 21, 2026 ▶ 27:56
Prediction Not checkable as stated
Dubois: AI community is still pretty far from real-world embodied interaction
“So I do feel like we will improve the common sense of our model by having them interact in the real world. But we are, we're still pretty far from that, I think. And by we, I mean just generally the academic community and the AI community seems pretty far from…”
Yann Dubois May 21, 2026 ▶ 29:14
Prediction Not checkable as stated
Dubois: Simulations will never fully eliminate the need for real-world AI training
“The problem is simulations are always going to be really hard and are not going to be truthful. So I think there will always need to be a certain, a little bit of training that will need to happen in the real world to make sure that the model realizes kind of …”
Yann Dubois May 21, 2026 ▶ 29:54
Insight
Dubois: Supervised fine-tuning cannot exceed human demonstrator capabilities
“The problem with this is that you will never get better than what your ground truth gives you. And humans are actually pretty limited in many sense, so you will never, like, overcome the human labelers that you, you're working with.”
Yann Dubois May 21, 2026 ▶ 34:36
Insight
Dubois: Starting post-training with RL without SFT is extremely inefficient
“Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically.”
Yann Dubois May 21, 2026 ▶ 36:48
Assertion Partly supported
Dubois: Models like Kimi and DeepSeek use ~1M RL data points
“Now when you look at reinforcement learning from models like Kimi or from DeepSeq models, it seems that they are closer to one million data points.”
Yann Dubois May 21, 2026 ▶ 38:08
Insight
Dubois: RL becomes effective once base models possess strong world priors
“It seems that after crossing a certain scale of models that know basically everything about the world, and what we call, like, good priors about the world, It seems that reinforcement learning just started to work, and this is not only with LMS. Robotics seems…”
Yann Dubois May 21, 2026 ▶ 40:08
Insight
Dubois: Agentic RL training suffers from sparse reward credit assignment
“When we are training more agentic systems, you only know whether you're correct at the end of your very long rollout. So you get very little information per token of whether you were correct or not. And it's hard to say it's hard to basically do attribution. I…”
Yann Dubois May 21, 2026 ▶ 41:15
Prediction Not checkable as stated
Dubois: AI models good at math competitions will probably excel at coding
“If you are really good at, like, math competitions, your model will probably be pretty good at, like, coding competitions.”
Yann Dubois May 21, 2026 ▶ 47:45
Insight
Dubois: AI performance in new verticals depends on domain expert focus
“And in general, I would say the performance of the model really depends on, like, the number of people who care about the final output of the model and who are looking at that model. So if they start looking more on specific verticals, like, these verticals wi…”
Yann Dubois May 21, 2026 ▶ 50:03
Insight
Dubois: AI model knowledge calibration generalizes across all domains
“When you have hallucination of LMs, if a model is really bad at saying that it doesn't know, that usually happens in every single domain. You won't have, like, one domain where the model is extremely calibrated about its knowledge, and another domain where it'…”
Yann Dubois May 21, 2026 ▶ 54:04
Insight
Dubois: Effective reinforcement learning pipelines prevent AI hallucinations caused by SFT
“So, so hallucination at least the intuition that people have is that it can come, for example, from SFT, and it can come from this, like, pursuing pipeline, but if you have good reinforcement in pipeline, that shouldn't happen too often.”
Yann Dubois May 21, 2026 ▶ 55:42
Insight
Dubois: Horizontal AI capabilities can conflict with one another
“You will have cases where basically these horizontal capabilities go against each other.”
Yann Dubois May 21, 2026 ▶ 57:59
Prediction Not checkable as stated
Dubois: Model capacity does not limit AI performance in legal or medical fields
“But there's nothing, I would say, in the capacity of the model That is constraining the model to be as good at legal and like medical and like other domains.”
Yann Dubois May 21, 2026 ▶ 59:58
Assertion Not checkable as stated
Dubois: AI outperforming humans in specific domains creates evaluation bottlenecks
“Models in specific axes are becoming better than the majority of humans, and so we have fewer and fewer humans that can actually evaluate these models in particular axes.”
Yann Dubois May 21, 2026 ▶ 1:01:24
Insight
Dubois: Quantifying AI improvements is as important as training models
“Finding issues and, like, making sure that we can quantify improvements is just as important, if not more important, but there's always this, like, cultural gap.”
Yann Dubois May 21, 2026 ▶ 1:01:45
Insight
Dubois: Creating AI evaluations inherently creates methods for building training datasets
“Every time you build an eval, you actually build a way to build training data sets.”
Yann Dubois May 21, 2026 ▶ 1:03:13
Insight
Dubois: AI models create a capability flywheel as better models become better teachers
“As we get, like, better models we have this self-reinforcing loop, and we have this, like, capability flywheel, where better models become better teachers for other models.”
Yann Dubois May 21, 2026 ▶ 1:03:45
Disclosure
Dubois: A large portion of OpenAI post-training team works on model-as-a-judge frameworks
“A lot of my team works on that, and I think it's really critical is to work on this model, model as a judge kind of framework.”
Yann Dubois May 21, 2026 ▶ 1:04:12
Prediction Not checkable as stated
Dubois: AI's coding discontinuity will permeate other verticals within two years
“Now the feeling of discontinuity will happen. It did happen three months ago with coding or four months ago with coding, and I think that will happen now in every other domains. Like most people are not feeling the same way Like the, like kind of the capabilit…”
Yann Dubois May 21, 2026 ▶ 1:04:53
Insight
Dubois: AI models outperform new employees initially but lack continual learning
“Right now, actually most models at day zero, if you just drop them in a company arguably they are more useful than most new employees. So they start higher at T zero. But then across time they are mostly constant because they don't really learn kind of company…”
Yann Dubois May 21, 2026 ▶ 1:06:39
Prediction Not checkable as stated
Dubois: General AI agent harnesses designed to endure will not work
“If you try to have, like, a general harness to, that will, like, sustain over time I don't think that will work.”
Yann Dubois May 21, 2026 ▶ 1:10:25
What-if
Dubois: Current models with optimized harnesses would feel like AGI
“If we froze the models that we have right now, and you really worked on the harness, and, like, maybe, like, we also spend more time, like, training with, like, a great harness I think people would really feel the AGI in every single domain, or could already f…”
Yann Dubois May 21, 2026 ▶ 1:10:51
Insight
Dubois: Last-mile integration is the main AI bottleneck, not raw intelligence
“I think most of the time, the bottleneck is the last mile.”
Yann Dubois May 21, 2026 ▶ 1:12:34
Prediction Not checkable as stated
Dubois: Horizontal AI model progress will not stop anytime soon
“Maybe one day when we stop making horizontal progress, which I don't think is anytime soon, maybe we will start focusing on that, but yeah, that's not what we're doing now.”
Yann Dubois May 21, 2026 ▶ 1:12:59
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.