May 21, 2026 · 1h 14m · mad
OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Yann Dubois, Post-Training Frontiers Co-Lead at OpenAI, about the technical mechanics behind GPT-5.5, the evolution of reinforcement learning from synthetic math problems to messy real-world tasks, and the distinction between horizontal AI capabilities and last-mile vertical application development.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 19.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Yann pushes back against academic intuition (including Yann LeCun's quote) that RL is just an overcomplicated cherry on top, explaining how pre-trained models provided the necessary world priors for RL to scale.
Hardest push from Matt ▶ 20:10 Reconciling efficiency claims with extended thinking timeMatt challenges Yann to reconcile the claims of increased per-token model efficiency with the reality of long-thinking models like Pro that require extended wait times.
Biggest teaching moment ▶ 54:30 Explaining SFT vs RL dynamics in model hallucinationsYann educates the host on fundamental machine learning mechanics, citing John Schulman's research to demonstrate how supervised fine-tuning forces models to hallucinate while reinforcement learning naturally suppresses it.
Matt holds his own ▶ 42:10 Citing state-of-the-art RL algorithms like GRPOMatt demonstrates deep technical familiarity with the current post-training landscape by bringing up modern algorithms like GRPO and questioning their practical implementation over older methods.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Unpacking the Step-Function Perception in AI Progress | 3 | 5 | 1 | 1 | Matt opens by framing recent releases like GPT-5.5 as unlocking a step-function jump in progress. Yann gently reframes this, explaining that while capability growth is actually continuous, hitting reliability thresholds creates the perception of a step function. | |
| Model Reliability and Error Rates in Agentic Systems | 3 | 5 | 1 | 1 | Matt inquires whether model reliability stems from applied engineering or core model improvements. Yann breaks down error probabilities over time in agentic workflows and describes internal sentiment cycles before launch. | |
| Pillars of GPT-5.5: Horizontal vs. Vertical Research Teams | 4 | 5 | 1 | 1 | Matt asks how OpenAI structures teams to achieve specialized capabilities across diverse tasks. Yann details the interplay between vertical domain teams and horizontal capability teams. | |
| Optimizing Model Efficiency and Test-Time Scaling Curves | 4 | 5 | 1 | 1 | Matt presses on how efficiency per token is optimized across AI research and engineering. Yann explains test-time scaling curves and how research shifts performance curves leftward. | |
| Yann Dubois' Career Journey from Word2vec to OpenAI | 2 | 4 | 0 | 0 | Matt asks Yann about his personal background and journey to OpenAI. Yann recounts discovering word2vec during his undergrad, working on low-resource NLP in Singapore, and doing his PhD at Stanford. | |
| Behind the Scenes of the Live GPT-5 Announcement Demo | 2 | 3 | 0 | 0 | Matt brings up Yann's appearance in the live GPT-5 demo video. Yann humorously recalls the high stress when the app failed during the final rehearsal right before going live. | |
| Test-Time Compute Scaling: GPT-5.5 Thinking vs. GPT-5.5 Pro | 4 | 6 | 1 | 2 | Matt asks for the core difference between Thinking and Pro models. Yann explains that Pro pours logarithmically higher test-time compute for marginal gains, making it ideal for mathematicians rather than impatient users. | |
| Reasoning Efficiency: The Undergrad vs. Domain Expert Metaphor | 4 | 6 | 1 | 2 | Matt asks Yann to reconcile per-token efficiency with thinking longer and how reasoning gets smarter. Yann uses an undergrad versus domain expert analogy to explain how better priors prune useless paths. | |
| Embodied AI, Physical Intuition, and World Models | 3 | 5 | 2 | 1 | Matt asks about data frontiers like multimodal data and embodied AI. Yann notes that while video and physical world interaction build intuitive common sense, simulated world models often suffer from over-optimization past utility. | |
| Defining Mid-Training: Overweighting High-Quality Curated Data | 3 | 6 | 1 | 1 | Matt introduces mid-training, asking why it is distinct from pre-training and post-training. Yann contrasts raw web data ingesting with overweighting high-quality sources like Wikipedia or code. | |
| Why Scaling Reinforcement Learning Is Infrastructure-Intensive and Hard | 3 | 6 | 2 | 1 | Matt asks why scaling reinforcement learning took so long and why it is notoriously difficult. Yann counters old academic skepticism (referencing LeCun's cherry-on-top view) and outlines infra costs and credit assignment problems in long rollouts. | |
| Modern RL Algorithms and AI Development: Science vs. Alchemy | 5 | 5 | 1 | 1 | Matt cites specific modern RL techniques like GRPO and asks about the balance of science versus alchemy. Yann explains why simple, scalable sampling methods triumph over overly complex frameworks. | |
| Fast Post-Training Iterations and Horizontal Skill Classes | 4 | 5 | 1 | 1 | Matt asks whether domain spikes in models stem from specific dataset targeting or core architecture. Yann highlights fast post-training iteration loops and explains that performance gains map to horizontal skill classes rather than narrow topics. | |
| Expanding AI Alignment Across Broader Economic Sectors | 3 | 5 | 1 | 1 | Matt asks how progress expands from math and coding into wider economic benchmarks like GDPval. Yann explains that domain prioritization is limited by human expert availability and curated data collection. | |
| Capability Generalization and Mitigating Hallucinations via RL | 4 | 7 | 1 | 1 | Matt explores capability generalization across domains and hallucination mitigation. Yann details why SFT incentivizes guessing and hallucination while RL actively penalizes incorrect sampling choices. | |
| Horizontal Capability Trade-Offs and Real-World Domain Tractability | 4 | 6 | 2 | 1 | Matt asks if trade-offs exist where getting better at one domain harms another. Yann explains tension between explicit instruction following and implicit intuition, as well as domain tractability based on verifiable feedback. | |
| Challenges in AI Model Evaluation (Evals) | 4 | 6 | 1 | 1 | Matt shifts to model evaluation (evals) and why measuring frontier performance is notoriously hard. Yann points out that open-ended real-world tasks lack single ground truths and human experts capable of grading them are scarce. | |
| Model as a Judge and the Capability Flywheel | 4 | 6 | 1 | 1 | Matt asks about the flywheel effect of AI evaluating AI via model-as-a-judge frameworks. Yann explains that building high-quality evals inherently creates training datasets, creating a virtuous automated loop. | |
| Continual Learning and the Enterprise Utility Curve | 4 | 6 | 1 | 2 | Matt asks about continual learning and automated loops. Yann introduces a enterprise utility curve over time, expressing surprise that three years after ChatGPT, models still cannot continually adapt to enterprise contexts on the fly. | |
| The Role and Future of Agent Harnesses | 4 | 5 | 1 | 1 | Matt brings up the debate on whether foundational models will absorb external agent harnesses. Yann advises builders to use harnesses for immediate vertical needs while expecting to retune them as models evolve. | |
| Building Last-Mile Vertical AI Applications | 3 | 4 | 0 | 0 | Matt asks if startups should still build vertical application software in an era of improving base models. Yann strongly encourages building last-mile vertical applications, calling integrations and permissions the true bottleneck rather than base intelligence. |