Nov 26, 2025 · 1h 5m · mad
What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews OpenAI Lead Research Scientist Łukasz Kaiser about the trajectory of artificial intelligence, the transition from pre-training compute scaling to reasoning models, the creation of the Transformer architecture, and the future of multimodal AI.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.1% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
When Matt suggests GPT-5.1 represents a massive capability jump over GPT-5, Łukasz directly rejects the premise, stating 'I think less than you think' and explaining that pre-training was mostly about cost reduction.
Hardest push from Matt ▶ 53:17 Challenging transformer sufficiency with alternative architecturesMatt challenges the assumption that scaling transformers and reasoning is sufficient, explicitly pressing Łukasz on whether alternative paradigms like Yann LeCun's JEPA are required for true generalization.
Biggest teaching moment ▶ 48:34 Visual dot puzzle demonstrating frontier model limitationsŁukasz educates Matt on the jagged frontier of reasoning models by presenting a simple primary-school visual puzzle that stumps both Gemini 3 and GPT-5.1 despite their Olympiad-level math capabilities.
Matt holds his own ▶ 56:08 Detailed technical knowledge of Codex Max and context compactionMatt demonstrates deep technical fluency by asking a highly detailed question referencing GPT-5.1 Codex Max's long-running agentic execution and multi-context window compaction mechanisms.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Debunking the AI Slowdown Narrative and GPU Scaling Laws | 3 | 4 | 1 | 1 | Matt asks an informed question citing recent model drops to challenge the AI slowdown narrative. Łukasz explains how pre-training scaling laws still hold log-linearly while the new reasoning paradigm offers superior compute-to-capability gains. | |
| Practical Capability Leap: Web Browsing and Practical Real-World LLMs | 1 | 5 | 1 | 0 | Łukasz details how reasoning models search the web dynamically (e.g., finding zoo hours) rather than hallucinating from outdated memory. Matt listens as Łukasz highlights the gap between public perception and actual current capabilities. | |
| Developer Adoption and the Shift in Coding Workflows | 3 | 4 | 1 | 1 | Matt asks Łukasz to detail obvious low-hanging fruit in frontier model development. Łukasz outlines non-obvious internal engineering bugs, data cleaning challenges, and multimodal improvements. | |
| Deep Dive into Reasoning Models and Chain of Thought | 2 | 6 | 0 | 0 | Matt prompts an educational breakdown of reasoning models. Łukasz explains how non-differentiable chain-of-thought tokens require reinforcement learning and verifiable rewards rather than simple next-token gradient descent. | |
| Reinforcement Learning Paradigms: From RLHF to Outcome Verifiers | 3 | 5 | 0 | 0 | Matt raises the distinction between pre-training and post-training RL. Łukasz explains the shift from easily hackable RLHF preference models to robust outcome verifiers in technical domains. | |
| Generalization of Reasoning and Future Visual Chain of Thought | 4 | 5 | 1 | 1 | Matt references OpenAI's GDPVal benchmark to ask about expanding RL across economic sectors. Łukasz emphasizes how messy raw internet pre-training data is and why visual reasoning remains undertrained. | |
| How Chain of Thought Operates and Summary Generation | 3 | 6 | 0 | 0 | Matt inquires whether the displayed chain of thought matches actual processing. Łukasz clarifies that user-facing thinking is a distinct summary of raw, messy model tokens and explains self-correction behavior. | |
| Łukasz Kaiser’s Academic Journey: From Mathematics to Google Brain | 2 | 3 | 0 | 0 | Matt asks about Łukasz's path into AI. Łukasz describes transitioning from theoretical computer science in Poland and Germany to Google, highlighting France's unique 10-year academic leave policy. | |
| Inside the Creation of the Transformer Architecture | 3 | 5 | 1 | 1 | Matt asks how the Transformer paper came together. Łukasz corrects the myth of all eight authors being in one room and recalls early pushback against building single models for multiple tasks. | |
| Moving from Google to OpenAI and Lab Cultures | 2 | 3 | 0 | 0 | Matt asks about Łukasz's move to OpenAI. Łukasz details Ilya Sutskever's repeated recruitments, COVID workplace dynamics, and scaling differences between Google Brain and early OpenAI. | |
| Organizational Structure and GPU Resource Allocation at OpenAI | 3 | 4 | 0 | 0 | Matt asks if researchers compete for GPU compute. Łukasz clarifies that resource distribution is primarily governed by technical requirements, with pre-training consuming the majority. | |
| The Resurgence of Pre-Training, Distillation, and Model Economics | 3 | 6 | 1 | 1 | Matt asks about the future of pre-training. Łukasz explains how consumer scale transformed model deployment economics, shifting focus to distillation and smaller, cheaper inference models. | |
| Interpretability and Model Inner Workings in Modern AI | 4 | 5 | 0 | 0 | Matt inquires whether modern hybrid models remain black boxes. Łukasz cites OpenAI's recent work on sparse weights while emphasizing fundamental limits in comprehending massive parallel systems. | |
| Evolution from GPT-4 to GPT-5.1 and Post-Training | 3 | 6 | 2 | 1 | Matt notes that GPT-5.1 feels like a massive leap over GPT-5. Łukasz reveals that underlying technical changes were smaller than perceived, with post-training refinement driving user experience. | |
| Post-Training Techniques for Tone and Persona Steering | 3 | 5 | 1 | 0 | Matt asks about tone customization in GPT-5.1. Łukasz explains post-training RL steering and discusses why model naming was decoupled from underlying technical pre-training runs. | |
| Controlling Thinking Time and the Jagged Nature of Reasoning | 3 | 7 | 1 | 0 | Matt asks how thinking time allocation works. Łukasz explains compute-scaling benefits while highlighting how reasoning models exhibit jagged capabilities that fail simple primary school puzzles. | |
| Visual Case Study: The Dot Counting Puzzle Benchmark | 3 | 7 | 1 | 1 | Łukasz presents a visual dot-counting puzzle where frontier models fail simple context reasoning. Matt asks what trips them up, and Łukasz points to undertrained multimodal reasoning. | |
| Alternative AI Architectures, ARC Prize, and AI Interns | 4 | 5 | 1 | 1 | Matt asks whether non-transformer architectures like Yann LeCun's JEPA are needed for true generalization. Łukasz evaluates ARC Prize progress and engineering constraints on experimental architectures. | |
| Agentic Coding with GPT-5.1 Codex Max and Context Compaction | 4 | 6 | 0 | 0 | Matt references the specs of GPT-5.1 Codex Max. Łukasz explains how quadratic attention memory limits long-running tasks and how context compaction enables multi-day agentic workflows. | |
| Future of Knowledge Work, Trust, and Economic Adaptation | 4 | 6 | 1 | 1 | Matt asks what value remains for product builders as general models expand. Łukasz uses the translation industry to argue that human trust and verification preserve demand for human oversight. | |
| Frontier Research Directions: Physical Robotics and Embodied AI | 3 | 4 | 0 | 0 | Matt asks about future research frontiers. Łukasz identifies general-data RL as his key focus and predicts embodied robotics will rapidly advance once multimodal reasoning matures. |