Dec 14, 2025 · 1h 25m · lennys-podcast

Inside OpenAI: 2026 is the year of agents, AI’s biggest bottleneck, and why compute isn’t the issue

Alexander Embirikos · 55m spoken Lenny Rachitsky · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI Codex product lead Alexander Embirikos joins Lenny Rachitsky to discuss how autonomous coding agents are transforming software engineering, why human verification—not compute—is AI's primary bottleneck, and how code execution powers universal agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 27.8% of the talking time here. How this is scored →

Lenny as informed peer 4.1 Guest teaching 5.4 Guest disagreement 0.7 Lenny pushing back 0.9
05100:0020:0040:001:00:001:20:005:14–11:34 · Lenny as informed peer 4/10 Operating Culture and Bottoms-Up Speed at OpenAI Lenny sets the stage by contrasting startup dynamics with OpenAI's unique scale and offers the 'ready fire aim' framing. Alexander politely refines the analogy, explaining that long-range vision is clear while immediate tactical product-market fit is discovered purely empirically.11:34–15:43 · Lenny as informed peer 4/10 Redefining Codex as a Software Engineering Teammate Lenny contrasts Codex with existing autocomplete tools like Cursor. Alexander educates Lenny on their broader vision of Codex as a full-lifecycle software teammate rather than an IDE-bound autocomplete widget.15:43–21:23 · Lenny as informed peer 5/10 Unlocking Growth: Moving from Cloud to Interactive Local Sandboxes Lenny references Andrej Karpathy's experiences debugging with Codex to explore what unlocked growth. Alexander explains the counterintuitive discovery that their fully cloud-hosted async architecture was too advanced for mainstream adoption compared to local sandbox interaction.21:23–25:14 · Lenny as informed peer 3/10 Codex Max Architecture and Context Compaction Lenny asks how Codex differentiates on core intelligence beyond models. Alexander delivers an in-depth technical explanation of the three-tier architecture powering Codex Max and how context compaction works across model, API, and harness layers.25:15–32:02 · Lenny as informed peer 5/10 Why Code Generation Is the Backbone of Universal Agents Lenny connects Codex's B2B agent nature with Nick Turley's consumer super assistant thesis. Alexander provides a thesis reframing: code execution is fundamentally the most reliable and composable primitive for any computer-using agent.32:02–39:12 · Lenny as informed peer 5/10 The Evolution of Development: Reviewing AI Code and New Workflows Lenny references Cursor CEO Michael Truell's spec-driven development concepts. Alexander pushes back gently on purely spec-driven workflows, noting developers dislike writing specs and suggesting chatter-driven and feedback-driven paradigms instead.39:12–43:15 · Lenny as informed peer 6/10 Mixed-Initiative AI and Contextual Desktop Workflows Lenny cites Block CTO Dan G's experience with the Goose agent watching engineer screens. Alexander agrees on contextual desktop intelligence but pinpoints code validation and review as the true critical bottleneck.43:17–53:35 · Lenny as informed peer 4/10 Sponsor Segment: Jira Product Discovery Following the Jira sponsor read, Alexander describes how PMs and designers prototype directly with Codex, highlighting the Sora Android app build in 28 days. Lenny expresses genuine disbelief at the tiny team size and compressed timeframe.53:36–58:09 · Lenny as informed peer 4/10 Evaluating Product Progress Beyond Synthetic Benchmarks Lenny inquires about internal KPIs and public benchmark reliance. Alexander clarifies that synthetic benchmarks are secondary to tracking D7 retention and reading gritty qualitative complaints on Reddit and Twitter.58:09–1:02:10 · Lenny as informed peer 5/10 Contextual Intelligence and the Strategy Behind Atlas Browser Lenny brings up his own public tweet criticizing Atlas browser's AI-only search UX. Alexander welcomes the friction and explains the strategic need for an AI-native browser rendering engine to enable unobtrusive contextual actions.1:02:10–1:05:38 · Lenny as informed peer 3/10 Practical Guidance for Deploying Codex on Hard Problems Lenny asks for tactical tips on getting started with Codex. Alexander advises against starting with trivial toy problems, urging users to immediately test Codex against their most difficult, gnarly bugs.1:05:38–1:10:35 · Lenny as informed peer 4/10 Essential Skills for Engineers and Self-Healing Infrastructure Lenny explores future essential engineering skills. Alexander highlights systems reasoning and self-healing infrastructure, sharing how Codex is currently trialed to be on call for babysitting its own expensive training runs.1:10:36–1:13:31 · Lenny as informed peer 4/10 AGI Timelines and Overcoming the Human Validation Bottleneck Lenny asks for Alexander's AGI timeline. Alexander offers an unconventional breakdown, asserting compute or model capability is not the current bottleneck, but rather human typing and verification throughput.1:13:32–1:15:53 · Lenny as informed peer 2/10 Recruiting for the Codex Team at OpenAI Alexander outlines recruiting needs for the Codex team, giving prospective candidates a specific mental rubric around visualizing engineer workflows over the next six months.1:15:54–1:24:46 · Lenny as informed peer 4/10 Lightning Round: Culture Sci-Fi, FSD UI, and Family Heritage In the lightning round, Alexander discusses Iain M. Banks' Culture series, Tesla's shared-control FSD UI as an agent design masterclass, and his family heritage in Andros, Greece.5:14–11:34 · Guest teaching 5/10 Operating Culture and Bottoms-Up Speed at OpenAI Lenny sets the stage by contrasting startup dynamics with OpenAI's unique scale and offers the 'ready fire aim' framing. Alexander politely refines the analogy, explaining that long-range vision is clear while immediate tactical product-market fit is discovered purely empirically.11:34–15:43 · Guest teaching 6/10 Redefining Codex as a Software Engineering Teammate Lenny contrasts Codex with existing autocomplete tools like Cursor. Alexander educates Lenny on their broader vision of Codex as a full-lifecycle software teammate rather than an IDE-bound autocomplete widget.15:43–21:23 · Guest teaching 6/10 Unlocking Growth: Moving from Cloud to Interactive Local Sandboxes Lenny references Andrej Karpathy's experiences debugging with Codex to explore what unlocked growth. Alexander explains the counterintuitive discovery that their fully cloud-hosted async architecture was too advanced for mainstream adoption compared to local sandbox interaction.21:23–25:14 · Guest teaching 7/10 Codex Max Architecture and Context Compaction Lenny asks how Codex differentiates on core intelligence beyond models. Alexander delivers an in-depth technical explanation of the three-tier architecture powering Codex Max and how context compaction works across model, API, and harness layers.25:15–32:02 · Guest teaching 6/10 Why Code Generation Is the Backbone of Universal Agents Lenny connects Codex's B2B agent nature with Nick Turley's consumer super assistant thesis. Alexander provides a thesis reframing: code execution is fundamentally the most reliable and composable primitive for any computer-using agent.32:02–39:12 · Guest teaching 5/10 The Evolution of Development: Reviewing AI Code and New Workflows Lenny references Cursor CEO Michael Truell's spec-driven development concepts. Alexander pushes back gently on purely spec-driven workflows, noting developers dislike writing specs and suggesting chatter-driven and feedback-driven paradigms instead.39:12–43:15 · Guest teaching 4/10 Mixed-Initiative AI and Contextual Desktop Workflows Lenny cites Block CTO Dan G's experience with the Goose agent watching engineer screens. Alexander agrees on contextual desktop intelligence but pinpoints code validation and review as the true critical bottleneck.43:17–53:35 · Guest teaching 5/10 Sponsor Segment: Jira Product Discovery Following the Jira sponsor read, Alexander describes how PMs and designers prototype directly with Codex, highlighting the Sora Android app build in 28 days. Lenny expresses genuine disbelief at the tiny team size and compressed timeframe.53:36–58:09 · Guest teaching 5/10 Evaluating Product Progress Beyond Synthetic Benchmarks Lenny inquires about internal KPIs and public benchmark reliance. Alexander clarifies that synthetic benchmarks are secondary to tracking D7 retention and reading gritty qualitative complaints on Reddit and Twitter.58:09–1:02:10 · Guest teaching 6/10 Contextual Intelligence and the Strategy Behind Atlas Browser Lenny brings up his own public tweet criticizing Atlas browser's AI-only search UX. Alexander welcomes the friction and explains the strategic need for an AI-native browser rendering engine to enable unobtrusive contextual actions.1:02:10–1:05:38 · Guest teaching 6/10 Practical Guidance for Deploying Codex on Hard Problems Lenny asks for tactical tips on getting started with Codex. Alexander advises against starting with trivial toy problems, urging users to immediately test Codex against their most difficult, gnarly bugs.1:05:38–1:10:35 · Guest teaching 6/10 Essential Skills for Engineers and Self-Healing Infrastructure Lenny explores future essential engineering skills. Alexander highlights systems reasoning and self-healing infrastructure, sharing how Codex is currently trialed to be on call for babysitting its own expensive training runs.1:10:36–1:13:31 · Guest teaching 7/10 AGI Timelines and Overcoming the Human Validation Bottleneck Lenny asks for Alexander's AGI timeline. Alexander offers an unconventional breakdown, asserting compute or model capability is not the current bottleneck, but rather human typing and verification throughput.1:13:32–1:15:53 · Guest teaching 4/10 Recruiting for the Codex Team at OpenAI Alexander outlines recruiting needs for the Codex team, giving prospective candidates a specific mental rubric around visualizing engineer workflows over the next six months.1:15:54–1:24:46 · Guest teaching 3/10 Lightning Round: Culture Sci-Fi, FSD UI, and Family Heritage In the lightning round, Alexander discusses Iain M. Banks' Culture series, Tesla's shared-control FSD UI as an agent design masterclass, and his family heritage in Andros, Greece.5:14–11:34 · Guest disagreement 1/10 Operating Culture and Bottoms-Up Speed at OpenAI Lenny sets the stage by contrasting startup dynamics with OpenAI's unique scale and offers the 'ready fire aim' framing. Alexander politely refines the analogy, explaining that long-range vision is clear while immediate tactical product-market fit is discovered purely empirically.11:34–15:43 · Guest disagreement 1/10 Redefining Codex as a Software Engineering Teammate Lenny contrasts Codex with existing autocomplete tools like Cursor. Alexander educates Lenny on their broader vision of Codex as a full-lifecycle software teammate rather than an IDE-bound autocomplete widget.15:43–21:23 · Guest disagreement 1/10 Unlocking Growth: Moving from Cloud to Interactive Local Sandboxes Lenny references Andrej Karpathy's experiences debugging with Codex to explore what unlocked growth. Alexander explains the counterintuitive discovery that their fully cloud-hosted async architecture was too advanced for mainstream adoption compared to local sandbox interaction.21:23–25:14 · Guest disagreement 0/10 Codex Max Architecture and Context Compaction Lenny asks how Codex differentiates on core intelligence beyond models. Alexander delivers an in-depth technical explanation of the three-tier architecture powering Codex Max and how context compaction works across model, API, and harness layers.25:15–32:02 · Guest disagreement 2/10 Why Code Generation Is the Backbone of Universal Agents Lenny connects Codex's B2B agent nature with Nick Turley's consumer super assistant thesis. Alexander provides a thesis reframing: code execution is fundamentally the most reliable and composable primitive for any computer-using agent.32:02–39:12 · Guest disagreement 1/10 The Evolution of Development: Reviewing AI Code and New Workflows Lenny references Cursor CEO Michael Truell's spec-driven development concepts. Alexander pushes back gently on purely spec-driven workflows, noting developers dislike writing specs and suggesting chatter-driven and feedback-driven paradigms instead.39:12–43:15 · Guest disagreement 0/10 Mixed-Initiative AI and Contextual Desktop Workflows Lenny cites Block CTO Dan G's experience with the Goose agent watching engineer screens. Alexander agrees on contextual desktop intelligence but pinpoints code validation and review as the true critical bottleneck.43:17–53:35 · Guest disagreement 0/10 Sponsor Segment: Jira Product Discovery Following the Jira sponsor read, Alexander describes how PMs and designers prototype directly with Codex, highlighting the Sora Android app build in 28 days. Lenny expresses genuine disbelief at the tiny team size and compressed timeframe.53:36–58:09 · Guest disagreement 1/10 Evaluating Product Progress Beyond Synthetic Benchmarks Lenny inquires about internal KPIs and public benchmark reliance. Alexander clarifies that synthetic benchmarks are secondary to tracking D7 retention and reading gritty qualitative complaints on Reddit and Twitter.58:09–1:02:10 · Guest disagreement 1/10 Contextual Intelligence and the Strategy Behind Atlas Browser Lenny brings up his own public tweet criticizing Atlas browser's AI-only search UX. Alexander welcomes the friction and explains the strategic need for an AI-native browser rendering engine to enable unobtrusive contextual actions.1:02:10–1:05:38 · Guest disagreement 1/10 Practical Guidance for Deploying Codex on Hard Problems Lenny asks for tactical tips on getting started with Codex. Alexander advises against starting with trivial toy problems, urging users to immediately test Codex against their most difficult, gnarly bugs.1:05:38–1:10:35 · Guest disagreement 0/10 Essential Skills for Engineers and Self-Healing Infrastructure Lenny explores future essential engineering skills. Alexander highlights systems reasoning and self-healing infrastructure, sharing how Codex is currently trialed to be on call for babysitting its own expensive training runs.1:10:36–1:13:31 · Guest disagreement 1/10 AGI Timelines and Overcoming the Human Validation Bottleneck Lenny asks for Alexander's AGI timeline. Alexander offers an unconventional breakdown, asserting compute or model capability is not the current bottleneck, but rather human typing and verification throughput.1:13:32–1:15:53 · Guest disagreement 0/10 Recruiting for the Codex Team at OpenAI Alexander outlines recruiting needs for the Codex team, giving prospective candidates a specific mental rubric around visualizing engineer workflows over the next six months.1:15:54–1:24:46 · Guest disagreement 0/10 Lightning Round: Culture Sci-Fi, FSD UI, and Family Heritage In the lightning round, Alexander discusses Iain M. Banks' Culture series, Tesla's shared-control FSD UI as an agent design masterclass, and his family heritage in Andros, Greece.5:14–11:34 · Lenny pushing back 2/10 Operating Culture and Bottoms-Up Speed at OpenAI Lenny sets the stage by contrasting startup dynamics with OpenAI's unique scale and offers the 'ready fire aim' framing. Alexander politely refines the analogy, explaining that long-range vision is clear while immediate tactical product-market fit is discovered purely empirically.11:34–15:43 · Lenny pushing back 1/10 Redefining Codex as a Software Engineering Teammate Lenny contrasts Codex with existing autocomplete tools like Cursor. Alexander educates Lenny on their broader vision of Codex as a full-lifecycle software teammate rather than an IDE-bound autocomplete widget.15:43–21:23 · Lenny pushing back 1/10 Unlocking Growth: Moving from Cloud to Interactive Local Sandboxes Lenny references Andrej Karpathy's experiences debugging with Codex to explore what unlocked growth. Alexander explains the counterintuitive discovery that their fully cloud-hosted async architecture was too advanced for mainstream adoption compared to local sandbox interaction.21:23–25:14 · Lenny pushing back 0/10 Codex Max Architecture and Context Compaction Lenny asks how Codex differentiates on core intelligence beyond models. Alexander delivers an in-depth technical explanation of the three-tier architecture powering Codex Max and how context compaction works across model, API, and harness layers.25:15–32:02 · Lenny pushing back 1/10 Why Code Generation Is the Backbone of Universal Agents Lenny connects Codex's B2B agent nature with Nick Turley's consumer super assistant thesis. Alexander provides a thesis reframing: code execution is fundamentally the most reliable and composable primitive for any computer-using agent.32:02–39:12 · Lenny pushing back 2/10 The Evolution of Development: Reviewing AI Code and New Workflows Lenny references Cursor CEO Michael Truell's spec-driven development concepts. Alexander pushes back gently on purely spec-driven workflows, noting developers dislike writing specs and suggesting chatter-driven and feedback-driven paradigms instead.39:12–43:15 · Lenny pushing back 1/10 Mixed-Initiative AI and Contextual Desktop Workflows Lenny cites Block CTO Dan G's experience with the Goose agent watching engineer screens. Alexander agrees on contextual desktop intelligence but pinpoints code validation and review as the true critical bottleneck.43:17–53:35 · Lenny pushing back 1/10 Sponsor Segment: Jira Product Discovery Following the Jira sponsor read, Alexander describes how PMs and designers prototype directly with Codex, highlighting the Sora Android app build in 28 days. Lenny expresses genuine disbelief at the tiny team size and compressed timeframe.53:36–58:09 · Lenny pushing back 1/10 Evaluating Product Progress Beyond Synthetic Benchmarks Lenny inquires about internal KPIs and public benchmark reliance. Alexander clarifies that synthetic benchmarks are secondary to tracking D7 retention and reading gritty qualitative complaints on Reddit and Twitter.58:09–1:02:10 · Lenny pushing back 2/10 Contextual Intelligence and the Strategy Behind Atlas Browser Lenny brings up his own public tweet criticizing Atlas browser's AI-only search UX. Alexander welcomes the friction and explains the strategic need for an AI-native browser rendering engine to enable unobtrusive contextual actions.1:02:10–1:05:38 · Lenny pushing back 0/10 Practical Guidance for Deploying Codex on Hard Problems Lenny asks for tactical tips on getting started with Codex. Alexander advises against starting with trivial toy problems, urging users to immediately test Codex against their most difficult, gnarly bugs.1:05:38–1:10:35 · Lenny pushing back 1/10 Essential Skills for Engineers and Self-Healing Infrastructure Lenny explores future essential engineering skills. Alexander highlights systems reasoning and self-healing infrastructure, sharing how Codex is currently trialed to be on call for babysitting its own expensive training runs.1:10:36–1:13:31 · Lenny pushing back 1/10 AGI Timelines and Overcoming the Human Validation Bottleneck Lenny asks for Alexander's AGI timeline. Alexander offers an unconventional breakdown, asserting compute or model capability is not the current bottleneck, but rather human typing and verification throughput.1:13:32–1:15:53 · Lenny pushing back 0/10 Recruiting for the Codex Team at OpenAI Alexander outlines recruiting needs for the Codex team, giving prospective candidates a specific mental rubric around visualizing engineer workflows over the next six months.1:15:54–1:24:46 · Lenny pushing back 0/10 Lightning Round: Culture Sci-Fi, FSD UI, and Family Heritage In the lightning round, Alexander discusses Iain M. Banks' Culture series, Tesla's shared-control FSD UI as an agent design masterclass, and his family heritage in Andros, Greece.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 59.7% · guest 40.3%0:00 · Lenny 59.7% · guest 40.3%3:00 · Lenny 92.8% · guest 7.2%3:00 · Lenny 92.8% · guest 7.2%6:00 · Lenny 12.4% · guest 87.6%6:00 · Lenny 12.4% · guest 87.6%9:00 · Lenny 34.1% · guest 65.9%9:00 · Lenny 34.1% · guest 65.9%12:00 · Lenny 9.5% · guest 90.5%12:00 · Lenny 9.5% · guest 90.5%15:00 · Lenny 32.3% · guest 67.7%15:00 · Lenny 32.3% · guest 67.7%18:00 · Lenny 10.4% · guest 89.6%18:00 · Lenny 10.4% · guest 89.6%21:00 · Lenny 14.1% · guest 85.9%21:00 · Lenny 14.1% · guest 85.9%24:00 · Lenny 8.3% · guest 91.7%24:00 · Lenny 8.3% · guest 91.7%27:00 · Lenny 21.8% · guest 78.2%27:00 · Lenny 21.8% · guest 78.2%30:00 · Lenny 13% · guest 87%30:00 · Lenny 13% · guest 87%33:00 · Lenny 25.2% · guest 74.8%33:00 · Lenny 25.2% · guest 74.8%36:00 · Lenny 20.5% · guest 79.5%36:00 · Lenny 20.5% · guest 79.5%39:00 · Lenny 39.2% · guest 60.8%39:00 · Lenny 39.2% · guest 60.8%42:00 · Lenny 54.3% · guest 45.7%42:00 · Lenny 54.3% · guest 45.7%45:00 · Lenny 0.1% · guest 99.9%45:00 · Lenny 0.1% · guest 99.9%48:00 · Lenny 19.8% · guest 80.2%48:00 · Lenny 19.8% · guest 80.2%51:00 · Lenny 18.1% · guest 81.9%51:00 · Lenny 18.1% · guest 81.9%54:00 · Lenny 26.4% · guest 73.6%54:00 · Lenny 26.4% · guest 73.6%57:00 · Lenny 29.3% · guest 70.7%57:00 · Lenny 29.3% · guest 70.7%1:00:00 · Lenny 17.6% · guest 82.4%1:00:00 · Lenny 17.6% · guest 82.4%1:03:00 · Lenny 27.7% · guest 72.3%1:03:00 · Lenny 27.7% · guest 72.3%1:06:00 · Lenny 6% · guest 94%1:06:00 · Lenny 6% · guest 94%1:09:00 · Lenny 25.8% · guest 74.2%1:09:00 · Lenny 25.8% · guest 74.2%1:12:00 · Lenny 42.4% · guest 57.6%1:12:00 · Lenny 42.4% · guest 57.6%1:15:00 · Lenny 41.5% · guest 58.5%1:15:00 · Lenny 41.5% · guest 58.5%1:18:00 · Lenny 23.2% · guest 76.8%1:18:00 · Lenny 23.2% · guest 76.8%1:21:00 · Lenny 31.5% · guest 68.5%1:21:00 · Lenny 31.5% · guest 68.5%1:24:00 · Lenny 85.8% · guest 14.2%1:24:00 · Lenny 85.8% · guest 14.2%
Sharpest disagreement ▶ 36:25 Pushing back against spec-driven development dogmas

Alexander challenges the popular industry enthusiasm for spec-driven development championed by other coding tools, asserting that engineers fundamentally dislike writing specs and suggesting chatter-driven development instead.

Hardest push from Lenny ▶ 58:48 Critiquing the Atlas AI browser search experience

Lenny challenges the OpenAI browser approach by sharing his tweet where he abandoned Atlas due to the friction of forced AI search over classic web search.

Biggest teaching moment ▶ 1:11:05 Reframing AGI bottlenecks around human validation

Alexander reframes the typical hardware and compute timeline debate by proving that human multitasking, typing, and verification speeds are the real pacing constraints holding back compound AI acceleration.

Lenny holds their own ▶ 41:25 Synthesizing real-world agent experiments from Block

Lenny demonstrates deep technical breadth by drawing on Block CTO Dan G's agent Goose to test Alexander's theories on automated desktop agent workflows.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Operating Culture and Bottoms-Up Speed at OpenAI 4512 Lenny sets the stage by contrasting startup dynamics with OpenAI's unique scale and offers the 'ready fire aim' framing. Alexander politely refines the analogy, explaining that long-range vision is clear while immediate tactical product-market fit is discovered purely empirically.
Redefining Codex as a Software Engineering Teammate 4611 Lenny contrasts Codex with existing autocomplete tools like Cursor. Alexander educates Lenny on their broader vision of Codex as a full-lifecycle software teammate rather than an IDE-bound autocomplete widget.
Unlocking Growth: Moving from Cloud to Interactive Local Sandboxes 5611 Lenny references Andrej Karpathy's experiences debugging with Codex to explore what unlocked growth. Alexander explains the counterintuitive discovery that their fully cloud-hosted async architecture was too advanced for mainstream adoption compared to local sandbox interaction.
Codex Max Architecture and Context Compaction 3700 Lenny asks how Codex differentiates on core intelligence beyond models. Alexander delivers an in-depth technical explanation of the three-tier architecture powering Codex Max and how context compaction works across model, API, and harness layers.
Why Code Generation Is the Backbone of Universal Agents 5621 Lenny connects Codex's B2B agent nature with Nick Turley's consumer super assistant thesis. Alexander provides a thesis reframing: code execution is fundamentally the most reliable and composable primitive for any computer-using agent.
The Evolution of Development: Reviewing AI Code and New Workflows 5512 Lenny references Cursor CEO Michael Truell's spec-driven development concepts. Alexander pushes back gently on purely spec-driven workflows, noting developers dislike writing specs and suggesting chatter-driven and feedback-driven paradigms instead.
Mixed-Initiative AI and Contextual Desktop Workflows 6401 Lenny cites Block CTO Dan G's experience with the Goose agent watching engineer screens. Alexander agrees on contextual desktop intelligence but pinpoints code validation and review as the true critical bottleneck.
Sponsor Segment: Jira Product Discovery 4501 Following the Jira sponsor read, Alexander describes how PMs and designers prototype directly with Codex, highlighting the Sora Android app build in 28 days. Lenny expresses genuine disbelief at the tiny team size and compressed timeframe.
Evaluating Product Progress Beyond Synthetic Benchmarks 4511 Lenny inquires about internal KPIs and public benchmark reliance. Alexander clarifies that synthetic benchmarks are secondary to tracking D7 retention and reading gritty qualitative complaints on Reddit and Twitter.
Contextual Intelligence and the Strategy Behind Atlas Browser 5612 Lenny brings up his own public tweet criticizing Atlas browser's AI-only search UX. Alexander welcomes the friction and explains the strategic need for an AI-native browser rendering engine to enable unobtrusive contextual actions.
Practical Guidance for Deploying Codex on Hard Problems 3610 Lenny asks for tactical tips on getting started with Codex. Alexander advises against starting with trivial toy problems, urging users to immediately test Codex against their most difficult, gnarly bugs.
Essential Skills for Engineers and Self-Healing Infrastructure 4601 Lenny explores future essential engineering skills. Alexander highlights systems reasoning and self-healing infrastructure, sharing how Codex is currently trialed to be on call for babysitting its own expensive training runs.
AGI Timelines and Overcoming the Human Validation Bottleneck 4711 Lenny asks for Alexander's AGI timeline. Alexander offers an unconventional breakdown, asserting compute or model capability is not the current bottleneck, but rather human typing and verification throughput.
Recruiting for the Codex Team at OpenAI 2400 Alexander outlines recruiting needs for the Codex team, giving prospective candidates a specific mental rubric around visualizing engineer workflows over the next six months.
Lightning Round: Culture Sci-Fi, FSD UI, and Family Heritage 4300 In the lightning round, Alexander discusses Iain M. Banks' Culture series, Tesla's shared-control FSD UI as an agent design masterclass, and his family heritage in Andros, Greece.

Statements from this episode (27)

Opinion
The tech industry is far behind on AI product development
“I believe that even if we had no more progress today with models, which is absolutely not the case, but if, even if we had no more progress, we are way behind on product. There's so much more product to build.”
Alexander Embirikos Dec 14, 2025 ▶ 7:41
Disclosure
Proactive agents are critical to achieving OpenAI's AGI mission
“One of our major goals with Codex is to like get to proactivity. I think this is like critically important to like achieve the mission of OpenAI, which is to deliver the benefits of AGI to all humanity.”
Alexander Embirikos Dec 14, 2025 ▶ 14:02
Insight
Current AI products are difficult to use because they require manual prompting
“They're actually like really hard to use because you have to like be very thoughtful about when it could help you. And if you're not prompting a model to help you, it's probably not helping you at that time. And if you think of how many times like the average …”
Alexander Embirikos Dec 14, 2025 ▶ 14:17
Assertion Not checkable as stated
Codex usage grew 20x since August, serving trillions of weekly tokens
“The last stat we shared there was, like, we were, like, well over 10 X since August. In fact, it's been, like, 20 X since then. Also, the codex models are serving many, many trillions of tokens a week now, and it's basically, like, our most served coding model…”
Alexander Embirikos Dec 14, 2025 ▶ 16:01
Assertion Not checkable as stated
Codex is the most-served coding model on the OpenAI API
“We've reached the point where actually the codex model is the most served coding model in the API as well.”
Alexander Embirikos Dec 14, 2025 ▶ 16:50
Insight
AI coding agents must start locally before users delegate asynchronously
“The key unlock is actually first you need to land with users in a way that's, like, much more intuitive and, like, trivial to get value from. So the way that most people discover, like, the vast majority of users discover Codex today is either they download an…”
Alexander Embirikos Dec 14, 2025 ▶ 18:39
Insight
Internal dogfooding at AI labs skews product signals for the general market
“It just turns out that this is one of those places where the signal we got from dogfooding is a little bit different from the signal you get from like the general market, because at OpenAI, you know, we train reasoning models all day. And so we're very used to…”
Alexander Embirikos Dec 14, 2025 ▶ 20:40
Disclosure
OpenAI Codex executes code directly via shell inside a secure sandbox
“The way that we built Codex is that it just uses the shell, but in order to make that like safer and secure, we have a sandbox that the model is used to operating in.”
Alexander Embirikos Dec 14, 2025 ▶ 24:36
Insight
The best way for AI models to use computers is writing code
“It turns out the best way for models to use computers is simply to write code.”
Alexander Embirikos Dec 14, 2025 ▶ 29:03
Prediction Open · timeframe Dec 2030
Ubiquitous code generation will increase demand for human software engineering skills
“And as you make code even more ubiquitous, it's actually just going to be used for many more purposes. And so there's just going to be a ton more need for people with this, like humans with this competency.”
Alexander Embirikos Dec 14, 2025 ▶ 33:51
Insight
Structuring markdown plans with verifiable steps extends AI agent autonomy
“Like often collaborate with it first to write like a plan on MD, like a markdown file. That's your plan. And once you're happy with that, Then you ask it to go off and do work, and if that plan has verifiable steps, it'll, like, work for much longer.”
Alexander Embirikos Dec 14, 2025 ▶ 36:36
Assertion Supported
The Codex brand name was first used for GitHub Copilot's model
“The first time we used the brand codex at OpenAI was actually the model powering GitHub Copilot.”
Alexander Embirikos Dec 14, 2025 ▶ 39:29
Insight
Mixed-initiative AI works when upside is high and failure is low-friction
“Part of what's so magical about it is that When the, it can surface like ideas for helping you really rapidly. When it's right, you're accelerated. When it's wrong, it's not like that annoying. It can be annoying, but it's not that annoying. And so you can cre…”
Alexander Embirikos Dec 14, 2025 ▶ 39:55
Disclosure
OpenAI data scientists use Codex in Slack to diagnose metric changes
“Data scientists often here are using Codex a ton to just, like, answer questions, like, why do you think this metric moved? What happened? So questions, you know, you get the answer right back in Slack.”
Alexander Embirikos Dec 14, 2025 ▶ 42:08
Insight
Code validation and review are the primary bottlenecks in AI workflows
“The real, like, I think bottleneck right now is like validating that the code worked and like writing code review.”
Alexander Embirikos Dec 14, 2025 ▶ 42:26
Assertion Not checkable as stated
OpenAI built the Sora Android app in 28 days using Codex
“The Sora Android app, right, like a fully new app, We built it in 18 days. It went from like zero to launch to employees, and then 10 days later, so 28 days total, we went to just like GA, to the public. And that was done just like with the help of Codex.”
Alexander Embirikos Dec 14, 2025 ▶ 47:11
Assertion Contradicted
OpenAI's latest model is its first to natively understand PowerShell
“Like, just the model we shipped last week is the first model that natively understands PowerShell.”
Alexander Embirikos Dec 14, 2025 ▶ 50:02
Insight
Deep customer domain understanding beats raw building skills for AI startups
“If you're starting a new company today, and you have, like, a really good understanding and, like, network of customers that are currently underserved by AI tools, I think you're, like, you're set. Whereas if you're, like, good at building, like, you know, web…”
Alexander Embirikos Dec 14, 2025 ▶ 54:56
Insight
The Codex team prioritizes Day 7 retention over power-user edge cases
“One of the things that I'm constantly reminding myself of is that a tool like Codex sort of naturally is a tool that you would, you know, become a power user of. And so we can accidentally spend a lot of our time thinking about features that are like very deep…”
Alexander Embirikos Dec 14, 2025 ▶ 55:51
Opinion
Reddit provides more authentic Codex user feedback than hype on X
“Especially, I think for Twitter X it's a little bit more hypey, and then Reddit is a little more negative, but real, actually. So I've started increasingly paying attention to, like, how people are talking about using codecs on Reddit, actually.”
Alexander Embirikos Dec 14, 2025 ▶ 57:17
Insight
A dedicated AI browser provides superior context compared to desktop screenshots
“A lot of work is done in the web, and if we could build a browser, then we could be contextual for you, but in a much more first class way. We weren't hacking like other desktop software, which have like very varied support for like what content they're render…”
Alexander Embirikos Dec 14, 2025 ▶ 59:39
Insight
The most effective way to test Codex is on your hardest tasks
“The best way to try codex is to give it your hardest tasks. Which is a little different than some of the other coding agents. Like, you know, some tools you might think, okay, let me like start easy or just like, you know, like vibe code something random and d…”
Alexander Embirikos Dec 14, 2025 ▶ 1:03:12
Insight
AI coding agents are shrinking the experience gap for junior engineers
“They actually have, like, less of a handicap than before versus a more senior career person because, you know, the divide is actually getting smaller because they've got these amazing coding agents now.”
Alexander Embirikos Dec 14, 2025 ▶ 1:06:50
Disclosure
OpenAI uses Codex to manage infrastructure and review its own training runs
“Codex writes a lot of the code that helps, like, manage its training runs, the key infrastructure. You know, we move pretty fast, and so we have a Codex code review is, like, catching a lot of mistakes. It's actually caught some, like, pretty interesting confi…”
Alexander Embirikos Dec 14, 2025 ▶ 1:09:12
Insight
Human prompt writing and validation are the main limits to AI productivity
“I think that the current limiting factor, I mean, there's many, but I think a current underappreciated limiting factor is like literally human typing speed or human multitasking speed on like writing prompts. And like, you know, you were talking about, it's li…”
Alexander Embirikos Dec 14, 2025 ▶ 1:11:06
Prediction Not checkable as stated
AI-driven developer productivity will hockey-stick in 2026
“I think starting next year, we're gonna see, like, early adopters, like, starting to, like, hockey stick their productivity. And then over the years that follow, we're gonna see larger and larger companies, like, hockey stick that productivity. And then somewh…”
Alexander Embirikos Dec 14, 2025 ▶ 1:12:23
Opinion
Tesla's FSD is a masterclass in human-in-the-loop AI agent design
“I find the Tesla software, like, quite inspiring. In particular, it has, like, the self-driving feature, and, you know, I, I've mentioned a few times, like, today, like, I think it's really interesting to think about how to build, like, mixed initiative softwa…”
Alexander Embirikos Dec 14, 2025 ▶ 1:19:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.