May 16, 2025 · 53m · latent-space

ChatGPT Codex: The Missing Manual

Alexander Embiricos · 22m spoken Josh Ma · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI engineers Josh and Alexander join the Latent Space Podcast to detail the architecture, developer best practices, and design philosophy behind ChatGPT Codex. They discuss shifting from synchronous code autocomplete to an asynchronous delegation model powered by cloud sandboxes, model-native reasoning, and standardized repository instructions.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.5 Guest teaching 3.0 Guest disagreement 1.3 The hosts pushing back 1.6
05100:0015:0030:0045:001:59–8:40 · The hosts as informed peer 2/10 Origins of Codex: From Local CLI to Cloud Agent Alexander and Josh share their personal backgrounds and how they joined OpenAI to work on agentic coding tools. The hosts offer supportive opening banter and prompt the guests' origin stories.8:41–14:10 · The hosts as informed peer 4/10 Codex Architecture: Independent Cloud Agents and Pull Requests Alessio asks about the differences between Codex CLI and the hosted cloud agent, sharing an early test run on a Rails codebase. Alexander and Josh describe how cloud agents autonomously run tests, infer style, and write verifiable PR summaries.14:10–21:27 · The hosts as informed peer 5/10 Best Practices: Formatting, Modularity, and Repository Setup Wix draws on software engineering experience to highlight how linters and commit hooks act as in-the-loop verifiers for AI agents. The guests detail naming conventions and architectural modularity strategies that assist agent exploration.21:28–25:38 · The hosts as informed peer 4/10 Instruction Design and the agents.md Standard Alessio presses the guests on why agents require a dedicated agents.md specification instead of conventional README or contributor docs. Josh and Alexander explain the difference between human onboarding and agent instruction scoping.25:39–36:34 · The hosts as informed peer 6/10 Pushing Intelligence into the Model Versus Hardcoded Scaffolding Wix questions context window limits and challenges the slow feedback loop of retraining models rather than writing deterministic scaffolding. Josh and Alexander explain OpenAI's core philosophy of pushing state machines and planning directly into model training.36:37–41:26 · The hosts as informed peer 6/10 Task Concurrency and the Abundance Mindset Wix references runtime metrics from the METR paper and Operator benchmarks to probe Codex task length limits. Alexander explains the need for an abundance mindset where users trigger dozens of parallel tasks rather than watching an agent code synchronously.41:28–47:18 · The hosts as informed peer 5/10 Compute Sandboxing, Execution Safety, and Single-Shot Autonomy Alessio and Wix ask about cloud environment access, REPL debugging, and security sandboxing. Josh explains network isolation during agent runs and contrasts their single-shot autonomous goal against multi-turn human-in-the-loop architectures.47:19–53:21 · The hosts as informed peer 4/10 Research Preview Roadmap and Community Call to Action The hosts ask about onboarding easter eggs and raise upcoming pricing concerns based on competitor moves. Alexander and Josh share their roadmap for environment customization and emphasize value-driven pricing.1:59–8:40 · Guest teaching 1/10 Origins of Codex: From Local CLI to Cloud Agent Alexander and Josh share their personal backgrounds and how they joined OpenAI to work on agentic coding tools. The hosts offer supportive opening banter and prompt the guests' origin stories.8:41–14:10 · Guest teaching 3/10 Codex Architecture: Independent Cloud Agents and Pull Requests Alessio asks about the differences between Codex CLI and the hosted cloud agent, sharing an early test run on a Rails codebase. Alexander and Josh describe how cloud agents autonomously run tests, infer style, and write verifiable PR summaries.14:10–21:27 · Guest teaching 3/10 Best Practices: Formatting, Modularity, and Repository Setup Wix draws on software engineering experience to highlight how linters and commit hooks act as in-the-loop verifiers for AI agents. The guests detail naming conventions and architectural modularity strategies that assist agent exploration.21:28–25:38 · Guest teaching 4/10 Instruction Design and the agents.md Standard Alessio presses the guests on why agents require a dedicated agents.md specification instead of conventional README or contributor docs. Josh and Alexander explain the difference between human onboarding and agent instruction scoping.25:39–36:34 · Guest teaching 5/10 Pushing Intelligence into the Model Versus Hardcoded Scaffolding Wix questions context window limits and challenges the slow feedback loop of retraining models rather than writing deterministic scaffolding. Josh and Alexander explain OpenAI's core philosophy of pushing state machines and planning directly into model training.36:37–41:26 · Guest teaching 3/10 Task Concurrency and the Abundance Mindset Wix references runtime metrics from the METR paper and Operator benchmarks to probe Codex task length limits. Alexander explains the need for an abundance mindset where users trigger dozens of parallel tasks rather than watching an agent code synchronously.41:28–47:18 · Guest teaching 3/10 Compute Sandboxing, Execution Safety, and Single-Shot Autonomy Alessio and Wix ask about cloud environment access, REPL debugging, and security sandboxing. Josh explains network isolation during agent runs and contrasts their single-shot autonomous goal against multi-turn human-in-the-loop architectures.47:19–53:21 · Guest teaching 2/10 Research Preview Roadmap and Community Call to Action The hosts ask about onboarding easter eggs and raise upcoming pricing concerns based on competitor moves. Alexander and Josh share their roadmap for environment customization and emphasize value-driven pricing.1:59–8:40 · Guest disagreement 0/10 Origins of Codex: From Local CLI to Cloud Agent Alexander and Josh share their personal backgrounds and how they joined OpenAI to work on agentic coding tools. The hosts offer supportive opening banter and prompt the guests' origin stories.8:41–14:10 · Guest disagreement 1/10 Codex Architecture: Independent Cloud Agents and Pull Requests Alessio asks about the differences between Codex CLI and the hosted cloud agent, sharing an early test run on a Rails codebase. Alexander and Josh describe how cloud agents autonomously run tests, infer style, and write verifiable PR summaries.14:10–21:27 · Guest disagreement 1/10 Best Practices: Formatting, Modularity, and Repository Setup Wix draws on software engineering experience to highlight how linters and commit hooks act as in-the-loop verifiers for AI agents. The guests detail naming conventions and architectural modularity strategies that assist agent exploration.21:28–25:38 · Guest disagreement 2/10 Instruction Design and the agents.md Standard Alessio presses the guests on why agents require a dedicated agents.md specification instead of conventional README or contributor docs. Josh and Alexander explain the difference between human onboarding and agent instruction scoping.25:39–36:34 · Guest disagreement 3/10 Pushing Intelligence into the Model Versus Hardcoded Scaffolding Wix questions context window limits and challenges the slow feedback loop of retraining models rather than writing deterministic scaffolding. Josh and Alexander explain OpenAI's core philosophy of pushing state machines and planning directly into model training.36:37–41:26 · Guest disagreement 1/10 Task Concurrency and the Abundance Mindset Wix references runtime metrics from the METR paper and Operator benchmarks to probe Codex task length limits. Alexander explains the need for an abundance mindset where users trigger dozens of parallel tasks rather than watching an agent code synchronously.41:28–47:18 · Guest disagreement 1/10 Compute Sandboxing, Execution Safety, and Single-Shot Autonomy Alessio and Wix ask about cloud environment access, REPL debugging, and security sandboxing. Josh explains network isolation during agent runs and contrasts their single-shot autonomous goal against multi-turn human-in-the-loop architectures.47:19–53:21 · Guest disagreement 1/10 Research Preview Roadmap and Community Call to Action The hosts ask about onboarding easter eggs and raise upcoming pricing concerns based on competitor moves. Alexander and Josh share their roadmap for environment customization and emphasize value-driven pricing.1:59–8:40 · The hosts pushing back 0/10 Origins of Codex: From Local CLI to Cloud Agent Alexander and Josh share their personal backgrounds and how they joined OpenAI to work on agentic coding tools. The hosts offer supportive opening banter and prompt the guests' origin stories.8:41–14:10 · The hosts pushing back 1/10 Codex Architecture: Independent Cloud Agents and Pull Requests Alessio asks about the differences between Codex CLI and the hosted cloud agent, sharing an early test run on a Rails codebase. Alexander and Josh describe how cloud agents autonomously run tests, infer style, and write verifiable PR summaries.14:10–21:27 · The hosts pushing back 1/10 Best Practices: Formatting, Modularity, and Repository Setup Wix draws on software engineering experience to highlight how linters and commit hooks act as in-the-loop verifiers for AI agents. The guests detail naming conventions and architectural modularity strategies that assist agent exploration.21:28–25:38 · The hosts pushing back 2/10 Instruction Design and the agents.md Standard Alessio presses the guests on why agents require a dedicated agents.md specification instead of conventional README or contributor docs. Josh and Alexander explain the difference between human onboarding and agent instruction scoping.25:39–36:34 · The hosts pushing back 4/10 Pushing Intelligence into the Model Versus Hardcoded Scaffolding Wix questions context window limits and challenges the slow feedback loop of retraining models rather than writing deterministic scaffolding. Josh and Alexander explain OpenAI's core philosophy of pushing state machines and planning directly into model training.36:37–41:26 · The hosts pushing back 2/10 Task Concurrency and the Abundance Mindset Wix references runtime metrics from the METR paper and Operator benchmarks to probe Codex task length limits. Alexander explains the need for an abundance mindset where users trigger dozens of parallel tasks rather than watching an agent code synchronously.41:28–47:18 · The hosts pushing back 1/10 Compute Sandboxing, Execution Safety, and Single-Shot Autonomy Alessio and Wix ask about cloud environment access, REPL debugging, and security sandboxing. Josh explains network isolation during agent runs and contrasts their single-shot autonomous goal against multi-turn human-in-the-loop architectures.47:19–53:21 · The hosts pushing back 2/10 Research Preview Roadmap and Community Call to Action The hosts ask about onboarding easter eggs and raise upcoming pricing concerns based on competitor moves. Alexander and Josh share their roadmap for environment customization and emphasize value-driven pricing.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 26:55 Rejecting context reification premise

Josh directly refutes Wix's assumption that agents.md is simply prepended to the system prompt, clarifying that the model actively searches and greps the repository dynamically.

Hardest push from the hosts ▶ 34:15 Challenging model-first iteration cycles

Wix challenges Alexander's pure model-first philosophy by arguing that without online learning, relying entirely on retraining makes product iteration painfully slow compared to software scaffolding.

Biggest teaching moment ▶ 27:30 Rejecting hardcoded scaffolding for model training

Josh educates Wix on OpenAI's internal philosophy, explaining how researchers resisted hardcoded tool guardrails and prompt heuristics in favor of training models to generalize.

The host holds their own ▶ 37:40 Citing the METR benchmark on autonomy horizon

Wix demonstrates technical domain expertise by connecting Codex's task duration limits directly to the METR evaluation paper and operator benchmark data.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Codex: From Local CLI to Cloud Agent 2100 Alexander and Josh share their personal backgrounds and how they joined OpenAI to work on agentic coding tools. The hosts offer supportive opening banter and prompt the guests' origin stories.
Codex Architecture: Independent Cloud Agents and Pull Requests 4311 Alessio asks about the differences between Codex CLI and the hosted cloud agent, sharing an early test run on a Rails codebase. Alexander and Josh describe how cloud agents autonomously run tests, infer style, and write verifiable PR summaries.
Best Practices: Formatting, Modularity, and Repository Setup 5311 Wix draws on software engineering experience to highlight how linters and commit hooks act as in-the-loop verifiers for AI agents. The guests detail naming conventions and architectural modularity strategies that assist agent exploration.
Instruction Design and the agents.md Standard 4422 Alessio presses the guests on why agents require a dedicated agents.md specification instead of conventional README or contributor docs. Josh and Alexander explain the difference between human onboarding and agent instruction scoping.
Pushing Intelligence into the Model Versus Hardcoded Scaffolding 6534 Wix questions context window limits and challenges the slow feedback loop of retraining models rather than writing deterministic scaffolding. Josh and Alexander explain OpenAI's core philosophy of pushing state machines and planning directly into model training.
Task Concurrency and the Abundance Mindset 6312 Wix references runtime metrics from the METR paper and Operator benchmarks to probe Codex task length limits. Alexander explains the need for an abundance mindset where users trigger dozens of parallel tasks rather than watching an agent code synchronously.
Compute Sandboxing, Execution Safety, and Single-Shot Autonomy 5311 Alessio and Wix ask about cloud environment access, REPL debugging, and security sandboxing. Josh explains network isolation during agent runs and contrasts their single-shot autonomous goal against multi-turn human-in-the-loop architectures.
Research Preview Roadmap and Community Call to Action 4212 The hosts ask about onboarding easter eggs and raise upcoming pricing concerns based on competitor moves. Alexander and Josh share their roadmap for environment customization and emphasize value-driven pricing.

Statements from this episode (15)

Prediction Not checkable as stated
Ma: Industry will build an agentic software engineer within two years
“Whether or not I was involved in the next two years, I think we were going, we are going to build an agentic software engineer.”
Josh Ma May 16, 2025 ▶ 6:59
Insight
Embiricos: Benchmark-Passing SWE Agent Outputs Are Often Unmergeable in Practice
“Because if you look at a lot of, like, Sweebench passing, like, outputs from, like, an agent, they're not really, like, PRs that you would merge, because, like, the code style might be, like, different. Like, it works, but the code style is different.”
Alexander Embiricos May 16, 2025 ▶ 10:36
Prediction Not checkable as stated
Embiricos: Majority of Future Code May Be Written by Parallel AI Agents
“In, in a future world that we imagine where actually you know, maybe the majority of code is actually being written by agents that we're delegating to, you know, doing tasks in parallel. It becomes, like, critically important that you can actually, like, integ…”
Alexander Embiricos May 16, 2025 ▶ 11:23
Assertion Supported
Ma: ChatGPT Codex recognizes sub-directory instruction hierarchies in agents.md
“We put a lot of effort into making sure the agent, like, understand this hierarchy of instructions, right? You can put them in sub-directories, and it'll understand which ones take precedence over which others.”
Josh Ma May 16, 2025 ▶ 15:01
Insight
Embiricos: Clean architecture is more important than ever with AI agents
“Good architecture is like even more important than ever. And like, I guess the fun thing is like for now, that's something that humans are really good at. So like, you know, kind of good, you know, important for the software engineers to do their job.”
Alexander Embiricos May 16, 2025 ▶ 19:52
Opinion
Josh Ma: Software engineering stays human-centric due to review and deployment
“As long as you see humans and AI writing it, like maybe there's a world where it's only AI's maintaining a code base and the assumptions change, but the moment you start to break that fourth wall and a human's coming in, doing code review, deploying the code, …”
Josh Ma May 16, 2025 ▶ 21:40
Disclosure
Embiricos: OpenAI open-sourced Codex CLI to standardize agent safety
“Part of why we made the Codex CLI open source is, like, a lot of problems, like, safety issues that you need to figure out for how to deploy these things safely, and no one should have to figure these out, like, more than once. So, that's why we went for, like…”
Alexander Embiricos May 16, 2025 ▶ 24:45
Insight
Josh Ma: AI agents naturally infer code style without explicit instructions
“For agents, I don't think you have really had to tell code style. It looks at your code base and just writes code that's consistent to that. Whereas like a human's not going to take its time sorry, their time to go through the code base and you know, follow al…”
Josh Ma May 16, 2025 ▶ 25:09
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Alexander Embiricos May 16, 2025 ▶ 30:37
Insight
Embiricos: Specialized domain training yields outsized returns in general models
“If you can, like, build, do something very specific for, like, a specific purpose, actually, when you bring that and you bring it into the generalized model, like, you might even get outsized returns on that. Because there's, like, transfer from all these diff…”
Alexander Embiricos May 16, 2025 ▶ 36:15
Disclosure
Ma: ChatGPT Codex enforces a hard one-hour runtime limit per task
“Our hard call is an hour right now, although don't hold us to that. It may change over time.”
Josh Ma May 16, 2025 ▶ 37:07
Insight
Embiricos: The best Codex users spend 30 seconds max prompt crafting
“The way we see people who, like, love Codex the most using it is they don't, they think for, like, maybe 30 seconds max about their prompt. It's just like, oh, I have this idea, like, boom. Oh, like, there's this thing I wanna do, like, boom. Oh, like, I just …”
Alexander Embiricos May 16, 2025 ▶ 39:48
Disclosure
Josh Ma: ChatGPT Codex cuts off internet access during agent execution
“Once the agent starts running, right what we actually do today, and we're hoping to like evolve on this, is we'll cut off internet access because we still don't fully understand what letting loose an agent in his own environment is going to do.”
Josh Ma May 16, 2025 ▶ 43:04
Disclosure
Josh Ma: Codex focuses on pushing single-shot autonomous software engineering
“I think what we see as, like, the role of Codex here is to really push Frontier on that sort of single-shot autonomous software engineering.”
Josh Ma May 16, 2025 ▶ 46:00
Disclosure
Embiricos: OpenAI Prioritizing Multimodal Inputs and Tool Integration for Codex
“Some of the items that are top of mind for me are, like, multimodal inputs. You know, we've talked, yeah, I know you find that, right? Yeah, like, another, another example would be, like, you know, just giving it a little bit more access to the world. You know…”
Alexander Embiricos May 16, 2025 ▶ 47:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.