Dec 26, 2025 · 27m · latent-space

⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI

Brian Fioca · 12m spoken Bill Chen · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI engineers Brian Fioca and Bill Chen join Latent Space to unpack the design, training, and evaluation behind Codex Max and GPT-5 agents. They explain how long-horizon autonomy, agent personality, sub-agent orchestration, and terminal-native tooling are transforming AI from code completion tools into universal autonomous assistants.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.1 Guest teaching 3.7 Guest disagreement 0.7 The hosts pushing back 2.9
05100:0010:0020:001:20–5:04 · The hosts as informed peer 3/10 Unveiling Codex Max and the Philosophy Behind the Name The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust.5:05–9:15 · The hosts as informed peer 4/10 Differentiating Codex from Mainline Models and Tool Habits The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general.9:15–11:55 · The hosts as informed peer 5/10 Navigating Agent Verbosity, Personality Steerability, and Headless Runs The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution.11:55–14:44 · The hosts as informed peer 3/10 Shifting Abstractions: Packaging Full Agents and Dynamic Tooling The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins.14:45–17:21 · The hosts as informed peer 3/10 Sub-Agent Orchestration, Production Evals, and Internal Adoption The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows.17:22–22:27 · The hosts as informed peer 6/10 The Crucial Role of Applied Evals and Multi-Turn Benchmarking The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies.22:27–25:01 · The hosts as informed peer 5/10 Expanding Coding Agents Beyond Code to Universal Terminal Automation The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents.1:20–5:04 · Guest teaching 3/10 Unveiling Codex Max and the Philosophy Behind the Name The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust.5:05–9:15 · Guest teaching 5/10 Differentiating Codex from Mainline Models and Tool Habits The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general.9:15–11:55 · Guest teaching 3/10 Navigating Agent Verbosity, Personality Steerability, and Headless Runs The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution.11:55–14:44 · Guest teaching 4/10 Shifting Abstractions: Packaging Full Agents and Dynamic Tooling The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins.14:45–17:21 · Guest teaching 4/10 Sub-Agent Orchestration, Production Evals, and Internal Adoption The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows.17:22–22:27 · Guest teaching 4/10 The Crucial Role of Applied Evals and Multi-Turn Benchmarking The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies.22:27–25:01 · Guest teaching 3/10 Expanding Coding Agents Beyond Code to Universal Terminal Automation The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents.1:20–5:04 · Guest disagreement 0/10 Unveiling Codex Max and the Philosophy Behind the Name The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust.5:05–9:15 · Guest disagreement 2/10 Differentiating Codex from Mainline Models and Tool Habits The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general.9:15–11:55 · Guest disagreement 1/10 Navigating Agent Verbosity, Personality Steerability, and Headless Runs The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution.11:55–14:44 · Guest disagreement 0/10 Shifting Abstractions: Packaging Full Agents and Dynamic Tooling The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins.14:45–17:21 · Guest disagreement 0/10 Sub-Agent Orchestration, Production Evals, and Internal Adoption The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows.17:22–22:27 · Guest disagreement 1/10 The Crucial Role of Applied Evals and Multi-Turn Benchmarking The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies.22:27–25:01 · Guest disagreement 1/10 Expanding Coding Agents Beyond Code to Universal Terminal Automation The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents.1:20–5:04 · The hosts pushing back 2/10 Unveiling Codex Max and the Philosophy Behind the Name The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust.5:05–9:15 · The hosts pushing back 4/10 Differentiating Codex from Mainline Models and Tool Habits The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general.9:15–11:55 · The hosts pushing back 5/10 Navigating Agent Verbosity, Personality Steerability, and Headless Runs The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution.11:55–14:44 · The hosts pushing back 1/10 Shifting Abstractions: Packaging Full Agents and Dynamic Tooling The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins.14:45–17:21 · The hosts pushing back 1/10 Sub-Agent Orchestration, Production Evals, and Internal Adoption The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows.17:22–22:27 · The hosts pushing back 4/10 The Crucial Role of Applied Evals and Multi-Turn Benchmarking The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies.22:27–25:01 · The hosts pushing back 3/10 Expanding Coding Agents Beyond Code to Universal Terminal Automation The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 5:20 Disentangling Codex specialization from mainline general models

Bill Chen firmly steps in to disentangle the host's premise regarding startup tool conflict, arguing opinionated harnesses are an advantage rather than a limitation.

Hardest push from the hosts ▶ 8:51 Challenging model habit formation as a positive feature

The host rejects the guest's praise of models developing human-like tool habits, insisting that true model capability should generalize rather than anchor to specific tool names.

Biggest teaching moment ▶ 8:19 Ripgrep naming discovery boosting tool call performance

Bill Chen and Brian Fioca reveal empirical findings that renaming tools from grep to rg significantly improved tool-call success due to terminal pretraining priors.

The host holds their own ▶ 21:34 Proposing batch multi-turn eval API architecture

The host leverages practical developer economics to argue why batch multi-turn APIs are necessary for cheap overnight evaluation runs, prompting the OpenAI team to accept it as direct product feedback.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Unveiling Codex Max and the Philosophy Behind the Name 3302 The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust.
Differentiating Codex from Mainline Models and Tool Habits 4524 The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general.
Navigating Agent Verbosity, Personality Steerability, and Headless Runs 5315 The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution.
Shifting Abstractions: Packaging Full Agents and Dynamic Tooling 3401 The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins.
Sub-Agent Orchestration, Production Evals, and Internal Adoption 3401 The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows.
The Crucial Role of Applied Evals and Multi-Turn Benchmarking 6414 The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies.
Expanding Coding Agents Beyond Code to Universal Terminal Automation 5313 The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents.

Statements from this episode (19)

Assertion Supported
Fioca: Codex Max can run continuously for 24 hours or more
“Max can run for a really long time. We can go 24 hours or more. I've actually, like, sort of had it gone for more than that”
Brian Fioca Dec 26, 2025 ▶ 1:45
Disclosure
Fioca: OpenAI evaluates GPT-5 coding models on behavioral software engineering practices
“And so these are just best software engineering practices that turn out to be behavior characteristics, and we can measure the model's performance on those behaviors and grade it that way.”
Brian Fioca Dec 26, 2025 ▶ 4:00
Disclosure
Fioca: OpenAI trains models to flexibly adapt across varied developer toolsets
“Initially, you know, our models are trained the way they were trained to use tools, and that kind of bakes in a habit, and so we've been getting the models better at using different types of tools.”
Brian Fioca Dec 26, 2025 ▶ 4:50
Assertion Not checkable as stated
Fioca: Codex is OpenAI's frontier coding model optimized for its harness
“Codex is, just to be clear, Codex is the frontier coding model that we have that is optimized for its harness.”
Brian Fioca Dec 26, 2025 ▶ 5:20
Assertion Supported
Fioca: OpenAI Codex harness is open-source and the model is API-accessible
“Yes, that's open source, and the model is available in the API, so, so that's what they focus on.”
Brian Fioca Dec 26, 2025 ▶ 5:39
Insight
Chen: Naming custom tools identically to terminal tools boosts Codex performance
“We found some, like, partners of ours, like, they discovered that what you can do is that you can actually still have a lot of the tools just named in the same way as the terminal tools, as well as having the same input and output. And all of a sudden, the too…”
Bill Chen Dec 26, 2025 ▶ 8:05
Assertion Not checkable as stated
Chen: Codex performs better when search tools are named 'rg' over 'grep'
“So if you call it grep, it actually does a little bit Worse. But if you call it RG, it actually does really well.”
Bill Chen Dec 26, 2025 ▶ 8:24
Insight
Fioca: AI models develop operational habits during training analogous to muscle memory
“This is one of the coolest things about, like, model training is literally, like, they develop habits. It's just like a person does. Like, if you're, like, working on some podcasting tool, right, you're really good at editing, and then somebody makes you use a…”
Brian Fioca Dec 26, 2025 ▶ 8:37
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Brian Fioca Dec 26, 2025 ▶ 10:39
Assertion Supported
Fioca: GPT-5.1 allows disabling preambles, unlike the reasoning-dependent Codex model
“So Five One, you can turn that off, you can prompt it not to do that, but the Codex model can't actually do that, and it relies on the reasoning summarizer to give you that update.”
Brian Fioca Dec 26, 2025 ▶ 11:26
Insight
Chen: AI software abstraction is shifting from raw models to packaged agents
“So we're actually shipping this Entirety, entire agent altogether, then you can actually build on top of that agent. That's one of the patterns that we're seeing here is rather than focusing on optimizing with every single model release, you're actually just b…”
Bill Chen Dec 26, 2025 ▶ 12:28
Insight
Fioca: Coding agents enable self-customizing software by writing integrations at runtime
“So now if it doesn't have a tool, it can make a tool that it needs to solve a problem, right? So that's like another layer of abstraction and it's not just coding. You can write software that has an agent that can spin up a codex instance and write a custom pl…”
Brian Fioca Dec 26, 2025 ▶ 13:40
Assertion Supported
Fioca: Codex Max manages its own context window to run indefinitely
“Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness.”
Brian Fioca Dec 26, 2025 ▶ 14:54
Disclosure
Fioca: Has not hand-written a single line of code in months
“I haven't written a single line of code by hand in months, because I know what I can trust it to do.”
Brian Fioca Dec 26, 2025 ▶ 16:06
Assertion Not checkable as stated
Chen: Approximately 50% of OpenAI employees adopted Codex at launch
“Initially when Codex first launched, it was around 50% of folks that open AI started using it.”
Bill Chen Dec 26, 2025 ▶ 16:28
Insight
Fioca: Raw frontier models resemble new PhD hires needing explicit job prompts
“I like to think of it as like we have, I mean, people say it's a PhD in, in an API, right? But you, if you know, you hire a PhD student, they don't know how to do the job. You have to give them a job description. Okay. That's a prompt, right? So now you have y…”
Brian Fioca Dec 26, 2025 ▶ 18:30
Assertion Supported
Chen: OpenAI's Batch API does not yet support multi-turn requests
“Batch multi-turn requests. I don't believe it. You can't do it yet.”
Bill Chen Dec 26, 2025 ▶ 21:52
Insight
Chen: Terminal coding agents are actually general-purpose computer-use agents
“So what would you think about it is are those coding agents are actually a computer use agent, but for the terminal. They're actually incredibly general.”
Bill Chen Dec 26, 2025 ▶ 24:42
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Bill Chen Dec 26, 2025 ▶ 25:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.