Dec 26, 2025 · 27m · latent-space
⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI engineers Brian Fioca and Bill Chen join Latent Space to unpack the design, training, and evaluation behind Codex Max and GPT-5 agents. They explain how long-horizon autonomy, agent personality, sub-agent orchestration, and terminal-native tooling are transforming AI from code completion tools into universal autonomous assistants.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Bill Chen firmly steps in to disentangle the host's premise regarding startup tool conflict, arguing opinionated harnesses are an advantage rather than a limitation.
Hardest push from the hosts ▶ 8:51 Challenging model habit formation as a positive featureThe host rejects the guest's praise of models developing human-like tool habits, insisting that true model capability should generalize rather than anchor to specific tool names.
Biggest teaching moment ▶ 8:19 Ripgrep naming discovery boosting tool call performanceBill Chen and Brian Fioca reveal empirical findings that renaming tools from grep to rg significantly improved tool-call success due to terminal pretraining priors.
The host holds their own ▶ 21:34 Proposing batch multi-turn eval API architectureThe host leverages practical developer economics to argue why batch multi-turn APIs are necessary for cheap overnight evaluation runs, prompting the OpenAI team to accept it as direct product feedback.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Unveiling Codex Max and the Philosophy Behind the Name | 3 | 3 | 0 | 2 | The host asks clarifying questions regarding the Max naming convention and training methodologies beyond vague RL claims. The guests elaborate on training models with behavioral characteristics like planning, communication, and tool habits to cultivate developer trust. | |
| Differentiating Codex from Mainline Models and Tool Habits | 4 | 5 | 2 | 4 | The host questions whether models developing tool-specific habits is actually desirable given generalisation goals. The guests explain how Codex is specialized around terminal tools like ripgrep and why mainline models remain more general. | |
| Navigating Agent Verbosity, Personality Steerability, and Headless Runs | 5 | 3 | 1 | 5 | The host challenges the premise of model personality in long-running headless agentic workflows where verbose communication wastes tokens. The guest clarifies that communicative preambles build developer trust during interactive sessions but can be suppressed in headless execution. | |
| Shifting Abstractions: Packaging Full Agents and Dynamic Tooling | 3 | 4 | 0 | 1 | The host queries broader industry trends for coding models. The guests explain the architectural shift from bare model APIs to full packaged agent harnesses that can dynamically write their own tools and plugins. | |
| Sub-Agent Orchestration, Production Evals, and Internal Adoption | 3 | 4 | 0 | 1 | The host asks about sub-agent orchestration and multi-agent systems. The guests explain how Codex Max manages its own context compaction and rollout traces to enable autonomous parallel sub-agent workflows. | |
| The Crucial Role of Applied Evals and Multi-Turn Benchmarking | 6 | 4 | 1 | 4 | The host pushes on the difficulty of benchmarking multi-turn evaluations and pitches a specific feature request for a batch multi-turn eval API to run cheap overnight jobs. The guests welcome the feedback and describe their rollout eval methodologies. | |
| Expanding Coding Agents Beyond Code to Universal Terminal Automation | 5 | 3 | 1 | 3 | The host describes building a Slack-based agent for non-coding workflows and critiques coding agents for not being vision-native enough. The guests reframe terminal coding agents as general-purpose computer use agents. |