Jun 19, 2025 · 1h 17m · latent-space
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI researcher Noam Brown joins the Latent Space Podcast to discuss the evolution of test-time reasoning models like o1 and o3, lessons from game-theoretic breakthroughs like Cicero and Libratus, and the future of scaling multi-agent civilizations.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Noam firmly rejects the host's framing that executing real-world actions and reverting after feedback counts as test-time compute, emphasizing that simulation before execution is fundamentally distinct.
Hardest push from the hosts ▶ 44:50 Pressing on flawed multi-agent approachesWhen Noam refuses to disclose OpenAI's current multi-agent work, Swyx immediately pushes back by demanding he specify exactly what approaches in the existing literature are misguided.
Biggest teaching moment ▶ 56:30 Explaining why self-play fails outside two-player zero-sum gamesNoam deconstructs the popular assumption that AlphaGo-style self-play is a direct path to AGI, demonstrating that minimax convergence guarantees collapse when applied to multiplayer or non-zero-sum domains.
The host holds their own ▶ 32:27 Citing David Luan on structural lab differencesSwyx demonstrates deep insider domain knowledge by citing former VP of Engineering David Luan to explain why Google Brain failed to scale models while OpenAI succeeded due to centralized compute pooling.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Reflections on Cicero, Diplomacy Championships, and AI Steerability | 6 | 4 | 1 | 2 | Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability. | |
| Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains | 5 | 5 | 2 | 2 | The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel. | |
| System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition | 5 | 4 | 1 | 2 | Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains. | |
| Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation | 6 | 6 | 3 | 3 | Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute. | |
| Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning | 5 | 5 | 2 | 2 | Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning. | |
| The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever | 6 | 5 | 2 | 2 | Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training. | |
| Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization | 7 | 4 | 1 | 2 | Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation. | |
| AI-Assisted Software Engineering, Codex, and Alignment Dimensions | 5 | 4 | 1 | 2 | The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment. | |
| Scaling Test-Time Compute to Multi-Agent Civilizations | 5 | 6 | 3 | 3 | Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson. | |
| Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling | 6 | 7 | 1 | 2 | Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO. | |
| World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games | 6 | 7 | 2 | 3 | Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains. | |
| Generative Media Trends and the Iteration Friction of Robotics Hardware | 6 | 3 | 1 | 1 | The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI. | |
| Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles | 5 | 5 | 2 | 2 | Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry. | |
| Social Deduction Games, Imperfect Information Scaling, and Farewell | 6 | 7 | 1 | 1 | Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states. |