Jun 19, 2025 · 1h 17m · latent-space

Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI

Noam Brown · 51m spoken Spooks (Swyx) · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researcher Noam Brown joins the Latent Space Podcast to discuss the evolution of test-time reasoning models like o1 and o3, lessons from game-theoretic breakthroughs like Cicero and Libratus, and the future of scaling multi-agent civilizations.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.6 Guest teaching 5.1 Guest disagreement 1.6 The hosts pushing back 2.1
05100:0020:0040:001:00:000:05–5:24 · The hosts as informed peer 6/10 Reflections on Cicero, Diplomacy Championships, and AI Steerability Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability.5:25–8:56 · The hosts as informed peer 5/10 Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel.8:57–13:57 · The hosts as informed peer 5/10 System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains.13:57–17:14 · The hosts as informed peer 6/10 Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute.17:15–21:50 · The hosts as informed peer 5/10 Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning.21:51–29:35 · The hosts as informed peer 6/10 The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training.29:35–33:22 · The hosts as informed peer 7/10 Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation.33:22–41:34 · The hosts as informed peer 5/10 AI-Assisted Software Engineering, Codex, and Alignment Dimensions The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment.41:35–45:08 · The hosts as informed peer 5/10 Scaling Test-Time Compute to Multi-Agent Civilizations Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson.45:08–52:07 · The hosts as informed peer 6/10 Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO.52:07–58:50 · The hosts as informed peer 6/10 World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains.58:52–1:04:23 · The hosts as informed peer 6/10 Generative Media Trends and the Iteration Friction of Robotics Hardware The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI.1:04:24–1:13:08 · The hosts as informed peer 5/10 Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry.1:13:09–1:17:36 · The hosts as informed peer 6/10 Social Deduction Games, Imperfect Information Scaling, and Farewell Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states.0:05–5:24 · Guest teaching 4/10 Reflections on Cicero, Diplomacy Championships, and AI Steerability Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability.5:25–8:56 · Guest teaching 5/10 Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel.8:57–13:57 · Guest teaching 4/10 System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains.13:57–17:14 · Guest teaching 6/10 Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute.17:15–21:50 · Guest teaching 5/10 Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning.21:51–29:35 · Guest teaching 5/10 The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training.29:35–33:22 · Guest teaching 4/10 Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation.33:22–41:34 · Guest teaching 4/10 AI-Assisted Software Engineering, Codex, and Alignment Dimensions The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment.41:35–45:08 · Guest teaching 6/10 Scaling Test-Time Compute to Multi-Agent Civilizations Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson.45:08–52:07 · Guest teaching 7/10 Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO.52:07–58:50 · Guest teaching 7/10 World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains.58:52–1:04:23 · Guest teaching 3/10 Generative Media Trends and the Iteration Friction of Robotics Hardware The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI.1:04:24–1:13:08 · Guest teaching 5/10 Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry.1:13:09–1:17:36 · Guest teaching 7/10 Social Deduction Games, Imperfect Information Scaling, and Farewell Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states.0:05–5:24 · Guest disagreement 1/10 Reflections on Cicero, Diplomacy Championships, and AI Steerability Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability.5:25–8:56 · Guest disagreement 2/10 Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel.8:57–13:57 · Guest disagreement 1/10 System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains.13:57–17:14 · Guest disagreement 3/10 Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute.17:15–21:50 · Guest disagreement 2/10 Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning.21:51–29:35 · Guest disagreement 2/10 The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training.29:35–33:22 · Guest disagreement 1/10 Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation.33:22–41:34 · Guest disagreement 1/10 AI-Assisted Software Engineering, Codex, and Alignment Dimensions The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment.41:35–45:08 · Guest disagreement 3/10 Scaling Test-Time Compute to Multi-Agent Civilizations Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson.45:08–52:07 · Guest disagreement 1/10 Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO.52:07–58:50 · Guest disagreement 2/10 World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains.58:52–1:04:23 · Guest disagreement 1/10 Generative Media Trends and the Iteration Friction of Robotics Hardware The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI.1:04:24–1:13:08 · Guest disagreement 2/10 Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry.1:13:09–1:17:36 · Guest disagreement 1/10 Social Deduction Games, Imperfect Information Scaling, and Farewell Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states.0:05–5:24 · The hosts pushing back 2/10 Reflections on Cicero, Diplomacy Championships, and AI Steerability Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability.5:25–8:56 · The hosts pushing back 2/10 Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel.8:57–13:57 · The hosts pushing back 2/10 System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains.13:57–17:14 · The hosts pushing back 3/10 Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute.17:15–21:50 · The hosts pushing back 2/10 Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning.21:51–29:35 · The hosts pushing back 2/10 The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training.29:35–33:22 · The hosts pushing back 2/10 Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation.33:22–41:34 · The hosts pushing back 2/10 AI-Assisted Software Engineering, Codex, and Alignment Dimensions The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment.41:35–45:08 · The hosts pushing back 3/10 Scaling Test-Time Compute to Multi-Agent Civilizations Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson.45:08–52:07 · The hosts pushing back 2/10 Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO.52:07–58:50 · The hosts pushing back 3/10 World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains.58:52–1:04:23 · The hosts pushing back 1/10 Generative Media Trends and the Iteration Friction of Robotics Hardware The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI.1:04:24–1:13:08 · The hosts pushing back 2/10 Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry.1:13:09–1:17:36 · The hosts pushing back 1/10 Social Deduction Games, Imperfect Information Scaling, and Farewell Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 16:48 Dismissing real-world rollback as test-time compute

Noam firmly rejects the host's framing that executing real-world actions and reverting after feedback counts as test-time compute, emphasizing that simulation before execution is fundamentally distinct.

Hardest push from the hosts ▶ 44:50 Pressing on flawed multi-agent approaches

When Noam refuses to disclose OpenAI's current multi-agent work, Swyx immediately pushes back by demanding he specify exactly what approaches in the existing literature are misguided.

Biggest teaching moment ▶ 56:30 Explaining why self-play fails outside two-player zero-sum games

Noam deconstructs the popular assumption that AlphaGo-style self-play is a direct path to AGI, demonstrating that minimax convergence guarantees collapse when applied to multiplayer or non-zero-sum domains.

The host holds their own ▶ 32:27 Citing David Luan on structural lab differences

Swyx demonstrates deep insider domain knowledge by citing former VP of Engineering David Luan to explain why Google Brain failed to scale models while OpenAI succeeded due to centralized compute pooling.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Reflections on Cicero, Diplomacy Championships, and AI Steerability 6412 Alessio and Swyx demonstrate strong context by citing Cicero's 2.7B parameter size and tournament gameplay history. Noam elaborates on bot debugging and why the safety community actually embraced Cicero's action-conditioned steerability.
Evolution of OpenAI's o-Series and Reasoning in Non-Verifiable Domains 5522 The hosts prompt Noam on whether reasoning models can generalize beyond verifiable domains like math and code. Noam points to Deep Research as an empirical existence proof that verifiable reward functions are not strictly necessary for reasoning models to excel.
System 1 vs. System 2 Reasoning, Emergence, and Multimodal Intuition 5412 Swyx probes the breakdown of the Thinking Fast and Slow analogy and asks whether reasoning capability is emergent. Noam explains that models require a foundational baseline of System 1 capability before System 2 compute yields meaningful gains.
Scaffolding, Tool Use, and the Boundaries of Test-Time Simulation 6633 Alessio and Swyx discuss harnesses and search mechanisms like MCTS. Noam takes a firm stance against complex scaffolding, asserting that ideal models should operate harness-free and that real-world trial-and-error cannot be conflated with test-time compute.
Model Routers, Scaling Scaffolds Away, and Reinforcement Fine-Tuning 5522 Swyx asks whether model routers require smart judges, prompting Noam to argue that both routers and scaffolding will eventually be rendered obsolete by scale. Alessio clarifies practical engineering focus by distinguishing between ephemeral harnesses and durable data collection via reinforcement fine-tuning.
The Genesis of Reasoning at OpenAI and Historical Scaling Bets with Ilya Sutskever 6522 Swyx brings up early OpenAI RL efforts (GPT-Zero) and Noam's past conversations with Ilya Sutskever. Noam clarifies that Ilya's intuition was right regarding scaling RL and details the internal controversies at OpenAI over allocating finite compute away from pre-training.
Sample Efficiency Gaps, Ilya's Research Philosophy, and Lab Organization 7412 Swyx raises sample efficiency differences between humans and machines, citing David Luan's comparison between Google Brain and OpenAI's centralized betting culture. Noam agrees and elaborates on how OpenAI's startup structure enabled concentrated resource allocation.
AI-Assisted Software Engineering, Codex, and Alignment Dimensions 5412 The conversation shifts to everyday workflows with Codex, Windsurf, and o3. Alessio and Swyx explore bottlenecks in the software engineering lifecycle like PR reviews, while Noam highlights the first-day-on-the-job problem and discusses safety versus user alignment.
Scaling Test-Time Compute to Multi-Agent Civilizations 5633 Alessio queries Noam about OpenAI's multi-agent team and references Voyager's skill library. Noam remains deliberately evasive about proprietary details while sharing his philosophical framework of AI civilizations progressing from cavemen to modern societies via the Bitter Lesson.
Game Theory in Poker and Diplomacy: GTO vs. Exploitative Opponent Modeling 6712 Alessio asks about Game Theory Optimal (GTO) play versus exploitative play in poker and opponent HUD tracking. Noam delivers a masterclass on minimax equilibrium limits in multiplayer settings, explaining why Diplomacy necessitated opponent modeling over pure GTO.
World Models, Theory of Mind, and the Breakdown of Self-Play Beyond Zero-Sum Games 6723 Swyx asks whether explicit world models and self-play like AlphaZero can generalize to all domains. Noam pushes back, explaining that self-play guarantees convergence to minimax only in two-player zero-sum games and breaks down in open-ended or collaborative domains.
Generative Media Trends and the Iteration Friction of Robotics Hardware 6311 The discussion covers generative image diffusion versus autoregressive models and humanoid robotics. Alessio references Richard Hamming's insights on technological shifts, while Noam explains why hardware iteration friction drove him away from robotics into pure software AI.
Research Practices, Wall-Clock Bottlenecks in Long Reasoning, and Model Lifecycles 5522 Swyx asks about scaling limits for test-time compute by 2030 and requests definitions for mid-training. Noam highlights the serial wall-clock time bottleneck when reasoning takes days or weeks and jokingly notes that mid-training definitions remain fuzzy across the industry.
Social Deduction Games, Imperfect Information Scaling, and Farewell 6711 Alessio asks how game complexity scales from 52-card poker to imperfect information games with massive state spaces like Magic: The Gathering. Noam breaks down the exact math behind poker's 1,326 hidden states versus Stratego's 40 factorial states.

Statements from this episode (30)

Insight
Brown: Debugging game AI requires deep mastery to spot novel brilliance
“When you work on these games, you kind of have to understand the game well enough to like be able to debug your bot because If the bot does something that's, like, really radical and, like, that humans typically wouldn't do, you're not sure if that's, like, a …”
Noam Brown Jun 19, 2025 ▶ 0:53
Assertion Supported
Brown won the 2025 World Diplomacy Championship
“When we released Cicero, we announced it in, like, late twenty-twenty-two, I still found the game, like, really fascinating, and so I, like, kept up with it, I, like, continued to play, and that led to me winning the championship in the World Championship in t…”
Noam Brown Jun 19, 2025 ▶ 1:32
Assertion Not checkable as stated
Brown: GPT-4o and o3 are passing the Turing test
“So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.”
Noam Brown Jun 19, 2025 ▶ 3:28
Prediction Not checkable as stated
Brown: Reasoning models will progress rapidly into agentic behavior
“I think that we're going to continue to see, as I said before, that we're going to see this paradigm continue to progress rapidly. And I think that that's true even today, that we saw that with like going from O-one preview to O-one to O-three, consistent prog…”
Noam Brown Jun 19, 2025 ▶ 6:07
Insight
Brown: Deep Research proves reasoning models work in unverifiable domains
“And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very like, what is the best research report that you could generate? And yet these models are doing extremely well at this domain. So I think that's like an ex…”
Noam Brown Jun 19, 2025 ▶ 7:32
Insight
Noam Brown: Models need baseline capabilities to benefit from test-time reasoning
“One thing that I think is underappreciated is that the models, the pre-trained models need a certain level of capability in order to really benefit from this, like, extra thinking.”
Noam Brown Jun 19, 2025 ▶ 9:22
What-if
Noam Brown: Reasoning paradigms would have failed on GPT-2
“If you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.”
Noam Brown Jun 19, 2025 ▶ 9:35
Assertion Supported
Noam Brown: GPT-4.5 makes Tic-Tac-Toe mistakes without System 2 reasoning
“With Tic-Tac-Toe, we see that, like, GPD-Four .5 falls over. You know, it plays decently well. I shouldn't say it falls over. It does reasonably well. You can draw the board. It can make legal moves, but it will make mistakes sometimes, and if you really need …”
Noam Brown Jun 19, 2025 ▶ 12:06
Insight
Noam Brown: The Ideal AI Agent Harness Is No Harness
“The ideal harness is no harness. Right. I think harnesses are like a crutch that eventually we're going to be able to move beyond.”
Noam Brown Jun 19, 2025 ▶ 14:03
Assertion Not checkable as stated
Brown: OpenAI's o3 Gets 'Not Very Far' Playing Pokémon Unharnessed
“How far does O three get without any harness? How far does it get playing Pokemon? And the answer is like, not very far, you know?”
Noam Brown Jun 19, 2025 ▶ 14:33
Prediction Not checkable as stated
Brown: Model routers will become obsolete as unified models emerge
“We've said pretty openly that we want to move to a world where there is a single unified model. And in that world, you shouldn't need a router on top of the model. So I think that the router issue Will eventually be solved also.”
Noam Brown Jun 19, 2025 ▶ 19:00
Insight
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26
Prediction Not checkable as stated
Brown: Pre-training scaling will hit economic limits before superintelligence without reasoning
“Like, we're gonna scale it, sure, we're gonna scale these things up by a few more orders of magnitude, they're gonna become more capable, but we're not gonna see superintelligence from just that. And like, yes, if we had a quadrillion dollars to train these mo…”
Noam Brown Jun 19, 2025 ▶ 23:20
Assertion Not checkable as stated
Brown: OpenAI saw conclusive proof of its reasoning paradigm in late 2023
“I think it was around, like, November, twenty-twenty-three, or October, twenty-twenty-three, when I think I was convinced that we had, like, very conclusive signs of life, that, like, oh, this was going to be, this is the paradigm, and it's going to be a big d…”
Noam Brown Jun 19, 2025 ▶ 25:17
Opinion
Noam Brown: Closing the human data efficiency gap is a top unsolved problem
“I think it's a fair statement to say that these models are less data efficient than humans. And I think that that's an unsolved research question and probably one of the most important unsolved research questions.”
Noam Brown Jun 19, 2025 ▶ 30:21
Assertion Not checkable as stated
Noam Brown: OpenAI succeeded early by betting on scaling over small experiments
“One of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab, you know, it was, like, difficult for them to scale. Back then, it was much more common to do, like, a lot of small experi…”
Noam Brown Jun 19, 2025 ▶ 32:07
Disclosure
Noam Brown: OpenAI o3 has basically replaced Google Search for me
“Like I've been using it day to day. It's basically replaced Google search for me. Like I just use it all the time.”
Noam Brown Jun 19, 2025 ▶ 35:46
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Noam Brown Jun 19, 2025 ▶ 38:35
Insight
Noam Brown: Aligned AI will outperform human virtual assistants on effort
“And so if you have an AI model that's, like, actually really aligned, To you and your preferences, then that could end up doing a way better job than a human could. Well, not, not that it's doing a better job than a human could, but like it's doing a better jo…”
Noam Brown Jun 19, 2025 ▶ 40:44
Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Noam Brown Jun 19, 2025 ▶ 41:59
Prediction Not checkable as stated
Brown: Multi-agent AI civilizations will far surpass current AI capabilities
“And I think that if you're able to have them cooperate and compete with billions of AIs over a long period of time and build up a civilization essentially, the things that they would be able to Produce and answer would be far beyond what is possible today with…”
Noam Brown Jun 19, 2025 ▶ 43:36
Opinion
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Noam Brown Jun 19, 2025 ▶ 44:58
Insight
Brown: Game Theory Optimal fails in collaborative games like Diplomacy
“Basically, when you're playing, like, the zero-sum games, like, poker, Game Theory Optimal works really well. When you're playing a game like Diplomacy, where there's, like, you need to collaborate and compete, and you need, there's room for collaboration, the…”
Noam Brown Jun 19, 2025 ▶ 49:43
Assertion Supported
Brown: Modern poker AIs stick to static GTO without player adaptation
“The way the Poker AI's work today, they're just kind of like sticking to their precomputed GTO strategy. And they're not adapting to the other players at the table.”
Noam Brown Jun 19, 2025 ▶ 51:37
Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31
Opinion
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Noam Brown Jun 19, 2025 ▶ 58:29
Insight
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Noam Brown Jun 19, 2025 ▶ 1:08:26
Disclosure
Brown: OpenAI models undergo mid-training and post-training before release
“For open AI models, like, they go through a mid-training step, and then they go through a post-training step, and then they're released, and they're a lot more useful. Like, frankly, if you interacted with the only pre-trained model, it would be super difficul…”
Noam Brown Jun 19, 2025 ▶ 1:11:38
Prediction Not checkable as stated
Noam Brown: A Superhuman Magic: The Gathering AI Is Feasible Today
“And my guess is that if somebody put in the effort, they could probably make a superhuman bot for Magic the Gathering now.”
Noam Brown Jun 19, 2025 ▶ 1:16:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.