Jun 26, 2026 · 36m · no-priors
Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI researcher Noam Brown joins host Sarah Guo on No Priors to discuss how large-scale test-time compute is reshaping artificial intelligence evaluation, frontier reasoning capabilities, and safety preparedness frameworks. The conversation explores the limitations of traditional benchmark grids, the dynamics of long-horizon autonomous scaffolding, and the future of multi-agent intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Noam politely rejects Sarah's premise that multi-agent research is underexplored, asserting it is widely explored but limited by current model architectures.
Hardest push from the hosts ▶ 25:52 Challenging fast takeoff assumptionsSarah directly confronts Noam with the sharp implication of his bottleneck argument, pushing him to state on record whether he rejects the fast takeoff thesis.
Biggest teaching moment ▶ 2:40 Deconstructing standard benchmark grid illusionsNoam educates the audience and host on why comparing single-number benchmark grids is fundamentally broken when newer models are more compute-efficient and do not quickly plateau.
The host holds their own ▶ 35:04 Synthesizing the routing vs scalar compute principleSarah demonstrates sharp domain grasp by instantly synthesizing Noam's thesis to evaluate routing architectures under the identical test-time compute cost scalar.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Evaluating AI Capabilities by Test-Time Compute Budget | 6 | 5 | 2 | 1 | Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading. | |
| Forecasting Long-Horizon Capability Curves and Inference Budgets | 6 | 4 | 2 | 2 | Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration. | |
| Benchmark Maxing, Scaffolding, and Private Evaluation Sets | 5 | 4 | 1 | 1 | Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations. | |
| The Safety Challenge of Budget-Scaled Hazardous Capabilities | 6 | 5 | 1 | 2 | Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets. | |
| Long-Horizon Scaffolding and Fast Model Release Cycles | 5 | 4 | 1 | 1 | The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles. | |
| Unlocking Latent Capabilities and Mathematical Discoveries | 5 | 6 | 2 | 1 | Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models. | |
| Frontier Research Priorities and Compute Scaling Limits | 7 | 5 | 3 | 4 | Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste). | |
| Recursive Self-Improvement and the Time Bottleneck Takeoff | 7 | 5 | 3 | 4 | Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck. | |
| The Multi-Agent Frontier and Compounding AI Knowledge | 5 | 5 | 4 | 2 | When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia. | |
| Frontier Competition Dynamics and Real-World AI Trust | 6 | 4 | 1 | 1 | Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents. | |
| Overcoming the Benchmark Grid Bad Equilibrium | 7 | 4 | 2 | 2 | Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets. |