Jun 26, 2026 · 36m · no-priors

Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown

Noam Brown · 25m spoken Sarah Guo · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researcher Noam Brown joins host Sarah Guo on No Priors to discuss how large-scale test-time compute is reshaping artificial intelligence evaluation, frontier reasoning capabilities, and safety preparedness frameworks. The conversation explores the limitations of traditional benchmark grids, the dynamics of long-horizon autonomous scaffolding, and the future of multi-agent intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →

The hosts as informed peer 5.9 Guest teaching 4.6 Guest disagreement 2.0 The hosts pushing back 1.9
05100:0010:0020:0030:000:00–4:20 · The hosts as informed peer 6/10 Evaluating AI Capabilities by Test-Time Compute Budget Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading.4:20–6:46 · The hosts as informed peer 6/10 Forecasting Long-Horizon Capability Curves and Inference Budgets Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration.6:47–11:26 · The hosts as informed peer 5/10 Benchmark Maxing, Scaffolding, and Private Evaluation Sets Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations.11:28–14:43 · The hosts as informed peer 6/10 The Safety Challenge of Budget-Scaled Hazardous Capabilities Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets.14:44–17:11 · The hosts as informed peer 5/10 Long-Horizon Scaffolding and Fast Model Release Cycles The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles.17:11–19:54 · The hosts as informed peer 5/10 Unlocking Latent Capabilities and Mathematical Discoveries Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models.19:55–23:52 · The hosts as informed peer 7/10 Frontier Research Priorities and Compute Scaling Limits Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste).23:59–27:10 · The hosts as informed peer 7/10 Recursive Self-Improvement and the Time Bottleneck Takeoff Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck.27:10–29:11 · The hosts as informed peer 5/10 The Multi-Agent Frontier and Compounding AI Knowledge When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia.29:11–31:51 · The hosts as informed peer 6/10 Frontier Competition Dynamics and Real-World AI Trust Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents.31:51–36:16 · The hosts as informed peer 7/10 Overcoming the Benchmark Grid Bad Equilibrium Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets.0:00–4:20 · Guest teaching 5/10 Evaluating AI Capabilities by Test-Time Compute Budget Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading.4:20–6:46 · Guest teaching 4/10 Forecasting Long-Horizon Capability Curves and Inference Budgets Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration.6:47–11:26 · Guest teaching 4/10 Benchmark Maxing, Scaffolding, and Private Evaluation Sets Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations.11:28–14:43 · Guest teaching 5/10 The Safety Challenge of Budget-Scaled Hazardous Capabilities Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets.14:44–17:11 · Guest teaching 4/10 Long-Horizon Scaffolding and Fast Model Release Cycles The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles.17:11–19:54 · Guest teaching 6/10 Unlocking Latent Capabilities and Mathematical Discoveries Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models.19:55–23:52 · Guest teaching 5/10 Frontier Research Priorities and Compute Scaling Limits Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste).23:59–27:10 · Guest teaching 5/10 Recursive Self-Improvement and the Time Bottleneck Takeoff Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck.27:10–29:11 · Guest teaching 5/10 The Multi-Agent Frontier and Compounding AI Knowledge When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia.29:11–31:51 · Guest teaching 4/10 Frontier Competition Dynamics and Real-World AI Trust Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents.31:51–36:16 · Guest teaching 4/10 Overcoming the Benchmark Grid Bad Equilibrium Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets.0:00–4:20 · Guest disagreement 2/10 Evaluating AI Capabilities by Test-Time Compute Budget Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading.4:20–6:46 · Guest disagreement 2/10 Forecasting Long-Horizon Capability Curves and Inference Budgets Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration.6:47–11:26 · Guest disagreement 1/10 Benchmark Maxing, Scaffolding, and Private Evaluation Sets Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations.11:28–14:43 · Guest disagreement 1/10 The Safety Challenge of Budget-Scaled Hazardous Capabilities Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets.14:44–17:11 · Guest disagreement 1/10 Long-Horizon Scaffolding and Fast Model Release Cycles The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles.17:11–19:54 · Guest disagreement 2/10 Unlocking Latent Capabilities and Mathematical Discoveries Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models.19:55–23:52 · Guest disagreement 3/10 Frontier Research Priorities and Compute Scaling Limits Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste).23:59–27:10 · Guest disagreement 3/10 Recursive Self-Improvement and the Time Bottleneck Takeoff Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck.27:10–29:11 · Guest disagreement 4/10 The Multi-Agent Frontier and Compounding AI Knowledge When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia.29:11–31:51 · Guest disagreement 1/10 Frontier Competition Dynamics and Real-World AI Trust Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents.31:51–36:16 · Guest disagreement 2/10 Overcoming the Benchmark Grid Bad Equilibrium Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets.0:00–4:20 · The hosts pushing back 1/10 Evaluating AI Capabilities by Test-Time Compute Budget Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading.4:20–6:46 · The hosts pushing back 2/10 Forecasting Long-Horizon Capability Curves and Inference Budgets Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration.6:47–11:26 · The hosts pushing back 1/10 Benchmark Maxing, Scaffolding, and Private Evaluation Sets Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations.11:28–14:43 · The hosts pushing back 2/10 The Safety Challenge of Budget-Scaled Hazardous Capabilities Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets.14:44–17:11 · The hosts pushing back 1/10 Long-Horizon Scaffolding and Fast Model Release Cycles The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles.17:11–19:54 · The hosts pushing back 1/10 Unlocking Latent Capabilities and Mathematical Discoveries Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models.19:55–23:52 · The hosts pushing back 4/10 Frontier Research Priorities and Compute Scaling Limits Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste).23:59–27:10 · The hosts pushing back 4/10 Recursive Self-Improvement and the Time Bottleneck Takeoff Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck.27:10–29:11 · The hosts pushing back 2/10 The Multi-Agent Frontier and Compounding AI Knowledge When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia.29:11–31:51 · The hosts pushing back 1/10 Frontier Competition Dynamics and Real-World AI Trust Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents.31:51–36:16 · The hosts pushing back 2/10 Overcoming the Benchmark Grid Bad Equilibrium Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 28.8% · guest 71.2%0:00 · the hosts 28.8% · guest 71.2%3:00 · the hosts 28.7% · guest 71.3%3:00 · the hosts 28.7% · guest 71.3%6:00 · the hosts 24.1% · guest 75.9%6:00 · the hosts 24.1% · guest 75.9%9:00 · the hosts 19.3% · guest 80.7%9:00 · the hosts 19.3% · guest 80.7%12:00 · the hosts 29.5% · guest 70.5%12:00 · the hosts 29.5% · guest 70.5%15:00 · the hosts 7.9% · guest 92.1%15:00 · the hosts 7.9% · guest 92.1%18:00 · the hosts 8.2% · guest 91.8%18:00 · the hosts 8.2% · guest 91.8%21:00 · the hosts 18.5% · guest 81.5%21:00 · the hosts 18.5% · guest 81.5%24:00 · the hosts 7.2% · guest 92.8%24:00 · the hosts 7.2% · guest 92.8%27:00 · the hosts 28.8% · guest 71.2%27:00 · the hosts 28.8% · guest 71.2%30:00 · the hosts 26.2% · guest 73.8%30:00 · the hosts 26.2% · guest 73.8%33:00 · the hosts 44.3% · guest 55.7%33:00 · the hosts 44.3% · guest 55.7%36:00 · the hosts 100% · guest 0%36:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 27:17 Reframing multi-agent exploration

Noam politely rejects Sarah's premise that multi-agent research is underexplored, asserting it is widely explored but limited by current model architectures.

Hardest push from the hosts ▶ 25:52 Challenging fast takeoff assumptions

Sarah directly confronts Noam with the sharp implication of his bottleneck argument, pushing him to state on record whether he rejects the fast takeoff thesis.

Biggest teaching moment ▶ 2:40 Deconstructing standard benchmark grid illusions

Noam educates the audience and host on why comparing single-number benchmark grids is fundamentally broken when newer models are more compute-efficient and do not quickly plateau.

The host holds their own ▶ 35:04 Synthesizing the routing vs scalar compute principle

Sarah demonstrates sharp domain grasp by instantly synthesizing Noam's thesis to evaluate routing architectures under the identical test-time compute cost scalar.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Evaluating AI Capabilities by Test-Time Compute Budget 6521 Sarah sets up the premise of Noam's influential essay on test-time compute and accurately summarizes the need for token budget limits. Noam explains the core technical dilemma where modern models do not plateau quickly, rendering standard benchmark grids misleading.
Forecasting Long-Horizon Capability Curves and Inference Budgets 6422 Sarah probes on whether users systematically under-allocate test-time compute or if models simply need to think faster. Noam clarifies the practical trade-off between long-horizon batch thinking and fast interactive iteration.
Benchmark Maxing, Scaffolding, and Private Evaluation Sets 5411 Sarah asks about practical evaluation methodologies beyond public benchmarks. Noam shares his personal methodology of testing models on poker bot solver construction across model generations.
The Safety Challenge of Budget-Scaled Hazardous Capabilities 6512 Sarah highlights the dangerous capability implications of non-asymptoting compute budgets and how safety evals lag behind rapid release cycles. Noam details how current responsible scaling policies neglect test-time compute budgets.
Long-Horizon Scaffolding and Fast Model Release Cycles 5411 The host asks whether Noam has tested running models indefinitely on tasks like poker solvers. Noam explains the tension between needing months to push a model to its capability ceiling and the 2-3 month frontier release cycles.
Unlocking Latent Capabilities and Mathematical Discoveries 5621 Sarah asks about latent capabilities in existing models, leading Noam to reveal how OpenAI disproved the Erdos unit distance conjecture using internal models and how general scaffolding could have extracted it from public models.
Frontier Research Priorities and Compute Scaling Limits 7534 Sarah provocatively asks if OpenAI researchers are just waiting for the next model release and pushes on limits of scaling. Noam articulates where compute scaling helps (reasoning/Sudoku) versus where it fundamentally fails (factual retrieval, research taste).
Recursive Self-Improvement and the Time Bottleneck Takeoff 7534 Sarah directly challenges Noam on whether his view implies we are far from an immediate fast takeoff. Noam presents his core thesis: because maximum capability requires large test-time compute, physical wall-clock time remains an unavoidable bottleneck.
The Multi-Agent Frontier and Compounding AI Knowledge 5542 When Sarah suggests multi-agent systems are an underexplored frontier, Noam pushes back that it is heavily explored but currently constrained by context horizons, comparing it to human cultural knowledge accumulation over millenia.
Frontier Competition Dynamics and Real-World AI Trust 6411 Sarah frames the competitive dynamics between frontier labs as a grounded grind rather than an runaway takeoff. Noam agrees and shares how he personally trusts frontier models for high-stakes decisions like legal and real estate documents.
Overcoming the Benchmark Grid Bad Equilibrium 7422 Sarah asks about specialized routing layers versus native model reasoning. Noam notes routing gains must still be evaluated against the baseline of giving the frontier model equivalent test-time compute budgets.

Statements from this episode (25)

Assertion Supported
Brown: GPT-5.5 is far more compute-efficient than GPT-5.4
“It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5. And once you control for the amount of thinking time, actually you can see t…”
Noam Brown Jun 26, 2026 ▶ 2:43
Assertion Open · timeframe Jun 2027
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Noam Brown Jun 26, 2026 ▶ 3:34
Insight
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Noam Brown Jun 26, 2026 ▶ 4:01
Assertion Supported
Brown: AISI evals show AI cyber capabilities improve past 100M tokens
“Actually the AISI in their evaluations has shown that the models continue to improve at A hundred million tokens. You know, if you run them for a hundred million tokens, they're still improving at beyond that point.”
Noam Brown Jun 26, 2026 ▶ 4:41
Insight
Brown: Test-time AI performance scales along a continuous, projectable slope
“You also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could probably do some kind of evaluation up to a certain budget an…”
Noam Brown Jun 26, 2026 ▶ 4:56
Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13
Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03
Insight
Brown: Poker bot creation is a superior AI reasoning evaluation
“I think it's a nice eval because there is very little open source code for making poker bots. And there's a lot of published essays, there's a lot of published papers on it, but you really have to reason through everything.”
Noam Brown Jun 26, 2026 ▶ 8:41
Prediction Not checkable as stated
Brown predicts AI will zero-shot his entire PhD thesis within one year
“And I wouldn't be surprised if, you know, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis in one go.”
Noam Brown Jun 26, 2026 ▶ 11:17
Assertion Supported
Brown: GPT-3 Capabilities Could Not Scale With Test-Time Compute Budget
“Like, with GPT-III, you couldn't scale test time compute. Like, if you gave it a budget of ten million dollars and said, okay, well, let's see what GPT-III can do, it really can't do that much, more than what you could do with, like, 10 dollars or one dollar.”
Noam Brown Jun 26, 2026 ▶ 12:42
Insight
Current AI safety frameworks fail to account for test-time compute scaling
“The preparedness frameworks and responsible scaling policies, they don't really account for the amount of tests I'm computed. They just say, okay, well, what's the capability of the model? The problem is we're in a world now where the capability of the model i…”
Noam Brown Jun 26, 2026 ▶ 12:55
Assertion Supported
Brown: Modern AI Models Can Run Scaffolded Experiments for Months
“We're seeing now with the most recent models that you can actually scaffold, for example, 5.5 into doing a series of experiments that can run for weeks, for months.”
Noam Brown Jun 26, 2026 ▶ 15:03
Insight
Brown: Rapid AI Release Cycles Obscure True Model Capability Ceilings
“The model release cycle is, look, we're releasing new models, like, every two or three months at this point, and so a model comes out, it takes two or three months to push it to its limits, and then you have another model come out, and so nobody actually knows…”
Noam Brown Jun 26, 2026 ▶ 16:10
Assertion Supported
OpenAI internal model reportedly disproved the Erdős unit distance conjecture
“We used an internal model at OpenAI a few weeks ago to disprove the unit Erdos unit distance conjecture.”
Noam Brown Jun 26, 2026 ▶ 17:20
Assertion Supported
Brown: GPT-5.5 can derive Erdős disproof with proper scaffolding
“After we announced the results, A bunch of people found that you could get the answer out of 5.5 as well. If, now, it's not as simple as just asking 5.5, hey, here's the Irish unit distance conjecture. What's the disproof? You had to scaffold it a bit. You had…”
Noam Brown Jun 26, 2026 ▶ 18:01
Insight
Complex task compute costs fall 10x to 100x per model release
“The model release cycle is every, every couple months we put out a new model that's even more powerful, and so the cost of disproving the Erdos unit distance gesture drops by, like, 10 or a hundred x with every model release cycle. Probably, in some cases, mor…”
Noam Brown Jun 26, 2026 ▶ 19:28
Disclosure
OpenAI discourages researchers from using current models on open math problems
“We are trying to encourage people to not spend all their time just, like, Going through all the mathematical open problems, physics problems, and just seeing, pushing the models to their limits to see what they can prove or disprove. Because we really think th…”
Noam Brown Jun 26, 2026 ▶ 20:14
Insight
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Noam Brown Jun 26, 2026 ▶ 21:43
Assertion Not checkable as stated
Brown says AI models optimized his PhD poker algorithms by 1,000x
“I was really impressed with the model's ability to optimize the algorithms that I had developed in my PhD. It was honestly, it was shocking to see how inefficient I was in retrospect, and they were able to make it like, you know, 1000 x faster.”
Noam Brown Jun 26, 2026 ▶ 24:00
Assertion Not checkable as stated
Brown: AI cannot invent novel algorithms better than existing research
“Go ahead and like look at all the published work and synthesize that and then try to come up with something novel and it's not able to do it. And I can give it a lot of time and it's still not able to do it.”
Noam Brown Jun 26, 2026 ▶ 24:30
Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Noam Brown Jun 26, 2026 ▶ 26:21
Insight
Noam Brown: AI Models Cannot Organically Accumulate Shared Knowledge Today
“We're not seeing that with AI models today. They kind of, they're born into a world for, and they exist for a very short context window, and then they just, like, disappear. And yeah, there are things that you can kind of do to, like, continue them, but it's v…”
Noam Brown Jun 26, 2026 ▶ 28:26
Opinion
Noam Brown: AI Model Outputs Are Arguably More Trustworthy Than Humans
“I use it day to day for a lot of this kind of stuff, and I think they're at a point now where They've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from a human.”
Noam Brown Jun 26, 2026 ▶ 31:36
Insight
Noam Brown: AI community stuck in bad equilibrium publishing static benchmark grids
“I would talk to researchers about we, it makes sense to show the benchmarks with an x-axis, whether it's tokens or cost or time, there should be an x-axis, and everybody would say, like, yeah, that makes sense, we should do that, but. Well, really, their respo…”
Noam Brown Jun 26, 2026 ▶ 32:27
Insight
Brown: Benchmark Gains From Routing May Fail in Real-World Use
“One issue you could run into is that you could optimize for certain benchmarks with the routing and then show like, oh yeah, we see this big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement.”
Noam Brown Jun 26, 2026 ▶ 35:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.