Insight certainty 4/5 debate potential 2/5

Brown: Extra test-time compute does not improve factual retrieval in AI

Noam Brown · Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown · Jun 26, 2026 · at 21:43

OpenAI researcher Noam Brown explains the fundamental scaling limits of test-time compute across different problem types.

0:00 / 0:26exact quote · 26.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't know the date, they could sit there, they could think about it for a week. If they don't have access to Wikipedia or something, they're not gonna be able to do better answering that question if they thought about it for a week compared to five seconds. Same with the model.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Noam Brown

Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Noam Brown Jun 26, 2026 ▶ 26:21 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Assertion Open · timeframe Jun 2027
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Noam Brown Jun 26, 2026 ▶ 3:34 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Opinion
Noam Brown: AI Model Outputs Are Arguably More Trustworthy Than Humans
“I use it day to day for a lot of this kind of stuff, and I think they're at a point now where They've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from a human.”
Noam Brown Jun 26, 2026 ▶ 31:36 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Insight
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Noam Brown Apr 25, 2023 ▶ 17:12 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta
Insight
Brown: AI benchmarks must control for test-time compute
“And so I think the proper way to, and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of…”
Noam Brown Jun 26, 2026 ▶ 4:01 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.