test-time compute

also referred to as: test time compute

6 statements across 2 episodes · 1 bullish · 3 bearish · 2 people on the record · first statement Oct 17, 2024 by Sarah Guo · across every show →

Everything said about test-time compute, oldest first

Oct 17, 2024 bullish
Prediction Not checkable as stated
Sarah Guo says test-time compute scaling unlocks new AI competition
“Another school of thought is which I do subscribe to, by the way, is you know, new scaling law, right? So will allow us to do an important range of new tasks, and how good it is exactly at this moment is not the important thing. It's a new dimension of competi…”
Sarah Guo Oct 17, 2024 ▶ 16:02 No Priors Ep. 86 | With Sarah Guo & Elad Gil
Jun 26, 2026 negative
Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026 bearish
Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Noam Brown Jun 26, 2026 ▶ 26:21 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026
Insight
Brown: Test-time AI performance scales along a continuous, projectable slope
“You also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could probably do some kind of evaluation up to a certain budget an…”
Noam Brown Jun 26, 2026 ▶ 4:56 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026 neutral
Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Jun 26, 2026 negative
Insight
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Noam Brown Jun 26, 2026 ▶ 21:43 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.