Test Time Compute

topic on 7 shows · 20 statements across 13 episodes

the Y Combinator Startup Podcast Latent Space No Priors Catalyst the MAD Podcast the a16z Podcast TBPN

20 statements about Test Time Compute, every show

NO PRIORS Insight
Brown: Test-time AI performance scales along a continuous, projectable slope
“You also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could probably do some kind of evaluation up to a certain budget an…”
Noam Brown Jun 26, 2026 ▶ 4:56 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Noam Brown Jun 26, 2026 ▶ 6:13 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Noam Brown Jun 26, 2026 ▶ 7:03 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Insight
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Noam Brown Jun 26, 2026 ▶ 21:43 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
NO PRIORS Prediction Not checkable as stated
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Noam Brown Jun 26, 2026 ▶ 26:21 Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
MAD Insight
Roberts: Powerful pre-trained models are necessary for effective RL and reasoning
“If you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to like think at use test time compute to for instance, solve, solve math problems that it wouldn't otherwise be able to do.”
Dan Roberts Jun 4, 2026 ▶ 27:15 OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Y COMBINATOR Assertion Not checkable as stated
Mukund Jha: Emergent invented multi-agent paradigms before academic papers were published
“There was a time when we sort of invented the multi-agent system. We invented memory. We invented like, how do we do agent to agent communication? How do you scale up test time compute? A lot of those things which like were sort of coming out, like we would di…”
Mukund Jha Mar 16, 2026 ▶ 3:44 AI Is Unlocking Millions Of New Builders · Y Combinator
CATALYST Assertion Supported
Chubuk: OpenAI o1 showed test-time compute improves results beyond training sets
“So what O-one showed is if you spend test time compute, you can get better results. So that was very exciting to me because there was one way of investing resources that was beyond the training set.”
Doge Chubuk Nov 6, 2025 ▶ 6:30 Inside a $300 million bet on AI for physical R&D
a16z Prediction Open · timeframe Oct 2030
Wang: AI image models will adopt multi-hour test-time compute when needed
“Yeah, absolutely. If it's necessary, yeah.”
Oliver Wang Oct 28, 2025 ▶ 19:07 Google DeepMind Developers: How Nano Banana Was Made
MAD Insight
Douglas: Test-time compute executes reasoning while RL provides feedback on correctness
“One way of thinking about this is test time compute is doing a lot of reasoning, and then RL is the feedback signal on whether or not that reasoning was right or wrong.”
Sholto Douglas Oct 2, 2025 ▶ 50:58 Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
MAD Insight
Douglas: Test-time compute solves harder tasks before RL distills them into models
“Test time compute lets you do harder problems than you can currently do, than you can like currently do off the cuff, and RL Then allows you to sort of distill that back into the model.”
Sholto Douglas Oct 2, 2025 ▶ 51:52 Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
Morcos: Test-time compute fundamentally favors smaller models to cut multi-step inference costs
“Test time compute as a paradigm really pushes you towards smaller models, right? Because if your cost of solving a problem is cost of inference times number of thinking steps, and you have to do a lot of thinking steps. Well, now this is like a really like min…”
Ari Morcos Aug 29, 2025 ▶ 1:04:35 Better Data is All You Need — Ari Morcos, Datology
LATENT SPACE Disclosure
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Noam Brown Jun 19, 2025 ▶ 41:59 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Noam Brown Jun 19, 2025 ▶ 1:08:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
MAD Prediction Not checkable as stated
Howard: Test-time compute scaling will hit diminishing returns within two years
“It's a thing where you get most of the juice out of it in the first year or two, so we're still in that. Phase at the moment, and we'll start to hit the curve off point pretty soon. Just like we did for training.”
Jeremy Howard May 15, 2025 ▶ 9:56 Jeremy Howard on Building 5,000 AI Products with 14 People (Answer AI Deep-Dive)
Packer: Sleep-time compute during idle downtime is a major missed opportunity
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at t…”
Charles Packer Apr 21, 2025 ▶ 1:02 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
LATENT SPACE Assertion Not checkable as stated
Packer: Nobody in the AI industry is actively scaling sleep-time compute
“How much we really can kind of take advantage of sleep time compute, which is effectively completely on mine today. Like nobody's really Scaling in the sleep time compute direction.”
Charles Packer Apr 21, 2025 ▶ 12:38 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
TBPN Prediction Not checkable as stated
Knoop: Scaling Test-Time Compute Will Not Get Us to AGI
“And then there's a new story that's emerged over the last like five months, which is, oh, we're going to scale up this test time compute and that's going to get us to AGI. And I think what V two shows is that that's not quite either. We still need some structu…”
Mike Knoop Apr 6, 2025 ▶ 6:40 Mike Knoop (Arc Prize) on Why Scaling AI Won’t Get Us to AGI
TBPN Insight
Hays: Users Will Stop Caring About Exposed AI Reasoning Once Trust Builds
“I think there's this period of time where it's true. People want to see how it's working through something, but then eventually when you have that level of trust with the model or the app that you're using, you just want it done.”
Jordi Hays Jan 28, 2025 ▶ 1:35:58 DeepSeek Update, Market Crash, Timeline in Turmoil, Is VC Cooked, Zero Cope Policy
NO PRIORS Prediction Not checkable as stated
Sarah Guo says test-time compute scaling unlocks new AI competition
“Another school of thought is which I do subscribe to, by the way, is you know, new scaling law, right? So will allow us to do an important range of new tasks, and how good it is exactly at this moment is not the important thing. It's a new dimension of competi…”
Sarah Guo Oct 17, 2024 ▶ 16:02 No Priors Ep. 86 | With Sarah Guo & Elad Gil

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.