Test Time Compute
topic on 7 shows · 20 statements across 13 episodes
the Y Combinator Startup Podcast
Latent Space
No Priors
Catalyst
the MAD Podcast
the a16z Podcast
TBPN
20 statements about Test Time Compute, every show
Brown: Test-time AI performance scales along a continuous, projectable slope
“You also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could probably do some kind of evaluation up to a certain budget an…”
Brown: Long AI Deliberation Time Is Impractical for Real Workflows
“This idea that the models, you just let them think for a week or whatever, and then they respond, it's, it sounds nice, and yes, the benchmarks look great, but it's not very practical when working because like, okay, you ask the model a question, and then you …”
Brown: Scaffolding Easily Inflates AI Benchmark Scores Without Real Gains
“It's really easy to show you can do much better than previous benchmarks or previous, previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay, well, we're going to, instead of just running this model once…”
Brown: Extra test-time compute does not improve factual retrieval in AI
“There are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't kno…”
Overnight AI intelligence explosion unlikely due to test-time compute bottlenecks
“And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test time compute. In order to achieve their greatest intelligence. If you, if it requires so much test time compute to unlock the full capabil…”
Roberts: Powerful pre-trained models are necessary for effective RL and reasoning
“If you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to like think at use test time compute to for instance, solve, solve math problems that it wouldn't otherwise be able to do.”
Mukund Jha: Emergent invented multi-agent paradigms before academic papers were published
“There was a time when we sort of invented the multi-agent system. We invented memory. We invented like, how do we do agent to agent communication? How do you scale up test time compute? A lot of those things which like were sort of coming out, like we would di…”
Chubuk: OpenAI o1 showed test-time compute improves results beyond training sets
“So what O-one showed is if you spend test time compute, you can get better results. So that was very exciting to me because there was one way of investing resources that was beyond the training set.”
Wang: AI image models will adopt multi-hour test-time compute when needed
“Yeah, absolutely. If it's necessary, yeah.”
Douglas: Test-time compute executes reasoning while RL provides feedback on correctness
“One way of thinking about this is test time compute is doing a lot of reasoning, and then RL is the feedback signal on whether or not that reasoning was right or wrong.”
Douglas: Test-time compute solves harder tasks before RL distills them into models
“Test time compute lets you do harder problems than you can currently do, than you can like currently do off the cuff, and RL
Then allows you to sort of distill that back into the model.”
Morcos: Test-time compute fundamentally favors smaller models to cut multi-step inference costs
“Test time compute as a paradigm really pushes you towards smaller models, right? Because if your cost of solving a problem is cost of inference times number of thinking steps, and you have to do a lot of thinking steps. Well, now this is like a really like min…”
Brown: OpenAI team is scaling test-time compute to hours and days
“The team, in many ways, is actually a misnomer, because we're working on more than just multi-agent. Multi-agent is one of the things we're working on. Some other things we're working on is just like being able to scale up test time compute by a ton. So how, y…”
Brown: Reasoning models improve primarily through compute efficiency rather than longer thinking duration
“These models are becoming more efficient in the way they're thinking, as they're able to do more with the same amount of test time compute, and I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer…”
Howard: Test-time compute scaling will hit diminishing returns within two years
“It's a thing where you get most of the juice out of it in the first year or two, so we're still in that. Phase at the moment, and we'll start to hit the curve off point pretty soon. Just like we did for training.”
Packer: Sleep-time compute during idle downtime is a major missed opportunity
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at t…”
Packer: Nobody in the AI industry is actively scaling sleep-time compute
“How much we really can kind of take advantage of sleep time compute, which is effectively completely on mine today. Like nobody's really Scaling in the sleep time compute direction.”
Knoop: Scaling Test-Time Compute Will Not Get Us to AGI
“And then there's a new story that's emerged over the last like five months, which is, oh, we're going to scale up this test time compute and that's going to get us to AGI. And I think what V two shows is that that's not quite either. We still need some structu…”
Hays: Users Will Stop Caring About Exposed AI Reasoning Once Trust Builds
“I think there's this period of time where it's true. People want to see how it's working through something, but then eventually when you have that level of trust with the model or the app that you're using, you just want it done.”
Sarah Guo says test-time compute scaling unlocks new AI competition
“Another school of thought is which I do subscribe to, by the way, is you know, new scaling law, right? So will allow us to do an important range of new tasks, and how good it is exactly at this moment is not the important thing. It's a new dimension of competi…”