Inference Time Compute
topic on 4 shows · 10 statements across 9 episodes
the Y Combinator Startup Podcast
Latent Space
No Priors
TBPN
10 statements about Inference Time Compute, every show
Jeff Dean: Inference-time compute search improves reliability in agent workflows
“Inference time compute to perform search over plausible ways of solving the problem that can get much, much higher performance or much more reliability in, Long-running agent flows.”
Nadella: Inference-time compute adds another massive AI scaling law
“Then this inference time compute seems to have really added in another massive scaling law.”
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Kilpatrick: Reasoning and test-time compute will scale faster medium-term
“And it feels like that is. More likely in the medium term going to be the thing that like continues to just like rapidly scale up relative to the other capabilities.”
Coogan: Inference-Time Compute Establishes an Independent Scaling Law from Pre-Training
“This development creates a new scaling law that is totally independent of the original pre-training scaling law. Now you still want to train the best model you can by clearly leveraging as much compute as you can as many trillion tokens of high quality trainin…”
Inference-Time Compute Lets Labs Scale Intelligence Without Doubling Supercomputers
“I don't need to go double the size of my supercomputer to hit a requisite intelligence threshold. I can just double the amount of inference time compute that my customers pay for.”
Inference-Time Compute Does Not Require Densely Interconnected Supercomputers
“If we have a new avenue, which is inference time compute, That doesn't require this densely interconnected supercomputer. It's fine to have nodes. You can do a lot more locally and less distributed.”
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Steinberger: AI performance requires trading off training compute against inference compute
“Well, so you can think of model performance as some function of training compute times some function of inference time compute. Now those are specific functions that are just scaling law things that you can like model, but the general Way to think about it is …”