Inference Time Compute

topic on 4 shows · 10 statements across 9 episodes

the Y Combinator Startup Podcast Latent Space No Priors TBPN

10 statements about Inference Time Compute, every show

Jeff Dean: Inference-time compute search improves reliability in agent workflows
“Inference time compute to perform search over plausible ways of solving the problem that can get much, much higher performance or much more reliability in, Long-running agent flows.”
Jeff Dean Jul 30, 2026 ▶ 24:13 Jeff Dean: The 1% Rule for Building in AI · Y Combinator
Nadella: Inference-time compute adds another massive AI scaling law
“Then this inference time compute seems to have really added in another massive scaling law.”
Satya Nadella Jun 25, 2025 ▶ 16:10 Satya Nadella: Microsoft's AI Bets, Hyperscaling, Quantum Computing Breakthroughs · Y Combinator
LATENT SPACE Prediction Not checkable as stated
Agarwal: Logit Distillation Can Match Giant Teacher Models on Reasoning
“My hunch is that the logic-based distillation can go even further, and you might be able to even close the gap with the biggest of the teachers you have, because I don't think you need a huge number of parameters, because the reasoning process is very, very, l…”
Rishabh Agarwal Mar 23, 2025 ▶ 42:39 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
LATENT SPACE Prediction Not checkable as stated
Kilpatrick: Reasoning and test-time compute will scale faster medium-term
“And it feels like that is. More likely in the medium term going to be the thing that like continues to just like rapidly scale up relative to the other capabilities.”
Logan Kilpatrick Feb 28, 2025 ▶ 5:36 Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
TBPN Insight
Coogan: Inference-Time Compute Establishes an Independent Scaling Law from Pre-Training
“This development creates a new scaling law that is totally independent of the original pre-training scaling law. Now you still want to train the best model you can by clearly leveraging as much compute as you can as many trillion tokens of high quality trainin…”
John Coogan Jan 28, 2025 ▶ 20:06 DeepSeek Update, Market Crash, Timeline in Turmoil, Is VC Cooked, Zero Cope Policy
NO PRIORS Insight
Inference-Time Compute Lets Labs Scale Intelligence Without Doubling Supercomputers
“I don't need to go double the size of my supercomputer to hit a requisite intelligence threshold. I can just double the amount of inference time compute that my customers pay for.”
Aidan Gomez Nov 21, 2024 ▶ 28:12 No Priors Ep. 91 | With Cohere Co-Founder and CEO Aidan Gomez
NO PRIORS Assertion Not checkable as stated
Inference-Time Compute Does Not Require Densely Interconnected Supercomputers
“If we have a new avenue, which is inference time compute, That doesn't require this densely interconnected supercomputer. It's fine to have nodes. You can do a lot more locally and less distributed.”
Aidan Gomez Nov 21, 2024 ▶ 29:28 No Priors Ep. 91 | With Cohere Co-Founder and CEO Aidan Gomez
NO PRIORS Opinion
Mehta: AlphaProof's RL scaling and test-time compute generalize across domains
“Some of the sort of tech we developed here of like, you know, like scaling RL and like figuring out how to spend a lot of inference time compute stuff like this feels like it's Quite generally applicable to many other problems.”
Rishi Mehta Nov 14, 2024 ▶ 21:49 No Priors Ep. 90 | With Google's DeepMind's AlphaProof Team
NO PRIORS Insight
Steinberger: AI performance requires trading off training compute against inference compute
“Well, so you can think of model performance as some function of training compute times some function of inference time compute. Now those are specific functions that are just scaling law things that you can like model, but the general Way to think about it is …”
Eric Steinberger Aug 30, 2024 ▶ 7:32 No Priors Ep. 79 | With Magic.dev CEO and Co-Founder Eric Steinberger
NO PRIORS Insight
Inference-time compute is the missing scaling dimension for AI reasoning
“This is why I'm interested in the reasoning direction, because I think there's this whole other dimension. That people are not scaling right now, which is the amount of compute at inference time.”
Noam Brown Apr 25, 2023 ▶ 17:12 No Priors Ep. 1 | With Noam Brown, Research Scientist at Meta

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.