inference

13 statements across 10 episodes · 7 bullish · 2 bearish · 8 people on the record · first statement Apr 25, 2023 by Matei Zaharia · across every show →

Everything said about inference, oldest first

Apr 25, 2023 bearish
Opinion
Zaharia: Trillion-parameter models are computationally inefficient for knowledge retrieval
“I think actually, I think from a computation perspective, it's very inefficient to have like a trillion parameters and have to actually load them all and add and multiply by them. Each time you make an inference, because they're just encoding knowledge, most o…”
Matei Zaharia Apr 25, 2023 ▶ 32:43 No Priors Ep. 11 | With Matei Zaharia, CTO of Databricks
Aug 10, 2023
Assertion Not checkable as stated
Guo: Inference already dominates OpenAI compute usage over training
“Inference, like inference already dominates open AI compute usage, right?”
Sarah Guo Aug 10, 2023 ▶ 7:29 No Priors Ep. 27 | With Sarah Guo & Elad Gil
Sep 7, 2023 bullish
Prediction Not checkable as stated
Feldman: Commercial AI will settle on 3B to 13B parameter models
“And so we're going to be down at three and at six billion and at thirteen billion, because that's, I can get pretty good, pretty darn good, and not break the bank with free inference.”
Andrew Feldman Sep 7, 2023 ▶ 27:28 No Priors Ep. 31 | With Cerebras CEO Andrew Feldman
Sep 14, 2023 bullish
Insight
Polosukhin: Inference demands vastly more aggregate compute than AI model training
“I think an inference is really interesting because we do need so much more compute for inference than we need for training, right? Like it's a very interesting like economy of scale. You train once, like Lama trained once and then everybody runs it everywhere.”
Illia Polosukhin Sep 14, 2023 ▶ 23:23 No Priors Ep. 32 | With NEAR’s Illia Polosukhin
Mar 21, 2024 neutral
Insight
Srivastava: AI inference demands strict uptime, while training tolerates node terminations
“Resiliency and reliability matters a lot more. You know, downtime is unacceptable from an input perspective. Nodes get terminated all the time from a training perspective.”
Tuhin Srivastava Mar 21, 2024 ▶ 6:01 No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava
Mar 21, 2024 neutral
Insight
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
Tuhin Srivastava Mar 21, 2024 ▶ 4:57 No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava
May 9, 2024 neutral
Assertion Not checkable as stated
Sarah Guo says massive frontier AI models are impossible to serve commercially
“Over time, applications are going to want efficient inference, and, like, really large models are impossible today to serve for the vast majority of use cases from a cost and speed perspective”
Sarah Guo May 9, 2024 ▶ 15:07 No Priors Ep. 63 | With Sarah Guo and Elad Gil
Aug 29, 2024 bullish
Prediction Not checkable as stated
AI inference will become a core cloud computing primitive like storage
“As you move forward, generative AI honestly becomes one of the compute building blocks that you think about. You're going to need storage, you need compute, you need databases, you need inference, if you will, for your application, largely. And I think that's …”
Matt Garman Aug 29, 2024 ▶ 40:49 No Priors Ep. 78 | With AWS CEO Matt Garman
Jan 9, 2025 positive
Disclosure
Bernhardsson: Modal Is Expanding into Bursty Experimental AI Training
“Traditionally, most of modal has always been inference. Like that's been our main use case, but we're really interested also in training. So in particular, like probably focused more on these like shorter, like very bursty sort of experimental training runs, n…”
Erik Bernhardsson Jan 9, 2025 ▶ 6:58 No Priors Ep. 96 | With Modal CEO and Founder Erik Bernhardsson
Aug 7, 2025 bullish
Prediction Not checkable as stated
Prince: In-network edge inference will handle models too large for end devices
“We believe that a lot of inference is going to happen on your end device, but there will always be some model which is too big or too resource intensive. And in that case, the next best place to run it is going to be on at the inside the network at the edge.”
Matthew Prince Aug 7, 2025 ▶ 23:27 No Priors Ep. 126 | With Cloudfare CEO Matthew Prince
May 1, 2026 bullish
Insight
Srivastava: Even After AGI Is Achieved, Inference Is All That Remains
“Even if there's AGI, all that's left is inference.”
Tuhin Srivastava May 1, 2026 ▶ 40:21 Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
May 1, 2026 bullish
Insight
Srivastava: Open-Source Baseline and Post-Training Enable In-House Inference
“The open source models have crossed some sort of chasm in terms of their baseline. Capability, and then I think RL techniques and post-training is for specialized models has become mainstream enough, and, you know, there's enough examples of its work, of it wo…”
Tuhin Srivastava May 1, 2026 ▶ 1:10 Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
May 1, 2026 negative
Insight
Srivastava: Raw GPU hosting is an unsticky commodity compared to inference software
“GPUs as a service is not sticky. I think that's been seen. Like, customers generally just see that as commodity. Inference with the software layer included is incredibly sticky.”
Tuhin Srivastava May 1, 2026 ▶ 24:46 Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.