inference
13 statements across 10 episodes · 7 bullish · 2 bearish · 8 people on the record · first statement Apr 25, 2023 by Matei Zaharia · across every show →
Everything said about inference, oldest first
Apr 25, 2023 bearish
Zaharia: Trillion-parameter models are computationally inefficient for knowledge retrieval
“I think actually, I think from a computation perspective, it's very inefficient to have like a trillion parameters and have to actually load them all and add and multiply by them. Each time you make an inference, because they're just encoding knowledge, most o…”
Aug 10, 2023
Sep 7, 2023 bullish
Sep 14, 2023 bullish
Polosukhin: Inference demands vastly more aggregate compute than AI model training
“I think an inference is really interesting because we do need so much more compute for inference than we need for training, right? Like it's a very interesting like economy of scale. You train once, like Lama trained once and then everybody runs it everywhere.”
Mar 21, 2024 neutral
Mar 21, 2024 neutral
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
May 9, 2024 neutral
Aug 29, 2024 bullish
AI inference will become a core cloud computing primitive like storage
“As you move forward, generative AI honestly becomes one of the compute building blocks that you think about. You're going to need storage, you need compute, you need databases, you need inference, if you will, for your application, largely. And I think that's …”
Jan 9, 2025 positive
Bernhardsson: Modal Is Expanding into Bursty Experimental AI Training
“Traditionally, most of modal has always been inference. Like that's been our main use case, but we're really interested also in training. So in particular, like probably focused more on these like shorter, like very bursty sort of experimental training runs, n…”
Aug 7, 2025 bullish
Prince: In-network edge inference will handle models too large for end devices
“We believe that a lot of inference is going to happen on your end device, but there will always be some model which is too big or too resource intensive. And in that case, the next best place to run it is going to be on at the inside the network at the edge.”
May 1, 2026 bullish
May 1, 2026 bullish
Srivastava: Open-Source Baseline and Post-Training Enable In-House Inference
“The open source models have crossed some sort of chasm in terms of their baseline. Capability, and then I think RL techniques and post-training is for specialized models has become mainstream enough, and, you know, there's enough examples of its work, of it wo…”