inference

13 statements across 11 episodes · 4 bullish · 4 bearish · 8 people on the record · first statement Jan 22, 2025 by George Sivulka · across every show →

Everything said about inference, oldest first

Jan 22, 2025 bearish
Prediction Held up
Sivulka: AI shift from training to inference will destabilize Nvidia's dominance
“The shift away from training to inference as a fundamental, like, almost macro shift in how people deploy AI. I actually think that will destabilize slightly the dominance of NVIDIA chips.”
George Sivulka Jan 22, 2025 ▶ 53:40 George Sivulka, Co-Founder & CEO @Hebbia: The Future of Foundation Models | E1250 · 20VC with Harry Stebbings
Jan 29, 2025 bullish
Prediction Open · timeframe Jan 2030
Ross: AI inference will account for 95% of total compute demand
“I think, 95%.”
Jonathan Ross Jan 29, 2025 ▶ 20:58 Jonathan Ross: DeepSeek Special - How Should OpenAI and the US Government Respond | E1253 · 20VC with Harry Stebbings
Feb 24, 2025 bullish
Prediction Not checkable as stated
Morin: Nvidia will remain dominant in AI inference due to availability
“The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Availability that is right. So I think Nvidia is used to stay at least if not for the H-one hundre…”
Steeve Morin Feb 24, 2025 ▶ 25:33 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
Feb 24, 2025 negative
Assertion Contradicted
Morin: Doubling GPUs in AI inference yields only 10% performance gain
“If you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount o…”
Steeve Morin Feb 24, 2025 ▶ 45:22 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
Mar 24, 2025 negative
Opinion
Feldman: GPU off-chip memory architecture can be beaten in inference
“The fundamental architecture of the GPU with off-chip memory is not great for inference. Now, they will continue to do well in inference, but it can be beaten, and I think they know it.”
Andrew Feldman Mar 24, 2025 ▶ 0:17 Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
Mar 24, 2025 negative
Assertion Partly supported
Feldman: GPUs operate at only 5% to 7% utilization during inference
“In a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93% wasted.”
Andrew Feldman Mar 24, 2025 ▶ 25:34 Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
Sep 29, 2025
Insight
Ross: AI Training and Inference Form a Virtuous Hardware Demand Cycle
“The more inference you have, as mentioned before, the more you need to train the model to optimize for the inference. And the more training you have the more inference you want to deploy to optimize for the cost of that training, to amortize the cost of the tr…”
Jonathan Ross Sep 29, 2025 ▶ 44:28 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
Oct 6, 2025 neutral
Assertion Supported
Feldman: Vastly more people do AI inference than AI training
“To move people off GPUs in inference, and the number of people doing inference is vastly higher than the number of people doing training.”
Andrew Feldman Oct 6, 2025 ▶ 26:05 Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
Feb 5, 2026 neutral
Insight
Lemkin: AI inference costs are the new sales and marketing expense
“For simplistic folks, for founders, I say inference is the new sales and marketing.”
Jason Lemkin Feb 5, 2026 ▶ 12:56 The SaaS Massacre: Public Market Collapse |Microsoft Lost $360B & NVIDIA’s $100B Dispute with OpenAI
Feb 16, 2026 bullish
Opinion
Stebbings: Data centers are today's most under-invested technology category
“Data center is the one that's the most under-invested categories today. When you look at inference needing to run for 24 hours a day for most of the knowledge worker population, and it running for like one percent of knowledge worker population today, I'm like…”
Harry Stebbings Feb 16, 2026 ▶ 58:35 Klarna CEO: SaaS is Dead: Why Systems of Record Will Die in an Agentic World · 20VC with Harry Stebbings
Mar 19, 2026 bullish
Prediction Not checkable as stated
Lemkin: Global AI inference volume will grow 1,000x in five years
“Yeah, I mean, I think it's got to be three orders of magnitude more inference we run in the next five years.”
Jason Lemkin Mar 19, 2026 ▶ 7:05 NVIDIA Predicts $1TRN in Revenue: Everything You Need to Know From GTC & Anduril Lands $20B Contract · 20VC with Harry Stebbings
Jun 8, 2026
Insight
Chernin: Bare metal AI compute has only a dozen global customers
“On bare metal level, you have maybe a dozen of the customers in the world that you can work with. On managed infrastructure, there are hundreds. On inference, there are thousands. On agentic, there will be tens of thousands of new developers that build it, rig…”
Roman Chernin Jun 8, 2026 ▶ 20:08 Nebius Co-Founder on AI Infrastructure Bubbles | How Price Elastic is Demand for Compute · 20VC with Harry Stebbings
Aug 2, 2026
Insight
Angelopoulos: Model labs must enter application layer to avoid commoditization
“If, like, inference is going to commoditize, then, of course, the next best thing is for the model providers to be moving up the application layer in order to own more of the application stack so that they ensure that they're not commoditized and they're getti…”
Anastasios Angelopoulos Aug 2, 2026 ▶ 56:20 Arena CEO: There Will be a $100BN US Open-Source Model & Data is a Trillion Dollar Market
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.