AI Inference

topic on 15 shows · 57 statements across 44 episodes

the Y Combinator Startup Podcast BG2 Pod the Knowledge Project Latent Space the Neon Show No Priors Invest Like the Best Sourcery Catalyst the MAD Podcast the a16z Podcast Big Technology All-In TBPN 20VC

57 statements about AI Inference, every show

BIG TECHNOLOGY Assertion Supported
Kedrosky: GPU failure rates are much higher in training than inference
“The failure rates of GPUs used so intensively for training purposes are much higher than inference specific usage.”
Paul Kedrosky Aug 12, 2026 ▶ 18:05 Why The AI Bubble Will Burst: The Most Logical Case — With Paul Kedrosky
20VC Insight
Atallah: Rare Disease Research Has Been Constrained by an Inference Bottleneck
“One is rare disease research, which I think is one of those things that has been intelligence bottleneck or really just the inference bottleneck. Like it involves like trying out lots of ideas and seeing if they work.”
Alex Atallah Aug 9, 2026 ▶ 1:06:42 OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic
Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Misha Smulyanski Jul 29, 2026 ▶ 47:35 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
TBPN Assertion Supported
Tae Kim: Morgan Stanley estimates AI inference yields 60-80% margins
“Like Morgan Stanley says, if you do inference, it's 60 to 80% profit margins, right?”
Tae Kim Jul 29, 2026 ▶ 15:11 RSI Is Closer Than People Think, Per Tae Kim
MAD Insight
Feldman: AI inference is bottlenecked by data movement, causing GPU slowness
“In inference in AI, it's the exact opposite. You move a huge amount of data, all the weights, from memory to compute, and you need one calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data…”
Andrew Feldman Jul 23, 2026 ▶ 32:16 Cerebras CEO: Why GPUs Can't Do Fast Inference
SOURCERY Prediction Open · timeframe Jul 2031
Liang: AI inference chip deployments will dwarf training by orders of magnitude
“Because at scale, the number of chips deployed for inferencing will be orders of magnitude greater than whatever you're doing for training.”
Rodrigo Liang Jul 17, 2026 ▶ 4:37 Inference 101: SambaNova CEO Rodrigo Liang · Sourcery with Molly O'Shea
SOURCERY Prediction Not checkable as stated
Liang: Demand will converge on the fastest, most accurate large models
“And so, as the cost of delivering fast goes down, you're going to see most people switch over to the fastest. And this is why I feel like, you know, the premium inference, which is large models, which equals the most accurate. The most accurate models and fast…”
Rodrigo Liang Jul 17, 2026 ▶ 20:13 Inference 101: SambaNova CEO Rodrigo Liang · Sourcery with Molly O'Shea
SOURCERY Assertion Partly supported
Feldman: Cerebras AI inference is 20x faster than competitors
“And we're the fastest, not by a little bit, but by 20 x.”
Andrew Feldman Jul 13, 2026 ▶ 4:02 Andrew Feldman on Building a Chip 58x Larger Than Nvidia's · Sourcery with Molly O'Shea
INVEST LIKE THE BEST Prediction Not checkable as stated
Wachen: Software COGS per incremental user will become high due to inference
“Like, the cost structure of every software company, the COGS is not going to be, like, zero anymore for an incremental user. It's going to be, like, quite high, and it's going to be a function of inference.”
Rob Wachen Jun 30, 2026 ▶ 21:55 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
INVEST LIKE THE BEST Prediction Open · timeframe Jun 2036
Wachen: AI inference will become the biggest market in the world
“It feels like we are on a, like, decade march for inference to become, you know, the biggest market in the world.”
Rob Wachen Jun 30, 2026 ▶ 22:08 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
INVEST LIKE THE BEST Prediction Not checkable as stated
Wachen: AI inference will eventually make up the majority of global GDP
“I firmly believe we are on a global march of inference becoming majority of global GDP. And it may take more than 10 years, but it's going to happen.”
Rob Wachen Jun 30, 2026 ▶ 1:24:40 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
NEON SHOW Prediction Didn’t hold up
Ambati: AI inference market will reach $1.3T-$1.4T by 2030
“The inference market is one of the largest or so the fastest growing CAGR. So what we are seeing is about 94% CAGR. And then, you know, today, I think around 20, 26, we are looking at an eighty billion dollar sort of market. But soon by 2030 it's gonna be a 1.…”
Vamshi Ambati Jun 9, 2026 ▶ 17:46 94% CAGR: What the Inference Boom means for your AI costs | Vamshi Ambati
20VC Prediction Not checkable as stated
Feldman predicts there will be zero market for slow AI inference
“Why do we believe that inference will be any different? There'll be zero marking for slowing them.”
Andrew Feldman May 26, 2026 ▶ 22:19 Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China · 20VC with Harry Stebbings
NO PRIORS Assertion Not checkable as stated
Srivastava: AI-Native Startups Represent 99% of Total Inference Call Volume
“I think if you look by inference count, it'd be 99% the full.”
Tuhin Srivastava May 1, 2026 ▶ 5:06 Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
Hassabis doubts inference will ever be free due to Jevons' paradox
“Yeah, I'm not sure inference will ever be essentially free. I mean, there's sort of Jevon's paradox and other things about like, I think we'll just end up using All of us will end up using whatever we can get our hands on, and you could imagine millions of age…”
Demis Hassabis Apr 29, 2026 ▶ 24:11 Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough · Y Combinator
Y COMBINATOR Prediction Not checkable as stated
Hassabis: Chip manufacturing will bottleneck AI inference for decades
“Certainly the energy, if we solve fusion or, you know, superconductors or, you know, optimal batteries or some set of those things, which I think we will do with material science, energy costs will be essentially zero, but there'll still be the physical creati…”
Demis Hassabis Apr 29, 2026 ▶ 24:51 Demis Hassabis: Agents, AGI & The Next Big Scientific Breakthrough · Y Combinator
CATALYST Insight
Vahdat: AI inference does not require gigawatt-scale data center capacity
“You don't strictly need a gigawatt of capacity to be able to do useful work. You probably don't even need a hundred megawatts of capacity. It gets a little bit more interesting than that because of, let's say, co-located compute and storage and networking and …”
Amin Vahdat Apr 23, 2026 ▶ 6:25 Inside Google’s massive AI capex (live)
ALL-IN Opinion
Huang: AI inference pipeline is the most complicated computing problem today
“The processing pipeline of inference is extremely complicated. And in fact, it is the most complicated computing problem today.”
Jensen Huang Mar 19, 2026 ▶ 1:46 Jensen Huang: Nvidia's Future, Physical AI, Rise of the Agent, Inference Explosion, AI PR Crisis
NO PRIORS Prediction Not checkable as stated
Tiwari: AI inference compute will shift toward decentralized clusters
“What you're seeing with inference is in many use cases, as this becomes more ubiquitous, you're going to have more and more decentralized inference clusters.”
Neil Tiwari Feb 26, 2026 ▶ 19:29 Who's Actually Funding the AI Buildout?
NO PRIORS Assertion Not checkable as stated
Guo: AI inference clouds are growing 1,000x in compute consumption
“I think to the point of like real data, the inference clouds are growing a thousand X in terms of consumption, right? And then they're getting more efficient. So revenue grows at some lower rate than that, but it's wild.”
Sarah Guo Feb 19, 2026 ▶ 19:44 The AI Code Slop: Risk or Opportunity?
TBPN Insight
Coogan: AI inference will be load-balanced across diverse chip architectures
“It does feel like we're going to enter a world where inference is load balanced across a variety of semiconductor stacks, and so, For really fast things, you might be going to a Grok or a Cerebris, or, you know, you might, for more basic stuff, you might be go…”
John Coogan Jan 23, 2026 ▶ 22:00 Capital One Acquires Brex for $5B, Davos’s $500M business, Apple’s Siri rebuild | Diet TBPN
BIG TECHNOLOGY Prediction Held up
Amon: AI inference will eventually surpass training in data centers
“Eventually inference is going to take over training because Just think about that for a second. If you're a company spending billions of dollars building a data center for training, you expect to get a return on that investment. So when you start putting AI in…”
Cristiano Amon Jan 20, 2026 ▶ 34:03 Qualcomm CEO Cristiano Amon: Future Of AI Devices, AI Fashion, Blending Reality and Computing
BIG TECHNOLOGY Disclosure
Amon: Qualcomm is building post-GPU chips for data center inference
“So we're building what we believe is post-GPU. When you started to do inference and you need the dedicated engines, We're building that. I actually believe that the NVIDIA acquisition of Croc validates that you different engines for different things, and I thi…”
Cristiano Amon Jan 20, 2026 ▶ 37:50 Qualcomm CEO Cristiano Amon: Future Of AI Devices, AI Fashion, Blending Reality and Computing
20VC Prediction Not checkable as stated
Lovinsky: Continuous background AI inference won't reach general knowledge work by 2026
“Whether it happens widely in knowledge work, By 20, 26, I think is, you know, I'll take a bet with you on that. Certainly for coding. I think in a lot of places we're already there, right? I mean, you have like the Ralph Wiggum stuff that, that popped up over,…”
Noam Lovinsky Jan 16, 2026 ▶ 32:35 20Product: Is the Design Phase Dead in a World of AI | Has Claude Code Crushed Anthropic Already | What Roles of a PM Are Less and More Important with AI | How the Best Product Leaders Tell Stories with Noam Lovinsky, CPO @ Superhuman
BIG TECHNOLOGY Assertion Not checkable as stated
Venturo: No difference between training and inference infrastructure for modern AI
“So there's really no difference between training infrastructure we deployed to build those capabilities and what our customers are ultimately using to serve them.”
Brian Venturo Jan 7, 2026 ▶ 14:40 Coreweave: AI Bubble Poster Child Or The Next Tech Giant? — With Michael Intrator and Brian Venturo
BIG TECHNOLOGY Assertion Not checkable as stated
Venturo: CoreWeave compute demand has shifted to roughly 50-50 training and inference
“Six months ago, I would have said it was two-thirds training and one-third inference. It's probably close to fifty-fifty now.”
Brian Venturo Jan 7, 2026 ▶ 14:59 Coreweave: AI Bubble Poster Child Or The Next Tech Giant? — With Michael Intrator and Brian Venturo
CATALYST Prediction Not checkable as stated
AI inference spending must surge to justify current adoption optimism
“If the optimism about AI is to be justified, you're going to have to see inference costs go way up because that will be an indicator that adoption has gone up in a fairly significant way, both among individual users, but also among companies and enterprise use…”
Dr. Ben Lee Dec 18, 2025 ▶ 12:26 Will inference move to the edge?
CATALYST Insight
AI training compute cycles cause massive power fluctuations unlike inference workloads
“And some of the people in the energy space may know that there are massive energy fluctuations or power fluctuations we will see in data center usage when the GPUs go from this computational intensive phase where you're learning the model weights to this commu…”
Dr. Ben Lee Dec 18, 2025 ▶ 18:09 Will inference move to the edge?
CATALYST Prediction Not checkable as stated
Apple is the most likely company to move AI inference on-device
“It's not hard to picture that, like, if somebody's gonna move a lot of this inference on device, it's gonna be Apple.”
Shayle Kann Dec 18, 2025 ▶ 35:46 Will inference move to the edge?
CATALYST Prediction Open · timeframe Dec 2030
Software agents, rather than human queries, will drive most AI inference workloads
“I think increasingly most of the inference workload will come from other software agents.”
Dr. Ben Lee Dec 18, 2025 ▶ 46:40 Will inference move to the edge?
20VC Insight
Randle: Investors should follow momentum when AI inference demand is unprecedented
“When you have demand like this, like you have for the initial GC, like the initial hyperscaler clouds. And I think we're seeing an even greater cohorted demand curve for AI inference. Sometimes you just got to shut your mind up and invest with the momentum.”
Everett Randle Nov 10, 2025 ▶ 27:15 Benchmark's GP, Everett Randle on Why Mega Funds Will Not Produce Good Returns · 20VC with Harry Stebbings
Corbitt: AI inference could be 10x larger if reliability issues are solved
“I think that there is today, like. 10 times as much AI inference that could exist than is existing right now, just Purely with projects that are like sitting in the proof of concept stage and have not been deployed because there's like huge bucket of those. An…”
Kyle Corbitt Oct 16, 2025 ▶ 1:04:50 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
20VC Insight
Feldman: AI inference growth compounds across users, frequency, and compute per query
“The greater growth of inference is the number of people who use it, Times the frequency of use. Times the amount of compute needed per use. Right? It is three different variables multiplied by each other. The problem is they're all growing fast.”
Andrew Feldman Oct 6, 2025 ▶ 26:40 Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
20VC Insight
Ross: Lowering AI chip prices 50% leads customers to buy double
“If we lower what we charge 50%, people are gonna buy twice as much. They're spending as much as they're making because whatever they spend increases the quality of the output.”
Jonathan Ross Sep 29, 2025 ▶ 1:21:05 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
20VC Opinion
Jeff Lawson: AI inference is not complex enough to sustain premium margins
“I believe inference itself is not such a hard algorithmic solve that, you know, you need to pay someone else to do it for you, but clearly training a model is”
Jeff Lawson Sep 11, 2025 ▶ 1:22:45 OpenAI’s $10BN Secondary Sale, Ramp Hits $1BN ARR & Brex Hits $700M · 20VC with Harry Stebbings
PyTorch democratized model training, but AI inference remains undemocratized
“And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. Inference…”
Chris Lattner Jun 13, 2025 ▶ 35:44 The Shape of Compute (Chris Lattner of Modular)
AI training scales with research teams; inference scales with customer base
“Training scales the size of your research team. Inference scales the size of your customer base.”
Chris Lattner Jun 13, 2025 ▶ 59:12 The Shape of Compute (Chris Lattner of Modular)
CATALYST Prediction Held up
Sharp: AI inference, not training, will drive future data center growth
“I would say that there's been a lot of growth in that, but we see that kind of leveling out where we really see the consumption of AI or inference that's driving that regional specific growth going forward. And I think that's where it's more embedded. And a lo…”
Chris Sharp Jun 12, 2025 ▶ 4:44 The state of play of data center development
CATALYST Assertion Not checkable as stated
Sharp: AI inference deployments require 5-megawatt blocks, unlike training arrays
“What we're really seeing is inference. Can come in, in like, five-ish megawatt blocks, and you can solve for it a bit differently. Now, the densification is still there, and then the private AI pieces, there's hot spots where it can be, you know, a couple of m…”
Chris Sharp Jun 12, 2025 ▶ 18:56 The state of play of data center development
$20/month local IDE pricing limits AI inference quality per task
“The cost when you are local first and your typical consumer is on a free plan or a Like, 20 dollar a month paid plan limits the amount of high quality inference you can do, and the scale or volume of inference you can do per, like, outcome.”
Eno Reyes May 29, 2025 ▶ 9:22 The AI Coding Factory
INVEST LIKE THE BEST Prediction Held up
Consumer tech products will increasingly tier pricing based on AI inference volume
“It's very likely that some consumers are going to want tons and tons of inference. And because that's a marginal cost, you're probably going to have to pay somehow for that as a consumer. So I think you're going to see more tiering of consumer products based o…”
Gustav Söderström May 20, 2025 ▶ 44:58 Spotify’s Journey To Profitability · Invest Like The Best
INVEST LIKE THE BEST Prediction Not checkable as stated
Overall demand for intelligence will outpace the falling costs of AI compute
“I think from a financial point of view when the cost of something drops, the demand usually increases more than the cost, than the drop. And I think that's, Bound to happen with intelligence.”
Gustav Söderström May 20, 2025 ▶ 47:46 Spotify’s Journey To Profitability · Invest Like The Best
Conrad: Custom chips make sense for inference, not training experimentation
“It only works if you really know which chip you're going to do. If you don't, then it's a little harder. So it makes, in my head, it makes more sense for inference where you've already established it, but for training there's so much, like, experimentation.”
Evan Conrad Apr 11, 2025 ▶ 20:17 SF Compute: Commoditizing Compute
MAD Prediction Not checkable as stated
Ramaswamy: AI inference will become significantly faster and cheaper in 2025
“The world of inference, not foundation models, the world of inference is going to go through so much change this year, both in terms of GPU availability, which is easing up quite a bit, But also in terms of people like Grok, I think the one with the K or the Q…”
Sridhar Ramaswamy Apr 10, 2025 ▶ 1:16:38 Snowflake CEO on Winning the AI Arms Race
20VC Assertion Supported
Boland: NVIDIA Is Re-Architecting GPUs to Optimize for Inference
“They are making a bunch of architectural changes to GPUs to make them better and better inference.”
Stan Boland Apr 10, 2025 ▶ 1:14:23 Tom Hulme & Stan Boland: Lessons from Jensen Huang & How to Fix the UK Tech Ecosystem · 20VC with Harry Stebbings
20VC Assertion Supported
Feldman: Nvidia Has Zero CUDA Software Lock-In for AI Inference
“In inference, it's not real at all. There's no CUDA lock-in in inference. None. Well, you can move from OpenAI on an NVIDIA GPU to Cerebrus to Fireworks Service on something else to Together to perplexity with 10 keystrokes. I mean, anybody who actually uses A…”
Andrew Feldman Mar 24, 2025 ▶ 42:10 Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
CATALYST Assertion Supported
Microsoft AI screened 32 million battery electrolyte candidates down to 100
“What you're referring to is a project where AI was leveraged by our research teams to screen over thirty-two million candidates for a new electrolyte for advanced batteries. Thanks to AI inference, you know, the R&D scientists were able to go from that number …”
Laurent Bueneau Mar 11, 2025 ▶ 7:08 How AI is solving real utility challenges [partner content]
KNOWLEDGE PROJECT Prediction Open · timeframe Mar 2030
Wolfe: 30% to 50% of AI inference will occur on-device
“I think that you're going to end up doing a lot of inference on device. Meaning, instead of going, ah, to the cloud and typing a query, like, 30 to 50% of your inference may be on an Apple or an Android device.”
Josh Wolfe Mar 4, 2025 ▶ 9:16 Josh Wolfe: Human Advantage in the World of AI
20VC Insight
Morin: Interconnect dependency is the core difference between training and inference
“In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of…”
Steeve Morin Feb 24, 2025 ▶ 17:58 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
20VC Assertion Not checkable as stated
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Steeve Morin Feb 24, 2025 ▶ 21:35 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
20VC Assertion Not checkable as stated
Ross: The AI industry mistakenly believed training was costlier than inference
“When we started, the first misconception, which people don't hold anymore, is that training was more expensive than inference.”
Jonathan Ross Feb 17, 2025 ▶ 9:16 Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260 · 20VC with Harry Stebbings
a16z Prediction Not checkable as stated
Ulrich: Proprietary data and inference will differentiate future AI applications
“I do think that the use of data becomes a more critical differentiator. Inference becomes a more critical differentiator.”
Greg Ulrich Feb 5, 2025 ▶ 29:26 How AI is Powering Payments, with Greg Ulrich of Mastercard
BG2 Prediction Not checkable as stated
Jensen Huang: AI inference demand will scale up by one billion times
“It's gonna go up a billion times.”
Jensen Huang Jan 11, 2025 ▶ 17:23 Market Predictions, Rates & Inflation, DOGE, CES, AI Compute | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
BG2 Prediction Not checkable as stated
Huang: AI inference compute is about to increase by a billion times
“It's about to go up by a billion times.”
Jensen Huang Oct 13, 2024 ▶ 58:25 Ep17. Welcome Jensen Huang | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Zhang: Next 10x AI inference gain requires multi-layer co-optimization
“If you only push on one direction, you are going to reach diminution return really, really quickly. Yeah, there's only that much you can do on the system side, only that much you can do on the algorithm side. And since the only big thing that's going to happen…”
Ce Zhang Feb 8, 2024 ▶ 40:16 Building an open AI company - with Ce and Vipul of Together AI
Doshi: AI inference DevOps is similar to scaling high-volume API servers
“I don't find, I find the DevOps for inference to be relatively easy. It doesn't feel that different than, you know, I think we had thousands and thousands of servers at Mixpanel just for dealing with the API had such huge quantities of volume that I didn't fin…”
Suhail Doshi Jan 2, 2024 ▶ 1:02:35 The AI-First Graphics Editor - with Suhail Doshi of Playground AI
a16z Assertion Not checkable as stated
Baszucki says Roblox will provide free AI inference to its creators.
“That the more we can run inference jobs on these, we can run super high volume inference at high quality at low cost and make this, you know, just freely available so the creators don't worry about it.”
David Baszucki Sep 25, 2023 ▶ 7:18 Leveling Up with Roblox's David Baszucki

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.