AI inference

12 statements across 11 episodes · 6 bullish · 3 bearish · 8 people on the record · first statement Feb 17, 2025 by Jonathan Ross · across every show →

Everything said about AI inference, oldest first

Feb 17, 2025 neutral
Assertion Not checkable as stated
Ross: The AI industry mistakenly believed training was costlier than inference
“When we started, the first misconception, which people don't hold anymore, is that training was more expensive than inference.”
Jonathan Ross Feb 17, 2025 ▶ 9:16 Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260 · 20VC with Harry Stebbings
Feb 24, 2025
Insight
Morin: Interconnect dependency is the core difference between training and inference
“In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of…”
Steeve Morin Feb 24, 2025 ▶ 17:58 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
Feb 24, 2025 positive
Assertion Not checkable as stated
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Steeve Morin Feb 24, 2025 ▶ 21:35 Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
Mar 24, 2025 negative
Assertion Supported
Feldman: Nvidia Has Zero CUDA Software Lock-In for AI Inference
“In inference, it's not real at all. There's no CUDA lock-in in inference. None. Well, you can move from OpenAI on an NVIDIA GPU to Cerebrus to Fireworks Service on something else to Together to perplexity with 10 keystrokes. I mean, anybody who actually uses A…”
Andrew Feldman Mar 24, 2025 ▶ 42:10 Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
Apr 10, 2025 positive
Assertion Supported
Boland: NVIDIA Is Re-Architecting GPUs to Optimize for Inference
“They are making a bunch of architectural changes to GPUs to make them better and better inference.”
Stan Boland Apr 10, 2025 ▶ 1:14:23 Tom Hulme & Stan Boland: Lessons from Jensen Huang & How to Fix the UK Tech Ecosystem · 20VC with Harry Stebbings
Sep 11, 2025 bearish
Opinion
Jeff Lawson: AI inference is not complex enough to sustain premium margins
“I believe inference itself is not such a hard algorithmic solve that, you know, you need to pay someone else to do it for you, but clearly training a model is”
Jeff Lawson Sep 11, 2025 ▶ 1:22:45 OpenAI’s $10BN Secondary Sale, Ramp Hits $1BN ARR & Brex Hits $700M · 20VC with Harry Stebbings
Sep 29, 2025 neutral
Insight
Ross: Lowering AI chip prices 50% leads customers to buy double
“If we lower what we charge 50%, people are gonna buy twice as much. They're spending as much as they're making because whatever they spend increases the quality of the output.”
Jonathan Ross Sep 29, 2025 ▶ 1:21:05 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
Oct 6, 2025 bullish
Insight
Feldman: AI inference growth compounds across users, frequency, and compute per query
“The greater growth of inference is the number of people who use it, Times the frequency of use. Times the amount of compute needed per use. Right? It is three different variables multiplied by each other. The problem is they're all growing fast.”
Andrew Feldman Oct 6, 2025 ▶ 26:40 Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
Nov 10, 2025 bullish
Insight
Randle: Investors should follow momentum when AI inference demand is unprecedented
“When you have demand like this, like you have for the initial GC, like the initial hyperscaler clouds. And I think we're seeing an even greater cohorted demand curve for AI inference. Sometimes you just got to shut your mind up and invest with the momentum.”
Everett Randle Nov 10, 2025 ▶ 27:15 Benchmark's GP, Everett Randle on Why Mega Funds Will Not Produce Good Returns · 20VC with Harry Stebbings
Jan 16, 2026 bearish
Prediction Not checkable as stated
Lovinsky: Continuous background AI inference won't reach general knowledge work by 2026
“Whether it happens widely in knowledge work, By 20, 26, I think is, you know, I'll take a bet with you on that. Certainly for coding. I think in a lot of places we're already there, right? I mean, you have like the Ralph Wiggum stuff that, that popped up over,…”
Noam Lovinsky Jan 16, 2026 ▶ 32:35 20Product: Is the Design Phase Dead in a World of AI | Has Claude Code Crushed Anthropic Already | What Roles of a PM Are Less and More Important with AI | How the Best Product Leaders Tell Stories with Noam Lovinsky, CPO @ Superhuman
May 26, 2026 bullish
Prediction Not checkable as stated
Feldman predicts there will be zero market for slow AI inference
“Why do we believe that inference will be any different? There'll be zero marking for slowing them.”
Andrew Feldman May 26, 2026 ▶ 22:19 Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China · 20VC with Harry Stebbings
Aug 9, 2026 positive
Insight
Atallah: Rare Disease Research Has Been Constrained by an Inference Bottleneck
“One is rare disease research, which I think is one of those things that has been intelligence bottleneck or really just the inference bottleneck. Like it involves like trying out lots of ideas and seeing if they work.”
Alex Atallah Aug 9, 2026 ▶ 1:06:42 OpenRouter CEO: Why Chinese Open Models Are Beating the US | Why Enterprises Fear OpenAI & Anthropic
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.