Inference Cost

topic on 13 shows · 17 statements across 17 episodes

Another Podcast Latent Space Lenny's Podcast the Neon Show No Priors the Official SaaStr Podcast Invest Like the Best Sourcery the MAD Podcast the a16z Podcast Big Technology All-In 20VC

17 statements about Inference Cost, every show

INVEST LIKE THE BEST Assertion Not checkable as stated
Thompson: Chinese open models like Kimi have high marginal inference costs
“You still have to run inference like GLM or Kimi. Kimi is very expensive to serve. The cost per answer is significantly higher.”
Ben Thompson Aug 18, 2026 ▶ 6:25 What Happens When the AI Boom Runs Out of Money · Invest Like The Best
NEON SHOW Assertion Not checkable as stated
Tewari: AI Inference Is Glance's Largest Operating Cost
“In the world of AI, as we launch Glance, what is our biggest cost? Our biggest cost is inference cost. Every interaction that you do with Glance and try to find, use the intelligence of Glance to search for products and to look for product, there is an inferen…”
Naveen Tewari Jul 3, 2026 ▶ 1:01:17 He Built "GPT For Shopping" And Got 450M+ Users | Naveen Tewari
ALL-IN Assertion Supported
Gerstner: AI inference cost is down 90% year over year
“Inference cost is down by 90% year over year.”
Brad Gerstner Apr 10, 2026 ▶ 54:39 Anthropic’s $30B Ramp, Mythos Doomsday, OpenClaw Ankled, Iran War Ceasefire, Israel's Influence
20VC Prediction Open · timeframe Jan 2027
Lemkin: AI startups must model inference costs rising in 2026
“You need to model in your inference costs are going up this year, not down.”
Jason Lemkin Jan 29, 2026 ▶ 23:45 Anthropic Inference Costs Skyrocket |TikTok Deal Closes |The IPO Market:Wealthfront & EquipmentShare · 20VC with Harry Stebbings
Cameron: Inference economics incentivize larger, sparser AI models over dense architectures
“It's, I think, less about total parameters in many cases when thinking about inference costs and more around number of active parameters, and so there's a bit of an incentive towards larger, sparser models.”
George Cameron Jan 9, 2026 ▶ 38:16 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Wu: Cheaper inference does not cut developer spending due to surging demand
“What we realized is as we make it cheaper, you know, the demand for that goes up even more, and you end up, you know, still spending quite a bit”
Sherwin Wu Oct 7, 2025 ▶ 32:11 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
NEON SHOW Prediction Not checkable as stated
Ganesan: US AI firms will ignore cost reduction and chase AGI instead
“Even at the current inference costs I think that the savings in the west is so high that there will be, there is actually no motivation for American companies to further reduce inference costs going forward. So I think they're going to just forget about this c…”
Shivakumar Ganesan Sep 5, 2025 ▶ 32:02 8 Years Without Funding to $100M Raise & Now A Category Leader | Shivku Ganesan ,Exotel
SAASTR Assertion Not checkable as stated
Consumer Demand for Top AI Models Has Increased Inference Costs 100X
“Instead of models getting cheaper, yes, maybe the running the same model got cheaper. But people trained much bigger models that are much more expensive to run now, and people expect to use the best model. So running inference in general, maybe a hundred X in …”
Gorkem Yurtseven Sep 5, 2025 ▶ 8:58 Anthropic, Cursor, FaL & Bessemer Venture Partners: The Realities of Scaling AI
LATENT SPACE Prediction Not checkable as stated
Morcos: AI inference costs will skyrocket, penalizing oversized models
“The inference costs are going to skyrocket with these models. And if you use a general purpose model, then you constrain to say, hey, this model knows about everything, but now only do this one thing. That model is going to have a ton of parameters that do not…”
Ari Morcos Aug 29, 2025 ▶ 59:00 Better Data is All You Need — Ari Morcos, Datology
Kurian: Long-term AI economics depend primarily on inference cost, not training
“First and foremost, in the long run, if AI really scales, the cost you really want to care about is inference cost, because that's what's integrated into serving, and any company that wants to recover the cost of training has to have a large scale inference fo…”
Thomas Kurian Apr 9, 2025 ▶ 15:24 Google Cloud CEO Thomas Kurian on AI Competition, Agents, And Tariffs
SOURCERY Assertion Supported
Tunguz: 100x cost difference between small and largest AI models
“I was just looking at the analysis between the smallest models, which are about four or two to four billion parameters, and the very largest models, which about four or fifty billion parameters, you have a hundred X difference in inference costs.”
Tomasz Tunguz Jan 23, 2025 ▶ 11:32 Inside Theory Ventures: Tomasz Tunguz’s $688M Thesis on Go-To-Market Disruption · Sourcery with Molly O'Shea
LENNY'S PODCAST Assertion Supported
Komoroske: Advertising cannot cover high inference costs for AI consumer startups
“And so if you're going to do a consumer startup, it can't be based on advertising. It's just too expensive. Advertising cannot clear the inference cost even with it with inference costs declining.”
Alex Komoroske Oct 3, 2024 ▶ 11:49 Thinking like a gardener, slime mold, the adjacent possible: Product advice from Alex Komoroske
20VC Assertion Supported
Narayanan: Inference costs dominate training costs for popular AI models
“Over the lifetime of a model, when you have billions of people using it, the inference cost actually adds up, and for many of the popular models, that's the cost that dominates.”
Arvind Narayanan Aug 28, 2024 ▶ 15:33 Arvind Narayanan: AI Scaling Myths, The Core Bottlenecks in AI Today & The Future of Models | E1195 · 20VC with Harry Stebbings
ANOTHER PODCAST Assertion Not checkable as stated
Evans: High inference costs prevent free 100-million-user consumer AI apps
“Because at the moment you can't make a consumer app that's free and have a hundred million users because you can't afford the inference cost.”
Benedict Evans Jul 1, 2024 ▶ 4:42 The AI summer
a16z Insight
Mensch: Overtraining models past Chinchilla limits lowers inference costs
“If you take into account the fact that your model Should also be efficient at inference time. You probably want to go far beyond the Cinchilla scaling low. So it means you want to overtrain the model. So train on more tokens than would be optimal for performan…”
Arthur Mensch Dec 28, 2023 ▶ 7:15 Safety in Numbers: Keeping AI Open
NO PRIORS Insight
Mensch: Pure scientific model performance ignores crucial runtime inference costs
“And if you want to push the performance, the pure performance of models, you don't care about inference because you, well, you are not going to use the model. You're just going to see whether they're good or not. And that's really for scientific purposes. But …”
Arthur Mensch Nov 9, 2023 ▶ 7:14 No Priors Ep. 40 | With Arthur Mensch, CEO Mistral AI
MAD Disclosure
Pesenti: Inference remains Facebook's largest machine learning cost
“Actually, To be clear, the most costly thing we do in ML is still inference cost, because when you put a piece of content within Facebook, it's running hundreds of different ML based algorithm, and they all run on machine parallel, and it's using a huge number…”
Jerome Pesenti Jun 10, 2020 ▶ 34:33 Fireside Chat: Jerome Pesenti (Head of AI, Facebook) with Matt Turck (Partner, FirstMark)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.