Thompson: Chinese open models like Kimi have high marginal inference costs
“You still have to run inference like GLM or Kimi. Kimi is very expensive to serve. The cost per answer is significantly higher.”
Tewari: AI Inference Is Glance's Largest Operating Cost
“In the world of AI, as we launch Glance, what is our biggest cost? Our biggest cost is inference cost. Every interaction that you do with Glance and try to find, use the intelligence of Glance to search for products and to look for product, there is an inferen…”
Gerstner: AI inference cost is down 90% year over year
“Inference cost is down by 90% year over year.”
Lemkin: AI startups must model inference costs rising in 2026
“You need to model in your inference costs are going up this year, not down.”
Cameron: Inference economics incentivize larger, sparser AI models over dense architectures
“It's, I think, less about total parameters in many cases when thinking about inference costs and more around number of active parameters, and so there's a bit of an incentive towards larger, sparser models.”
Wu: Cheaper inference does not cut developer spending due to surging demand
“What we realized is as we make it cheaper, you know, the demand for that goes up even more, and you end up, you know, still spending quite a bit”
Ganesan: US AI firms will ignore cost reduction and chase AGI instead
“Even at the current inference costs I think that the savings in the west is so high that there will be, there is actually no motivation for American companies to further reduce inference costs going forward. So I think they're going to just forget about this c…”
Consumer Demand for Top AI Models Has Increased Inference Costs 100X
“Instead of models getting cheaper, yes, maybe the running the same model got cheaper. But people trained much bigger models that are much more expensive to run now, and people expect to use the best model. So running inference in general, maybe a hundred X in …”
Morcos: AI inference costs will skyrocket, penalizing oversized models
“The inference costs are going to skyrocket with these models. And if you use a general purpose model, then you constrain to say, hey, this model knows about everything, but now only do this one thing. That model is going to have a ton of parameters that do not…”
Kurian: Long-term AI economics depend primarily on inference cost, not training
“First and foremost, in the long run, if AI really scales, the cost you really want to care about is inference cost, because that's what's integrated into serving, and any company that wants to recover the cost of training has to have a large scale inference fo…”
Tunguz: 100x cost difference between small and largest AI models
“I was just looking at the analysis between the smallest models, which are about four or two to four billion parameters, and the very largest models, which about four or fifty billion parameters, you have a hundred X difference in inference costs.”
Komoroske: Advertising cannot cover high inference costs for AI consumer startups
“And so if you're going to do a consumer startup, it can't be based on advertising. It's just too expensive. Advertising cannot clear the inference cost even with it with inference costs declining.”
Narayanan: Inference costs dominate training costs for popular AI models
“Over the lifetime of a model, when you have billions of people using it, the inference cost actually adds up, and for many of the popular models, that's the cost that dominates.”
Evans: High inference costs prevent free 100-million-user consumer AI apps
“Because at the moment you can't make a consumer app that's free and have a hundred million users because you can't afford the inference cost.”
Mensch: Overtraining models past Chinchilla limits lowers inference costs
“If you take into account the fact that your model Should also be efficient at inference time. You probably want to go far beyond the Cinchilla scaling low. So it means you want to overtrain the model. So train on more tokens than would be optimal for performan…”
Mensch: Pure scientific model performance ignores crucial runtime inference costs
“And if you want to push the performance, the pure performance of models, you don't care about inference because you, well, you are not going to use the model. You're just going to see whether they're good or not. And that's really for scientific purposes. But …”
Pesenti: Inference remains Facebook's largest machine learning cost
“Actually, To be clear, the most costly thing we do in ML is still inference cost, because when you put a piece of content within Facebook, it's running hundreds of different ML based algorithm, and they all run on machine parallel, and it's using a huge number…”