Kedrosky: GPU failure rates are much higher in training than inference
“The failure rates of GPUs used so intensively for training purposes are much higher than inference specific usage.”
Atallah: Rare Disease Research Has Been Constrained by an Inference Bottleneck
“One is rare disease research, which I think is one of those things that has been intelligence bottleneck or really just the inference bottleneck. Like it involves like trying out lots of ideas and seeing if they work.”
Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Tae Kim: Morgan Stanley estimates AI inference yields 60-80% margins
“Like Morgan Stanley says, if you do inference, it's 60 to 80% profit margins, right?”
Feldman: AI inference is bottlenecked by data movement, causing GPU slowness
“In inference in AI, it's the exact opposite. You move a huge amount of data, all the weights, from memory to compute, and you need one calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data…”
Liang: AI inference chip deployments will dwarf training by orders of magnitude
“Because at scale, the number of chips deployed for inferencing will be orders of magnitude greater than whatever you're doing for training.”
Liang: Demand will converge on the fastest, most accurate large models
“And so, as the cost of delivering fast goes down, you're going to see most people switch over to the fastest. And this is why I feel like, you know, the premium inference, which is large models, which equals the most accurate. The most accurate models and fast…”
Feldman: Cerebras AI inference is 20x faster than competitors
“And we're the fastest, not by a little bit, but by 20 x.”
Wachen: Software COGS per incremental user will become high due to inference
“Like, the cost structure of every software company, the COGS is not going to be, like, zero anymore for an incremental user. It's going to be, like, quite high, and it's going to be a function of inference.”
Wachen: AI inference will become the biggest market in the world
“It feels like we are on a, like, decade march for inference to become, you know, the biggest market in the world.”
Wachen: AI inference will eventually make up the majority of global GDP
“I firmly believe we are on a global march of inference becoming majority of global GDP. And it may take more than 10 years, but it's going to happen.”
Ambati: AI inference market will reach $1.3T-$1.4T by 2030
“The inference market is one of the largest or so the fastest growing CAGR. So what we are seeing is about 94% CAGR. And then, you know, today, I think around 20, 26, we are looking at an eighty billion dollar sort of market. But soon by 2030 it's gonna be a 1.…”
Feldman predicts there will be zero market for slow AI inference
“Why do we believe that inference will be any different? There'll be zero marking for slowing them.”
Srivastava: AI-Native Startups Represent 99% of Total Inference Call Volume
“I think if you look by inference count, it'd be 99% the full.”
Hassabis doubts inference will ever be free due to Jevons' paradox
“Yeah, I'm not sure inference will ever be essentially free. I mean, there's sort of Jevon's paradox and other things about like, I think we'll just end up using All of us will end up using whatever we can get our hands on, and you could imagine millions of age…”
Hassabis: Chip manufacturing will bottleneck AI inference for decades
“Certainly the energy, if we solve fusion or, you know, superconductors or, you know, optimal batteries or some set of those things, which I think we will do with material science, energy costs will be essentially zero, but there'll still be the physical creati…”
Vahdat: AI inference does not require gigawatt-scale data center capacity
“You don't strictly need a gigawatt of capacity to be able to do useful work. You probably don't even need a hundred megawatts of capacity. It gets a little bit more interesting than that because of, let's say, co-located compute and storage and networking and …”
Huang: AI inference pipeline is the most complicated computing problem today
“The processing pipeline of inference is extremely complicated. And in fact, it is the most complicated computing problem today.”
Tiwari: AI inference compute will shift toward decentralized clusters
“What you're seeing with inference is in many use cases, as this becomes more ubiquitous, you're going to have more and more decentralized inference clusters.”
Guo: AI inference clouds are growing 1,000x in compute consumption
“I think to the point of like real data, the inference clouds are growing a thousand X in terms of consumption, right? And then they're getting more efficient. So revenue grows at some lower rate than that, but it's wild.”
Coogan: AI inference will be load-balanced across diverse chip architectures
“It does feel like we're going to enter a world where inference is load balanced across a variety of semiconductor stacks, and so, For really fast things, you might be going to a Grok or a Cerebris, or, you know, you might, for more basic stuff, you might be go…”
Amon: AI inference will eventually surpass training in data centers
“Eventually inference is going to take over training because Just think about that for a second. If you're a company spending billions of dollars building a data center for training, you expect to get a return on that investment. So when you start putting AI in…”
Amon: Qualcomm is building post-GPU chips for data center inference
“So we're building what we believe is post-GPU. When you started to do inference and you need the dedicated engines, We're building that. I actually believe that the NVIDIA acquisition of Croc validates that you different engines for different things, and I thi…”
Lovinsky: Continuous background AI inference won't reach general knowledge work by 2026
“Whether it happens widely in knowledge work, By 20, 26, I think is, you know, I'll take a bet with you on that. Certainly for coding. I think in a lot of places we're already there, right? I mean, you have like the Ralph Wiggum stuff that, that popped up over,…”
Venturo: No difference between training and inference infrastructure for modern AI
“So there's really no difference between training infrastructure we deployed to build those capabilities and what our customers are ultimately using to serve them.”
Venturo: CoreWeave compute demand has shifted to roughly 50-50 training and inference
“Six months ago, I would have said it was two-thirds training and one-third inference. It's probably close to fifty-fifty now.”
AI inference spending must surge to justify current adoption optimism
“If the optimism about AI is to be justified, you're going to have to see inference costs go way up because that will be an indicator that adoption has gone up in a fairly significant way, both among individual users, but also among companies and enterprise use…”
AI training compute cycles cause massive power fluctuations unlike inference workloads
“And some of the people in the energy space may know that there are massive energy fluctuations or power fluctuations we will see in data center usage when the GPUs go from this computational intensive phase where you're learning the model weights to this commu…”
Apple is the most likely company to move AI inference on-device
“It's not hard to picture that, like, if somebody's gonna move a lot of this inference on device, it's gonna be Apple.”
Software agents, rather than human queries, will drive most AI inference workloads
“I think increasingly most of the inference workload will come from other software agents.”
Randle: Investors should follow momentum when AI inference demand is unprecedented
“When you have demand like this, like you have for the initial GC, like the initial hyperscaler clouds. And I think we're seeing an even greater cohorted demand curve for AI inference. Sometimes you just got to shut your mind up and invest with the momentum.”
Corbitt: AI inference could be 10x larger if reliability issues are solved
“I think that there is today, like. 10 times as much AI inference that could exist than is existing right now, just Purely with projects that are like sitting in the proof of concept stage and have not been deployed because there's like huge bucket of those. An…”
Feldman: AI inference growth compounds across users, frequency, and compute per query
“The greater growth of inference is the number of people who use it, Times the frequency of use. Times the amount of compute needed per use. Right? It is three different variables multiplied by each other. The problem is they're all growing fast.”
Ross: Lowering AI chip prices 50% leads customers to buy double
“If we lower what we charge 50%, people are gonna buy twice as much. They're spending as much as they're making because whatever they spend increases the quality of the output.”
Jeff Lawson: AI inference is not complex enough to sustain premium margins
“I believe inference itself is not such a hard algorithmic solve that, you know, you need to pay someone else to do it for you, but clearly training a model is”
PyTorch democratized model training, but AI inference remains undemocratized
“And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. Inference…”
AI training scales with research teams; inference scales with customer base
“Training scales the size of your research team. Inference scales the size of your customer base.”
Sharp: AI inference, not training, will drive future data center growth
“I would say that there's been a lot of growth in that, but we see that kind of leveling out where we really see the consumption of AI or inference that's driving that regional specific growth going forward. And I think that's where it's more embedded. And a lo…”
Sharp: AI inference deployments require 5-megawatt blocks, unlike training arrays
“What we're really seeing is inference. Can come in, in like, five-ish megawatt blocks, and you can solve for it a bit differently. Now, the densification is still there, and then the private AI pieces, there's hot spots where it can be, you know, a couple of m…”
$20/month local IDE pricing limits AI inference quality per task
“The cost when you are local first and your typical consumer is on a free plan or a Like, 20 dollar a month paid plan limits the amount of high quality inference you can do, and the scale or volume of inference you can do per, like, outcome.”
Consumer tech products will increasingly tier pricing based on AI inference volume
“It's very likely that some consumers are going to want tons and tons of inference. And because that's a marginal cost, you're probably going to have to pay somehow for that as a consumer. So I think you're going to see more tiering of consumer products based o…”
Overall demand for intelligence will outpace the falling costs of AI compute
“I think from a financial point of view when the cost of something drops, the demand usually increases more than the cost, than the drop. And I think that's, Bound to happen with intelligence.”
Conrad: Custom chips make sense for inference, not training experimentation
“It only works if you really know which chip you're going to do. If you don't, then it's a little harder. So it makes, in my head, it makes more sense for inference where you've already established it, but for training there's so much, like, experimentation.”
Ramaswamy: AI inference will become significantly faster and cheaper in 2025
“The world of inference, not foundation models, the world of inference is going to go through so much change this year, both in terms of GPU availability, which is easing up quite a bit, But also in terms of people like Grok, I think the one with the K or the Q…”
Boland: NVIDIA Is Re-Architecting GPUs to Optimize for Inference
“They are making a bunch of architectural changes to GPUs to make them better and better inference.”
Feldman: Nvidia Has Zero CUDA Software Lock-In for AI Inference
“In inference, it's not real at all. There's no CUDA lock-in in inference. None. Well, you can move from OpenAI on an NVIDIA GPU to Cerebrus to Fireworks Service on something else to Together to perplexity with 10 keystrokes. I mean, anybody who actually uses A…”
Microsoft AI screened 32 million battery electrolyte candidates down to 100
“What you're referring to is a project where AI was leveraged by our research teams to screen over thirty-two million candidates for a new electrolyte for advanced batteries. Thanks to AI inference, you know, the R&D scientists were able to go from that number …”
Wolfe: 30% to 50% of AI inference will occur on-device
“I think that you're going to end up doing a lot of inference on device. Meaning, instead of going, ah, to the cloud and typing a query, like, 30 to 50% of your inference may be on an Apple or an Android device.”
Morin: Interconnect dependency is the core difference between training and inference
“In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of…”
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Ross: The AI industry mistakenly believed training was costlier than inference
“When we started, the first misconception, which people don't hold anymore, is that training was more expensive than inference.”
Ulrich: Proprietary data and inference will differentiate future AI applications
“I do think that the use of data becomes a more critical differentiator. Inference becomes a more critical differentiator.”
Jensen Huang: AI inference demand will scale up by one billion times
“It's gonna go up a billion times.”
Huang: AI inference compute is about to increase by a billion times
“It's about to go up by a billion times.”
Zhang: Next 10x AI inference gain requires multi-layer co-optimization
“If you only push on one direction, you are going to reach diminution return really, really quickly. Yeah, there's only that much you can do on the system side, only that much you can do on the algorithm side. And since the only big thing that's going to happen…”
Doshi: AI inference DevOps is similar to scaling high-volume API servers
“I don't find, I find the DevOps for inference to be relatively easy. It doesn't feel that different than, you know, I think we had thousands and thousands of servers at Mixpanel just for dealing with the API had such huge quantities of volume that I didn't fin…”
Baszucki says Roblox will provide free AI inference to its creators.
“That the more we can run inference jobs on these, we can run super high volume inference at high quality at low cost and make this, you know, just freely available so the creators don't worry about it.”