Trojanowski: Cheaper inference is replacing graphs and embeddings for agent ontologies
“As models are getting cheaper and cheaper, more and more of that actually can just be done using inference instead of using determined, like things like graphs or things like embeddings.”
Inference optimization is only solved once researchers report mere 1% speedups
“Like, you'll, you'll know that influence is pretty much solved when researchers start publishing about how they got one percent faster at something.”
Optimal inference parallelism cannot be mathematically calculated; it must be auto-tuned
“And with training, it's more of like a math, like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning, like GPU kernel auto-tuning... You shadow the same traffic, like real traffic, and you just see which co…”
Angelopoulos: Model labs must enter application layer to avoid commoditization
“If, like, inference is going to commoditize, then, of course, the next best thing is for the model providers to be moving up the application layer in order to own more of the application stack so that they ensure that they're not commoditized and they're getti…”
Altman: Massive Inference Revenue Will Easily Fund Frontier Model Training
“We will have so much usage of our models that we do not need to be a gigantically high margin business to be able to afford model training. Like, so much of our future compute plans Will be used to sell inference to customers that even if we can enjoy a modest…”
Inference likely accounts for the majority of OpenAI's total compute
“Inference is a big, perhaps even the majority on compute.”
Katti: Modern AI model training consists heavily of inference workloads
“We don't like to make a distinction between Training and infants, because a lot of training is now infants. So when we train a new model, we are generating synthetic data, for example. That's inference. When we train a new model, we are doing post-train, and t…”
Evans: AI Model Inference Alone Generates 40% to 50% Gross Margins
“And meanwhile, we sort of know that you have positive gross margins on inference alone of sort of 40, 50%, but you've got the cost of building the next model, and you don't know where the cost The cost will move, or the cost of the next model will move, and yo…”
Wachen: Over $1B in AI Compute Revenue Generated Every Day
“Every day there's over a billion dollars of revenue in this category, and a lot of it's inference.”
Randle: Monetizing inference margins unlocks unprecedented AI revenue scaling
“When you hear about all these companies, and you hear about all of these businesses that are going not like, you know, one to three to nine to 20, like they used to, but are going one to 20 to a hundred, or one to 30 to 300, all of that can be traced back to t…”
Prince: Cloudflare often avoids paying for space or power
“And because of where we deploy systems, we often are in places where
We don't have to pay for the space or the power, which allows us to then pass those savings on to customers.
So it can be significantly more cost effective to use inference with us than it …”
Ambati: AI infrastructure spending is now driven primarily by inference
“The whole big bet of open AI, you know, trying to build these target and other You know, big data centers is not that they just want to train on this, because that cost is a one-time cost, right? So you train ones, you sort of run inference for a lifetime. So …”
Chernin: Bare metal AI compute has only a dozen global customers
“On bare metal level, you have maybe a dozen of the customers in the world that you can work with. On managed infrastructure, there are hundreds. On inference, there are thousands. On agentic, there will be tens of thousands of new developers that build it, rig…”
Baker: Anthropic has strongly positive gross margins on inference
“Anthropic I'm sure is at very positive gross margins on inference today.”
Baker: Disaggregated inference will extend GPU lifespans to 10-15 years
“And the disaggregation of inference means that I think these GPUs are going to have 10 or 15 year lives.”
Rao: More efficient inference directly increases reinforcement learning efficiency
“If we're doing reinforcement learning on the model, it's basically inference within a sandbox with a reward function, right? And so if the model's better at more efficient inference, that RL is more efficient as well.”
Srivastava: Open-Source Baseline and Post-Training Enable In-House Inference
“The open source models have crossed some sort of chasm in terms of their baseline. Capability, and then I think RL techniques and post-training is for specialized models has become mainstream enough, and, you know, there's enough examples of its work, of it wo…”
Srivastava: Raw GPU hosting is an unsticky commodity compared to inference software
“GPUs as a service is not sticky. I think that's been seen. Like, customers generally just see that as commodity. Inference with the software layer included is incredibly sticky.”
Srivastava: Even After AGI Is Achieved, Inference Is All That Remains
“Even if there's AGI, all that's left is inference.”
Michael Intrator: AI inference represents the monetization of AI investment
“I always think of inference as the monetization of the investment in artificial intelligence.”
Lemkin: Global AI inference volume will grow 1,000x in five years
“Yeah, I mean, I think it's got to be three orders of magnitude more inference we run in the next five years.”
Pope: MatX will sell AI inference chips first due to lower risk
“Our product is both training and inference, but I think the first sales will be an inference. That's mostly just a market effect where It's easier to buy, like, it's not as big of a risk to go to buy an inference cluster than as a training cluster.”
Stebbings: Data centers are today's most under-invested technology category
“Data center is the one that's the most under-invested categories today. When you look at inference needing to run for 24 hours a day for most of the knowledge worker population, and it running for like one percent of knowledge worker population today, I'm like…”
Lemkin: AI inference costs are the new sales and marketing expense
“For simplistic folks, for founders, I say inference is the new sales and marketing.”
Feldman: GPUs face a major inference bottleneck due to slow memory speed
“The GPU, for example, has a lot of capacity of memory, but it's really slow. All right, and that's a huge bottleneck in inference. It's why they can't be fast.”
Fu predicts increasing hardware diversity, particularly for AI model inference
“I'm sure NVIDIA will still do great and still grow beyond their five trillion dollar company or whatever it is at the time of recording. But I think you're going to see a lot more diversity especially around, I think inference of the model.”
Angelopoulos: LMArena's top expenses are free-tier inference, hiring, and SF office
“Primarily inference that funds the free usage of the platform and then also hiring, of course, headcount. We have an office, you know. That's an SF.”
Catanzaro: Continual Learning Requires Making Stateless Inference Stateful
“If you must update weights, then, like, you know, weights become stateful, and today, like, inference is not stateful.”
Wolfe: 50% of AI Inference Will Run Locally On-Device
“I'm of the belief that 50% of your inference and your queries will be on device. They will not be on device going to the cloud. They will be on device. Locally hosted on models that are either etched into the silicon or using flash memory.”
Epoch AI: Most AI compute is spent on product inference, not training
“It does seem as if most compute gets spent on inference that companies don't so far regret like using to offer their products.”
Feldman: Vastly more people do AI inference than AI training
“To move people off GPUs in inference, and the number of people doing inference is vastly higher than the number of people doing training.”
Ross: AI Training and Inference Form a Virtuous Hardware Demand Cycle
“The more inference you have, as mentioned before, the more you need to train the model to optimize for the inference. And the more training you have the more inference you want to deploy to optimize for the cost of that training, to amortize the cost of the tr…”
Gerstner: Over 40% of NVIDIA's revenue comes from AI inference workloads
“Over 40% of your revenue today is inference.”
Morcos: AI total cost of ownership is dominated by inference
“When you think about the total cost of ownership of these models, it's gonna be very, very heavily weighted towards inference. It's all inference.”
Pedregal: Granola's inference costs will stay flat or rise as queries expand
“What I do expect is I expect the cost of inference to stay the same or go up as we allow users to do much more complicated queries over much larger data sets.”
Prince: In-network edge inference will handle models too large for end devices
“We believe that a lot of inference is going to happen on your end device, but there will always be some model which is too big or too resource intensive. And in that case, the next best place to run it is going to be on at the inside the network at the edge.”
Rizwan: AI coding tools should not monetize through inference markups
“Our thesis is inference is not the business. We, yeah, we want to give the end user total transparency into price, into, which I think is like incredibly important to, you know, even get comfortable with the idea of spending as much money as you do. I think th…”
Kantrowitz notes Nvidia CEO says AI inference needs 100x more compute
“And I mean, we have Jensen Wang, the CEO of Nvidia saying inference is going to take a hundred times more compute than traditional LLM inference.”
Feldman: GPUs operate at only 5% to 7% utilization during inference
“In a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93% wasted.”
Feldman: GPU off-chip memory architecture can be beaten in inference
“The fundamental architecture of the GPU with off-chip memory is not great for inference. Now, they will continue to do well in inference, but it can be beaten, and I think they know it.”
Morin: Nvidia will remain dominant in AI inference due to availability
“The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Availability that is right. So I think Nvidia is used to stay at least if not for the H-one hundre…”
Morin: Doubling GPUs in AI inference yields only 10% performance gain
“If you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount o…”
Kimber: AI inference will ultimately draw far more power than training
“I think now what we're seeing is that inference in the aggregate is actually probably a much larger draw than the training in the long run.”
Coogan: DeepSeek definitely succeeded in making AI inference cheaper
“They clearly did a bunch of optimizations to make the code run faster. And there's just no question that inference is cheaper. Like that is true.”
Ross: AI inference will account for 95% of total compute demand
“I think, 95%.”
Sivulka: AI shift from training to inference will destabilize Nvidia's dominance
“The shift away from training to inference as a fundamental, like, almost macro shift in how people deploy AI. I actually think that will destabilize slightly the dominance of NVIDIA chips.”
Bernhardsson: Modal Is Expanding into Bursty Experimental AI Training
“Traditionally, most of modal has always been inference. Like that's been our main use case, but we're really interested also in training. So in particular, like probably focused more on these like shorter, like very bursty sort of experimental training runs, n…”
Patel: NVIDIA's inference moat relies on hardware rather than software
“NVIDIA's moat in, in inference is actually A lot smaller on software but it's a lot bigger on, hey, they just have the best hardware.”
Patel: Modern AI training requires more inference compute than weight updates
“In fact, there's more inference in training than there is updating the model weights, because you have to generate hundreds of possibilities And then, oh, you only train on a couple of them, right?”
AI inference matters more than training because it scales with global population
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
Gerstner: Microsoft said almost all $10B AI revenue is inference
“Microsoft has said of its ten billion dollars in AI revenue, it's almost all inference.”
Gerstner: Nvidia stated half its revenue comes from inference
“We know Nvidia has said half of their revenue is inference”
Gerstner: 40% of NVIDIA's revenue already comes from inference
“40% of their revenues are already inference.”
Gerstner: AI inference costs dropped 90% over the past year
“What we know is that the cost of inference has fallen by 90% over the course of last year.”
Huang: AI Inference and Post-Training Are Now Just as Hard as Pre-Training
“People used to think that pre-training was hard, and inference was easy. Now everything is hard.”
Huang: NVIDIA's moat in inference will be greater than in training
“And I'm sure I said it would be greater.”
AI inference will become a core cloud computing primitive like storage
“As you move forward, generative AI honestly becomes one of the compute building blocks that you think about. You're going to need storage, you need compute, you need databases, you need inference, if you will, for your application, largely. And I think that's …”
Sarah Guo says massive frontier AI models are impossible to serve commercially
“Over time, applications are going to want efficient inference, and, like, really large models are impossible today to serve for the vast majority of use cases from a cost and speed perspective”
Chamath predicts AI inference market will be 100x larger than training
“AI is really two markets, training and inference is going to be a hundred times bigger than training.”
Chamath asserts Nvidia hardware is miscast for AI inference
“NVIDIA is really good at training and very miscast at inference.”