Everything Tuhin Srivastava said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Srivastava: The ease of running CUDA workloads on AMD is significantly overstated
“I've, I personally think that it's pretty overstated how easy it is to run something that looks like CUDA or CUDA in some form on an AMD chip seems, seems like a challenge to me.”
Srivastava: AI Defensibility Comes from Workflow Integration, Not Models Alone
“To the extent that that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows that is where they will be able to develop mode.”
Srivastava: Frontier Labs Lack the User Signal to Displace Vertical Apps
“My argument would be here is that actually, you know, it's very, very hard for a frontier model company to go to either way at that, because they just don't have access to that user signal, and what will happen over time is the folks who have access to that us…”
Srivastava: AI-Native Startups Represent 99% of Total Inference Call Volume
“I think if you look by inference count, it'd be 99% the full.”
Srivastava: Startups should not do post-training before achieving product-market fit
“Hey, go find, go prove to yourself with the best in class model that you have something worth optimizing. And I think, you know, A lot of, you know, if a customer comes to us, was that meme, which was like, it was like two years ago, it feels like there's no G…”
Srivastava: Global Compute Supply Will Fall Short of LLM Demand for Decade
“I think, like, there's no world in which there's enough compute to, you know, get the amount of value that we want to get out of our limbs in the next five to 10 years.”
Srivastava: AI will lead to more software rather than fewer software engineers
“There's all this stuff about there being less software engineers, and I think we just build more software. I think we just build a ton more software”
Srivastava: 95% of Tokens on Baseten Run on Custom Modified Models
“I'd say, 95% of the tokens today are on the first business, and almost all of them there's probably a, yeah, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case, and I think what's eve…”
Srivastava: Only 3 or 4 Global Cloud Providers Belong in Gold Tier
“There's probably, like, a dozen good, like, clouds, and I'd probably, like, put, like, three or four of them in, like, the gold tier and I think that just means that, like, supply, like, not only are we supply crunched, we're supplier and operationally crunche…”
Srivastava: 1,024 B200s Require 3-5 Year Contracts and 20-30% Prepay
“So if you wanted a thou, a thousand, 1024 B 200 which is, you know from a good cloud right now, you're not getting that less than a three to five year contract right now with a, probably a 20 to 30% TC, TCV prepay.”
Srivastava: Founder micromanagement indicates having the wrong team
“If you feel like you are micromanaging, if you feel like you need, if you feel like, you know, you have to be involved in everything, I think that's a bit of a cop-out as a founder, because you're just like, I just need to be involved in everything. It's like,…”
Srivastava: Even After AGI Is Achieved, Inference Is All That Remains
“Even if there's AGI, all that's left is inference.”
Srivastava: LLM inference performance will commoditize and increasingly run locally
“We think over time, It will get somewhat commoditized, the performance, especially, especially for language models, to be honest. I think, you know, more and more of that stuff should run locally to some degree, I think”
Srivastava: Shared model endpoints do not work for enterprise AI workloads
“Shared endpoints don't really work for large companies, it's not going to work. They're not definitely not going to work for enterprises.”
Srivastava: Efficient AI businesses can achieve healthy margins despite heavy compute spend
“And, you know, what, what's crazy about it though, is I think the most efficient businesses through markups and through software optimization can actually direct, like drive pretty healthy margins. And still have these really aggressive consumption contract.”
Srivastava: Hedge funds and financial services lag in generative AI adoption
“I don't actually think they're on the cusp of it as much as, you know, you'd think, like, I, you know, if you think about a lot of this, the big data stuff, like, 10 years ago you know, the hedge funds were all over that, right? They're like, hey, there's alph…”
Srivastava: Baseten is avoiding non-NVIDIA hardware investments in the short term
“I think short term you know, it's really hard for me to see how we make investments beyond NVIDIA, especially when there's a customer, where there's like customer A crunch on the other side from customers who are like, hey, we need this now.”
Srivastava: Open-Source Baseline and Post-Training Enable In-House Inference
“The open source models have crossed some sort of chasm in terms of their baseline. Capability, and then I think RL techniques and post-training is for specialized models has become mainstream enough, and, you know, there's enough examples of its work, of it wo…”
Srivastava: AI companies choose models by capability before optimizing cost
“There are a, there's a subset of tasks, which I think is small today, where people really start to start with cost. But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered,…”
Srivastava: Baseten Runs 90 Clusters Across 18 Clouds at Mid-90s Utilization
“We run them in, like, uncomfortably high utilization. You know, we, when I'm saying we're like mid-nineties utilization most of the time there is, we have made, we have, we sit in 18 different clouds now. We have 90 clusters around the world across 18 differen…”
Srivastava: Raw GPU hosting is an unsticky commodity compared to inference software
“GPUs as a service is not sticky. I think that's been seen. Like, customers generally just see that as commodity. Inference with the software layer included is incredibly sticky.”
Srivastava: Baseten Maintains 400% Annual NDR and Zero Top-30 Customer Churn
“None of our top 30 customers have ever churned. You know, we're talking Like, 400% annual NDR around our business, and so it's like very, it's very, very sticky.”
Srivastava: Inference-Specific and Decode-Specific AI Chips Will Emerge
“Yeah, and I think there will be inference-specific chips. I think you have, like, decode-specific chips, I think.”
Srivastava: KV-cache-aware routing is already becoming old technology
“Even stuff like KV cache away routing and, you know, that stuff's a bit old now”