Baseten CEO Tuhin Srivastava discusses how AI inference infrastructure requirements have shifted toward higher-end NVIDIA GPUs since late 2022.
Opinion
Srivastava: The ease of running CUDA workloads on AMD is significantly overstated
“I've, I personally think that it's pretty overstated how easy it is to run something that looks like CUDA or CUDA in some form on an AMD chip seems, seems like a challenge to me.”
Insight
Srivastava: AI Defensibility Comes from Workflow Integration, Not Models Alone
“To the extent that that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows that is where they will be able to develop mode.”
Prediction Not checkable as stated
Srivastava: Frontier Labs Lack the User Signal to Displace Vertical Apps
“My argument would be here is that actually, you know, it's very, very hard for a frontier model company to go to either way at that, because they just don't have access to that user signal, and what will happen over time is the folks who have access to that us…”
Assertion Not checkable as stated
Srivastava: AI-Native Startups Represent 99% of Total Inference Call Volume
“I think if you look by inference count, it'd be 99% the full.”
Insight
Srivastava: Startups should not do post-training before achieving product-market fit
“Hey, go find, go prove to yourself with the best in class model that you have something worth optimizing. And I think, you know, A lot of, you know, if a customer comes to us, was that meme, which was like, it was like two years ago, it feels like there's no G…”
Prediction Not checkable as stated
Srivastava: Global Compute Supply Will Fall Short of LLM Demand for Decade
“I think, like, there's no world in which there's enough compute to, you know, get the amount of value that we want to get out of our limbs in the next five to 10 years.”