May 1, 2026 · 42m · no-priors

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Tuhin Srivastava · 30m spoken Sarah Guo · 5m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Baseten CEO Tuhin Srivastava joins No Priors to discuss the rapid expansion of the AI inference cloud ecosystem, explaining how custom model post-training, specialized runtimes, and distributed infrastructure enable AI-native applications to scale despite global GPU shortages.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.5% of the talking time here. How this is scored →

The hosts as informed peer 5.2 Guest teaching 5.3 Guest disagreement 1.2 The hosts pushing back 1.7
05100:0015:0030:001:55–4:34 · The hosts as informed peer 4/10 Why the Application Layer Will Survive Foundation Model Labs Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate.4:34–7:55 · The hosts as informed peer 5/10 AI-Native Startups Versus Traditional Enterprise Adoption Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements.7:55–13:06 · The hosts as informed peer 7/10 Open Source Model Evolution and Geopolitical Dynamics The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin.13:07–18:35 · The hosts as informed peer 4/10 Custom Workloads and the Link Between Post-Training and Inference Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes.18:35–24:10 · The hosts as informed peer 6/10 Navigating the Global AI Compute Supply Crunch Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier.24:10–28:29 · The hosts as informed peer 5/10 Inference Moats and NVIDIA's Hardware Dominance Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term.28:29–33:49 · The hosts as informed peer 5/10 Inference Cloud Architectures and Scale Edge Cases Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput.33:49–38:20 · The hosts as informed peer 6/10 Scaling Organizational Leadership and Operational Culture Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture.38:21–42:36 · The hosts as informed peer 5/10 Jevons Paradox and the Future of Ubiquitous AI Agents The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation.1:55–4:34 · Guest teaching 5/10 Why the Application Layer Will Survive Foundation Model Labs Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate.4:34–7:55 · Guest teaching 4/10 AI-Native Startups Versus Traditional Enterprise Adoption Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements.7:55–13:06 · Guest teaching 6/10 Open Source Model Evolution and Geopolitical Dynamics The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin.13:07–18:35 · Guest teaching 6/10 Custom Workloads and the Link Between Post-Training and Inference Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes.18:35–24:10 · Guest teaching 7/10 Navigating the Global AI Compute Supply Crunch Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier.24:10–28:29 · Guest teaching 6/10 Inference Moats and NVIDIA's Hardware Dominance Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term.28:29–33:49 · Guest teaching 6/10 Inference Cloud Architectures and Scale Edge Cases Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput.33:49–38:20 · Guest teaching 4/10 Scaling Organizational Leadership and Operational Culture Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture.38:21–42:36 · Guest teaching 4/10 Jevons Paradox and the Future of Ubiquitous AI Agents The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation.1:55–4:34 · Guest disagreement 1/10 Why the Application Layer Will Survive Foundation Model Labs Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate.4:34–7:55 · Guest disagreement 1/10 AI-Native Startups Versus Traditional Enterprise Adoption Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements.7:55–13:06 · Guest disagreement 2/10 Open Source Model Evolution and Geopolitical Dynamics The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin.13:07–18:35 · Guest disagreement 1/10 Custom Workloads and the Link Between Post-Training and Inference Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes.18:35–24:10 · Guest disagreement 1/10 Navigating the Global AI Compute Supply Crunch Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier.24:10–28:29 · Guest disagreement 2/10 Inference Moats and NVIDIA's Hardware Dominance Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term.28:29–33:49 · Guest disagreement 1/10 Inference Cloud Architectures and Scale Edge Cases Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput.33:49–38:20 · Guest disagreement 1/10 Scaling Organizational Leadership and Operational Culture Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture.38:21–42:36 · Guest disagreement 1/10 Jevons Paradox and the Future of Ubiquitous AI Agents The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation.1:55–4:34 · The hosts pushing back 2/10 Why the Application Layer Will Survive Foundation Model Labs Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate.4:34–7:55 · The hosts pushing back 2/10 AI-Native Startups Versus Traditional Enterprise Adoption Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements.7:55–13:06 · The hosts pushing back 3/10 Open Source Model Evolution and Geopolitical Dynamics The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin.13:07–18:35 · The hosts pushing back 1/10 Custom Workloads and the Link Between Post-Training and Inference Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes.18:35–24:10 · The hosts pushing back 2/10 Navigating the Global AI Compute Supply Crunch Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier.24:10–28:29 · The hosts pushing back 2/10 Inference Moats and NVIDIA's Hardware Dominance Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term.28:29–33:49 · The hosts pushing back 1/10 Inference Cloud Architectures and Scale Edge Cases Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput.33:49–38:20 · The hosts pushing back 1/10 Scaling Organizational Leadership and Operational Culture Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture.38:21–42:36 · The hosts pushing back 1/10 Jevons Paradox and the Future of Ubiquitous AI Agents The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 31.8% · guest 68.2%0:00 · the hosts 31.8% · guest 68.2%3:00 · the hosts 19.6% · guest 80.4%3:00 · the hosts 19.6% · guest 80.4%6:00 · the hosts 23% · guest 77%6:00 · the hosts 23% · guest 77%9:00 · the hosts 30.5% · guest 69.5%9:00 · the hosts 30.5% · guest 69.5%12:00 · the hosts 20.6% · guest 79.4%12:00 · the hosts 20.6% · guest 79.4%15:00 · the hosts 17.1% · guest 82.9%15:00 · the hosts 17.1% · guest 82.9%18:00 · the hosts 8.9% · guest 91.1%18:00 · the hosts 8.9% · guest 91.1%21:00 · the hosts 10.4% · guest 89.6%21:00 · the hosts 10.4% · guest 89.6%24:00 · the hosts 33.9% · guest 66.1%24:00 · the hosts 33.9% · guest 66.1%27:00 · the hosts 12.7% · guest 87.3%27:00 · the hosts 12.7% · guest 87.3%30:00 · the hosts 14.5% · guest 85.5%30:00 · the hosts 14.5% · guest 85.5%33:00 · the hosts 22.3% · guest 77.7%33:00 · the hosts 22.3% · guest 77.7%36:00 · the hosts 31.3% · guest 68.7%36:00 · the hosts 31.3% · guest 68.7%39:00 · the hosts 22.6% · guest 77.4%39:00 · the hosts 22.6% · guest 77.4%42:00 · the hosts 34.1% · guest 65.9%42:00 · the hosts 34.1% · guest 65.9%
Sharpest disagreement ▶ 10:22 Tuhin dismisses security hysteria around Chinese open models

Tuhin firmly rejects exaggerated fears about Chinese open-source models, arguing that network bounding protects data and that missing out on efficient intelligence poses a far greater risk to US innovation.

Hardest push from the hosts ▶ 26:02 Sarah presses on alternative silicon and multi-chip demand

Sarah directly questions Tuhin on whether customer demand for alternative chips like Grok will break NVIDIA's lock-in.

Biggest teaching moment ▶ 18:42 Tuhin details why 95% of production tokens are customized

Tuhin educates the hosts on real-world inference usage, explaining that practically no serious production customer runs vanilla open-source weights without custom compilation or post-training.

The host holds their own ▶ 11:48 Elad reframes open model economics as foreign state subsidies

Elad demonstrates high-level strategic insight by pointing out that state-backed open model development in China effectively functions as an indirect financial subsidy to US technology enterprises.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Why the Application Layer Will Survive Foundation Model Labs 4512 Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate.
AI-Native Startups Versus Traditional Enterprise Adoption 5412 Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements.
Open Source Model Evolution and Geopolitical Dynamics 7623 The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin.
Custom Workloads and the Link Between Post-Training and Inference 4611 Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes.
Navigating the Global AI Compute Supply Crunch 6712 Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier.
Inference Moats and NVIDIA's Hardware Dominance 5622 Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term.
Inference Cloud Architectures and Scale Edge Cases 5611 Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput.
Scaling Organizational Leadership and Operational Culture 6411 Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture.
Jevons Paradox and the Future of Ubiquitous AI Agents 5411 The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation.

Statements from this episode (21)

Assertion Not checkable as stated
Guo: Baseten Has Grown 30x Over the Past Year
“You guys have grown 30 X over the last year.”
Sarah Guo May 1, 2026 ▶ 0:39
Insight
Srivastava: Open-Source Baseline and Post-Training Enable In-House Inference
“The open source models have crossed some sort of chasm in terms of their baseline. Capability, and then I think RL techniques and post-training is for specialized models has become mainstream enough, and, you know, there's enough examples of its work, of it wo…”
Tuhin Srivastava May 1, 2026 ▶ 1:10
Insight
Srivastava: AI Defensibility Comes from Workflow Integration, Not Models Alone
“To the extent that that is encoded in a model, I think a lot of their business will be at risk, but to the extent that it is encoded in workflows that is where they will be able to develop mode.”
Tuhin Srivastava May 1, 2026 ▶ 2:42
Prediction Not checkable as stated
Srivastava: Frontier Labs Lack the User Signal to Displace Vertical Apps
“My argument would be here is that actually, you know, it's very, very hard for a frontier model company to go to either way at that, because they just don't have access to that user signal, and what will happen over time is the folks who have access to that us…”
Tuhin Srivastava May 1, 2026 ▶ 3:34
Assertion Not checkable as stated
Srivastava: AI-Native Startups Represent 99% of Total Inference Call Volume
“I think if you look by inference count, it'd be 99% the full.”
Tuhin Srivastava May 1, 2026 ▶ 5:06
Insight
Srivastava: AI companies choose models by capability before optimizing cost
“There are a, there's a subset of tasks, which I think is small today, where people really start to start with cost. But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered,…”
Tuhin Srivastava May 1, 2026 ▶ 8:31
Opinion
Gil: Chinese Open-Source Subsidies Indirectly Subsidize US Enterprise AI Adopters
“At least for now, it looks like effectively the Chinese government is subsidizing at least a large subset of these models, and that subsidy or surplus is effectively just being passed on to US enterprises We're adopting these models. In other words, it's a way…”
Elad Gil May 1, 2026 ▶ 11:37
Assertion Not checkable as stated
Srivastava: 95% of Tokens on Baseten Run on Custom Modified Models
“I'd say, 95% of the tokens today are on the first business, and almost all of them there's probably a, yeah, for almost all of them, the customer is making some modifications to the model with their own data specialized for the use case, and I think what's eve…”
Tuhin Srivastava May 1, 2026 ▶ 13:52
Insight
Srivastava: Startups should not do post-training before achieving product-market fit
“Hey, go find, go prove to yourself with the best in class model that you have something worth optimizing. And I think, you know, A lot of, you know, if a customer comes to us, was that meme, which was like, it was like two years ago, it feels like there's no G…”
Tuhin Srivastava May 1, 2026 ▶ 17:43
Disclosure
Srivastava: Baseten Runs 90 Clusters Across 18 Clouds at Mid-90s Utilization
“We run them in, like, uncomfortably high utilization. You know, we, when I'm saying we're like mid-nineties utilization most of the time there is, we have made, we have, we sit in 18 different clouds now. We have 90 clusters around the world across 18 differen…”
Tuhin Srivastava May 1, 2026 ▶ 19:07
Opinion
Srivastava: Only 3 or 4 Global Cloud Providers Belong in Gold Tier
“There's probably, like, a dozen good, like, clouds, and I'd probably, like, put, like, three or four of them in, like, the gold tier and I think that just means that, like, supply, like, not only are we supply crunched, we're supplier and operationally crunche…”
Tuhin Srivastava May 1, 2026 ▶ 21:14
Assertion Not checkable as stated
Srivastava: 1,024 B200s Require 3-5 Year Contracts and 20-30% Prepay
“So if you wanted a thou, a thousand, 1024 B 200 which is, you know from a good cloud right now, you're not getting that less than a three to five year contract right now with a, probably a 20 to 30% TC, TCV prepay.”
Tuhin Srivastava May 1, 2026 ▶ 22:45
Insight
Srivastava: Raw GPU hosting is an unsticky commodity compared to inference software
“GPUs as a service is not sticky. I think that's been seen. Like, customers generally just see that as commodity. Inference with the software layer included is incredibly sticky.”
Tuhin Srivastava May 1, 2026 ▶ 24:46
Assertion Not checkable as stated
Srivastava: Baseten Maintains 400% Annual NDR and Zero Top-30 Customer Churn
“None of our top 30 customers have ever churned. You know, we're talking Like, 400% annual NDR around our business, and so it's like very, it's very, very sticky.”
Tuhin Srivastava May 1, 2026 ▶ 24:59
Prediction Not checkable as stated
Srivastava: Inference-Specific and Decode-Specific AI Chips Will Emerge
“Yeah, and I think there will be inference-specific chips. I think you have, like, decode-specific chips, I think.”
Tuhin Srivastava May 1, 2026 ▶ 26:47
Opinion
Srivastava: KV-cache-aware routing is already becoming old technology
“Even stuff like KV cache away routing and, you know, that stuff's a bit old now”
Tuhin Srivastava May 1, 2026 ▶ 29:04
Assertion Not checkable as stated
Srivastava: Disentangling Pre-fill and Decode Yields Massive Performance Gains
“Somewhat disentangling pre-fill and decode and starting to treat them as separate problems. I think that's, you know, something we are very focused on, and we're seeing massive gains there.”
Tuhin Srivastava May 1, 2026 ▶ 29:13
Prediction Not checkable as stated
Srivastava: Global Compute Supply Will Fall Short of LLM Demand for Decade
“I think, like, there's no world in which there's enough compute to, you know, get the amount of value that we want to get out of our limbs in the next five to 10 years.”
Tuhin Srivastava May 1, 2026 ▶ 33:36
Insight
Srivastava: Founder micromanagement indicates having the wrong team
“If you feel like you are micromanaging, if you feel like you need, if you feel like, you know, you have to be involved in everything, I think that's a bit of a cop-out as a founder, because you're just like, I just need to be involved in everything. It's like,…”
Tuhin Srivastava May 1, 2026 ▶ 35:15
Insight
Srivastava: Even After AGI Is Achieved, Inference Is All That Remains
“Even if there's AGI, all that's left is inference.”
Tuhin Srivastava May 1, 2026 ▶ 40:21
Prediction Not checkable as stated
Srivastava: AI will lead to more software rather than fewer software engineers
“There's all this stuff about there being less software engineers, and I think we just build more software. I think we just build a ton more software”
Tuhin Srivastava May 1, 2026 ▶ 41:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.