May 1, 2026 · 42m · no-priors
Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Baseten CEO Tuhin Srivastava joins No Priors to discuss the rapid expansion of the AI inference cloud ecosystem, explaining how custom model post-training, specialized runtimes, and distributed infrastructure enable AI-native applications to scale despite global GPU shortages.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.5% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Tuhin firmly rejects exaggerated fears about Chinese open-source models, arguing that network bounding protects data and that missing out on efficient intelligence poses a far greater risk to US innovation.
Hardest push from the hosts ▶ 26:02 Sarah presses on alternative silicon and multi-chip demandSarah directly questions Tuhin on whether customer demand for alternative chips like Grok will break NVIDIA's lock-in.
Biggest teaching moment ▶ 18:42 Tuhin details why 95% of production tokens are customizedTuhin educates the hosts on real-world inference usage, explaining that practically no serious production customer runs vanilla open-source weights without custom compilation or post-training.
The host holds their own ▶ 11:48 Elad reframes open model economics as foreign state subsidiesElad demonstrates high-level strategic insight by pointing out that state-backed open model development in China effectively functions as an indirect financial subsidy to US technology enterprises.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Why the Application Layer Will Survive Foundation Model Labs | 4 | 5 | 1 | 2 | Sarah sets up the existential debate over whether the application layer can survive frontier model labs. Tuhin clearly articulates why proprietary user signal and deeply integrated clinician workflows (such as Abridge) create sustainable moats that foundation labs cannot easily replicate. | |
| AI-Native Startups Versus Traditional Enterprise Adoption | 5 | 4 | 1 | 2 | Elad asks about the current split between AI-native startups and traditional enterprises adopting AI. Tuhin humorously notes it is still 99% AI natives by inference volume, explaining that serving frontier AI apps indirectly prepares them for enterprise requirements. | |
| Open Source Model Evolution and Geopolitical Dynamics | 7 | 6 | 2 | 3 | The conversation shifts to open-source Chinese models and geopolitical concerns. Elad demonstrates sharp economic expertise by framing Chinese model subsidies as an indirect subsidy to U.S. enterprise efficiency, while Tuhin emphasizes the operational necessity of running the best available frontier open models regardless of origin. | |
| Custom Workloads and the Link Between Post-Training and Inference | 4 | 6 | 1 | 1 | Tuhin reveals that 95% of inference on Baseten comes from customized models rather than off-the-shelf weights. He educates the hosts on the tight architectural feedback loop linking post-training evaluation loops, quantization, and specialized inference runtimes. | |
| Navigating the Global AI Compute Supply Crunch | 6 | 7 | 1 | 2 | Tuhin details the brutal realities of the AI compute supply crunch, including multi-year commitments, prepays, and operational deficits among new data center providers. Elad actively engages on capital structure implications and whether this forces infrastructure startups to IPO earlier. | |
| Inference Moats and NVIDIA's Hardware Dominance | 5 | 6 | 2 | 2 | Sarah pushes on whether custom ASICs or alternative chips will challenge NVIDIA. Tuhin explains why NVIDIA's mature CUDA developer ecosystem and unmatched supply chain make them virtually unassailable in the near term. | |
| Inference Cloud Architectures and Scale Edge Cases | 5 | 6 | 1 | 1 | Sarah brings up hitting scale ceilings on major hyperscalers, and Tuhin shares technical edge cases like kernel panics from logging daemons and the relative immaturity of LLM inference runtimes at extreme throughput. | |
| Scaling Organizational Leadership and Operational Culture | 6 | 4 | 1 | 1 | Tuhin reflects on scaling organizational leadership, crediting Elad with convincing him to move past an overly flat engineering-only culture toward hiring executive leaders and cultivating a serious PagerDuty operations culture. | |
| Jevons Paradox and the Future of Ubiquitous AI Agents | 5 | 4 | 1 | 1 | The hosts and guest discuss Jevons paradox in AI intelligence. Tuhin explains that cheaper inference directly leads to longer-running agent loops and higher token consumption rather than saturation. |