Mar 21, 2024 · 38m · no-priors
No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Baseten CEO and co-founder Tuhin Srivastava joins Sarah Guo and Elad Gil on No Priors to discuss the technical complexities of machine learning inference, GPU hardware constraints, and the shifting unit economics of AI-native software.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Tuhin rejects industry optimism regarding running CUDA on AMD chips, calling the ease of adoption overstated and highlighting the harsh reality of node-level debugging.
Hardest push from the hosts ▶ 19:05 Challenging the near-term enterprise AI adoption timelineTuhin reframes Elad's optimistic 10x adoption wave prediction, arguing that top-down enterprise spend in the next 12-18 months risks repeating the empty ML hype cycles of 2018-2020.
Biggest teaching moment ▶ 9:30 Granular breakdown of inference speed and kernel optimizationTuhin educates the hosts on the exact technical barriers in LLM inference serving, explaining TRT-LLM forks, low-level kernel rewriting, and speculative decoding.
The host holds their own ▶ 24:27 Historical analysis of revenue ramps and oligopolistic moatsElad demonstrates deep market expertise by contextualizing current AI revenue spikes against 1990s telecom buildouts and explaining contractual locking in oligopoly markets.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Origins of Baseten and the Efficient Code Philosophy | 4 | 3 | 1 | 0 | Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health. | |
| Key Differences Between AI Training and Inference Workloads | 5 | 6 | 1 | 0 | Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training. | |
| Market Acceleration, Velocity, and Evolving GPU Hardware Demands | 5 | 4 | 0 | 0 | Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage. | |
| Inference Performance Benchmarking and Low-Level Kernel Optimization | 6 | 7 | 1 | 0 | Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques. | |
| Optimizing Multimodal Speech and Diffusion Model Workloads | 6 | 6 | 1 | 0 | Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices. | |
| Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs | 5 | 6 | 0 | 0 | Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy. | |
| Enterprise AI Adoption Trajectories and Long-Term Value Creation | 6 | 5 | 2 | 1 | Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value. | |
| Shifting SaaS Unit Economics and Rising Compute Spend | 7 | 3 | 0 | 0 | Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll. | |
| Startup Revenue Ramps, Defensibility, and Oligopolistic Markets | 8 | 3 | 0 | 0 | Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies. | |
| Industry Disruption Patterns and Defensibility of Venture Capital | 8 | 2 | 1 | 1 | After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data. | |
| GPU Availability Dynamics and the Challenges of Hardware Heterogeneity | 6 | 7 | 4 | 0 | Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes. | |
| The Build Versus Buy Decision in AI Infrastructure | 5 | 6 | 2 | 0 | Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions. | |
| Episode Conclusion, Host Farewell, and Show Subscription Information | 0 | 0 | 0 | 0 | Standard show wrap-up, thank yous, and subscription call to action. |