Mar 21, 2024 · 38m · no-priors

No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava

Tuhin Srivastava · 25m spoken Sarah Guo · 6m spoken Elad Gil · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Baseten CEO and co-founder Tuhin Srivastava joins Sarah Guo and Elad Gil on No Priors to discuss the technical complexities of machine learning inference, GPU hardware constraints, and the shifting unit economics of AI-native software.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.6% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 4.5 Guest disagreement 1.0 The hosts pushing back 0.1
05100:0010:0020:0030:000:24–4:09 · The hosts as informed peer 4/10 Origins of Baseten and the Efficient Code Philosophy Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health.4:10–6:12 · The hosts as informed peer 5/10 Key Differences Between AI Training and Inference Workloads Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training.6:12–8:44 · The hosts as informed peer 5/10 Market Acceleration, Velocity, and Evolving GPU Hardware Demands Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage.8:44–12:00 · The hosts as informed peer 6/10 Inference Performance Benchmarking and Low-Level Kernel Optimization Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques.12:00–15:41 · The hosts as informed peer 6/10 Optimizing Multimodal Speech and Diffusion Model Workloads Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices.15:41–17:58 · The hosts as informed peer 5/10 Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy.17:58–21:24 · The hosts as informed peer 6/10 Enterprise AI Adoption Trajectories and Long-Term Value Creation Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value.21:25–24:27 · The hosts as informed peer 7/10 Shifting SaaS Unit Economics and Rising Compute Spend Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll.24:27–27:22 · The hosts as informed peer 8/10 Startup Revenue Ramps, Defensibility, and Oligopolistic Markets Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies.27:22–31:10 · The hosts as informed peer 8/10 Industry Disruption Patterns and Defensibility of Venture Capital After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data.31:10–35:14 · The hosts as informed peer 6/10 GPU Availability Dynamics and the Challenges of Hardware Heterogeneity Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes.35:14–38:08 · The hosts as informed peer 5/10 The Build Versus Buy Decision in AI Infrastructure Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions.38:08–38:28 · The hosts as informed peer 0/10 Episode Conclusion, Host Farewell, and Show Subscription Information Standard show wrap-up, thank yous, and subscription call to action.0:24–4:09 · Guest teaching 3/10 Origins of Baseten and the Efficient Code Philosophy Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health.4:10–6:12 · Guest teaching 6/10 Key Differences Between AI Training and Inference Workloads Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training.6:12–8:44 · Guest teaching 4/10 Market Acceleration, Velocity, and Evolving GPU Hardware Demands Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage.8:44–12:00 · Guest teaching 7/10 Inference Performance Benchmarking and Low-Level Kernel Optimization Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques.12:00–15:41 · Guest teaching 6/10 Optimizing Multimodal Speech and Diffusion Model Workloads Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices.15:41–17:58 · Guest teaching 6/10 Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy.17:58–21:24 · Guest teaching 5/10 Enterprise AI Adoption Trajectories and Long-Term Value Creation Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value.21:25–24:27 · Guest teaching 3/10 Shifting SaaS Unit Economics and Rising Compute Spend Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll.24:27–27:22 · Guest teaching 3/10 Startup Revenue Ramps, Defensibility, and Oligopolistic Markets Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies.27:22–31:10 · Guest teaching 2/10 Industry Disruption Patterns and Defensibility of Venture Capital After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data.31:10–35:14 · Guest teaching 7/10 GPU Availability Dynamics and the Challenges of Hardware Heterogeneity Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes.35:14–38:08 · Guest teaching 6/10 The Build Versus Buy Decision in AI Infrastructure Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions.38:08–38:28 · Guest teaching 0/10 Episode Conclusion, Host Farewell, and Show Subscription Information Standard show wrap-up, thank yous, and subscription call to action.0:24–4:09 · Guest disagreement 1/10 Origins of Baseten and the Efficient Code Philosophy Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health.4:10–6:12 · Guest disagreement 1/10 Key Differences Between AI Training and Inference Workloads Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training.6:12–8:44 · Guest disagreement 0/10 Market Acceleration, Velocity, and Evolving GPU Hardware Demands Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage.8:44–12:00 · Guest disagreement 1/10 Inference Performance Benchmarking and Low-Level Kernel Optimization Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques.12:00–15:41 · Guest disagreement 1/10 Optimizing Multimodal Speech and Diffusion Model Workloads Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices.15:41–17:58 · Guest disagreement 0/10 Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy.17:58–21:24 · Guest disagreement 2/10 Enterprise AI Adoption Trajectories and Long-Term Value Creation Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value.21:25–24:27 · Guest disagreement 0/10 Shifting SaaS Unit Economics and Rising Compute Spend Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll.24:27–27:22 · Guest disagreement 0/10 Startup Revenue Ramps, Defensibility, and Oligopolistic Markets Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies.27:22–31:10 · Guest disagreement 1/10 Industry Disruption Patterns and Defensibility of Venture Capital After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data.31:10–35:14 · Guest disagreement 4/10 GPU Availability Dynamics and the Challenges of Hardware Heterogeneity Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes.35:14–38:08 · Guest disagreement 2/10 The Build Versus Buy Decision in AI Infrastructure Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions.38:08–38:28 · Guest disagreement 0/10 Episode Conclusion, Host Farewell, and Show Subscription Information Standard show wrap-up, thank yous, and subscription call to action.0:24–4:09 · The hosts pushing back 0/10 Origins of Baseten and the Efficient Code Philosophy Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health.4:10–6:12 · The hosts pushing back 0/10 Key Differences Between AI Training and Inference Workloads Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training.6:12–8:44 · The hosts pushing back 0/10 Market Acceleration, Velocity, and Evolving GPU Hardware Demands Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage.8:44–12:00 · The hosts pushing back 0/10 Inference Performance Benchmarking and Low-Level Kernel Optimization Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques.12:00–15:41 · The hosts pushing back 0/10 Optimizing Multimodal Speech and Diffusion Model Workloads Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices.15:41–17:58 · The hosts pushing back 0/10 Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy.17:58–21:24 · The hosts pushing back 1/10 Enterprise AI Adoption Trajectories and Long-Term Value Creation Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value.21:25–24:27 · The hosts pushing back 0/10 Shifting SaaS Unit Economics and Rising Compute Spend Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll.24:27–27:22 · The hosts pushing back 0/10 Startup Revenue Ramps, Defensibility, and Oligopolistic Markets Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies.27:22–31:10 · The hosts pushing back 1/10 Industry Disruption Patterns and Defensibility of Venture Capital After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data.31:10–35:14 · The hosts pushing back 0/10 GPU Availability Dynamics and the Challenges of Hardware Heterogeneity Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes.35:14–38:08 · The hosts pushing back 0/10 The Build Versus Buy Decision in AI Infrastructure Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions.38:08–38:28 · The hosts pushing back 0/10 Episode Conclusion, Host Farewell, and Show Subscription Information Standard show wrap-up, thank yous, and subscription call to action.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 20.9% · guest 79.1%0:00 · the hosts 20.9% · guest 79.1%3:00 · the hosts 12% · guest 88%3:00 · the hosts 12% · guest 88%6:00 · the hosts 16.1% · guest 83.9%6:00 · the hosts 16.1% · guest 83.9%9:00 · the hosts 1.5% · guest 98.5%9:00 · the hosts 1.5% · guest 98.5%12:00 · the hosts 34.8% · guest 65.2%12:00 · the hosts 34.8% · guest 65.2%15:00 · the hosts 6.6% · guest 93.4%15:00 · the hosts 6.6% · guest 93.4%18:00 · the hosts 20% · guest 80%18:00 · the hosts 20% · guest 80%21:00 · the hosts 46% · guest 54%21:00 · the hosts 46% · guest 54%24:00 · the hosts 58.9% · guest 41.1%24:00 · the hosts 58.9% · guest 41.1%27:00 · the hosts 69.5% · guest 30.5%27:00 · the hosts 69.5% · guest 30.5%30:00 · the hosts 56.3% · guest 43.7%30:00 · the hosts 56.3% · guest 43.7%33:00 · the hosts 2.2% · guest 97.8%33:00 · the hosts 2.2% · guest 97.8%36:00 · the hosts 12.3% · guest 87.7%36:00 · the hosts 12.3% · guest 87.7%
Sharpest disagreement ▶ 32:55 Pushback on hardware heterogeneity and AMD viability

Tuhin rejects industry optimism regarding running CUDA on AMD chips, calling the ease of adoption overstated and highlighting the harsh reality of node-level debugging.

Hardest push from the hosts ▶ 19:05 Challenging the near-term enterprise AI adoption timeline

Tuhin reframes Elad's optimistic 10x adoption wave prediction, arguing that top-down enterprise spend in the next 12-18 months risks repeating the empty ML hype cycles of 2018-2020.

Biggest teaching moment ▶ 9:30 Granular breakdown of inference speed and kernel optimization

Tuhin educates the hosts on the exact technical barriers in LLM inference serving, explaining TRT-LLM forks, low-level kernel rewriting, and speculative decoding.

The host holds their own ▶ 24:27 Historical analysis of revenue ramps and oligopolistic moats

Elad demonstrates deep market expertise by contextualizing current AI revenue spikes against 1990s telecom buildouts and explaining contractual locking in oligopoly markets.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Baseten and the Efficient Code Philosophy 4310 Sarah introduces Baseten and asks about their philosophy of efficient code versus no-code. Tuhin explains why engineers need code-level abstractions rather than rigid no-code tools, sharing concrete customer examples like Planned AI and Picnic Health.
Key Differences Between AI Training and Inference Workloads 5610 Sarah asks whether training and inference workloads are fundamentally different given recent massive GPU spend. Tuhin provides a detailed breakdown of SLA requirements, intra-node networking dependencies, and strict uptime expectations in inference versus training.
Market Acceleration, Velocity, and Evolving GPU Hardware Demands 5400 Elad inquires about what surprised Baseten most as the market accelerated. Tuhin notes the dramatic shift from the quiet 2019-2022 era to rapid GPU shifts (T4 to H100) and how startup velocity has become the dominant competitive advantage.
Inference Performance Benchmarking and Low-Level Kernel Optimization 6710 Sarah highlights Baseten's top placement on Artificial Analysis benchmarks and asks what drives low latency. Tuhin explains low-level kernel rewriting, TRT-LLM integration with Nvidia, and speculative decoding techniques.
Optimizing Multimodal Speech and Diffusion Model Workloads 6610 Elad asks about optimization trends across non-text modalities like speech and image generation. Tuhin discusses Whisper on TRT, batching techniques, and the growing trend toward localized, distilled models running on client devices.
Deployment Patterns: Shared Endpoints, Dedicated Cloud, and VPCs 5600 Sarah asks how teams decide between public endpoints, dedicated cloud instances, and self-hosting. Tuhin maps out the lifecycle from shared APIs to dedicated instances to customer-VPC deployments driven by cost commitments and data privacy.
Enterprise AI Adoption Trajectories and Long-Term Value Creation 6521 Elad suggests enterprises are on the verge of a 10x wave of AI adoption. Tuhin offers a counter-perspective, warning that near-term enterprise spending risks repeating the top-down ML hype trap of 2018-2020, while remaining bullish on 3-5 year compounding value.
Shifting SaaS Unit Economics and Rising Compute Spend 7300 Sarah explains how shifting COGS and compute expenses are replacing headcount in modern AI companies. Tuhin validates this with an anecdote of a customer whose compute bill surpassed everything except payroll.
Startup Revenue Ramps, Defensibility, and Oligopolistic Markets 8300 Elad draws historical parallels to 1990s dot-com telecom buildouts, questioning whether sudden zero-to-ten-million ramps represent durable moats or commoditized demand, analyzing contractual lock-in and oligopolies.
Industry Disruption Patterns and Defensibility of Venture Capital 8211 After Tuhin questions whether hedge funds or software jobs are safe, Sarah delivers a detailed monologue analyzing why early-stage venture capital is structurally defensible against AI automation due to sparse, out-of-distribution human data.
GPU Availability Dynamics and the Challenges of Hardware Heterogeneity 6740 Sarah inquires about GPU shortages and hardware heterogeneity like AMD alternatives. Tuhin pushes back skeptically against claims of seamless non-Nvidia adoption, detailing the severe practical challenges of porting CUDA and debugging unstable nodes.
The Build Versus Buy Decision in AI Infrastructure 5620 Elad asks about the build-versus-buy decision for infrastructure. Tuhin argues strongly that in-house infrastructure is a distraction that risks catastrophic downtime, citing customers that abandoned internal systems to move to managed solutions.
Episode Conclusion, Host Farewell, and Show Subscription Information 0000 Standard show wrap-up, thank yous, and subscription call to action.

Statements from this episode (19)

Disclosure
Srivastava: Baseten powers AI features for Descript and Patreon
“We work with companies like Descript where AI is very, very core to the product experience. We power a lot of AI features that Patreon has shipped.”
Tuhin Srivastava Mar 21, 2024 ▶ 2:34
Insight
Srivastava: Inter-rack networking matters less for AI inference than training
“Even the GPU clusters themselves, like, you know, the full training networking is a very, very important Piece to have networking on the racks themselves with inference and matters a little less because you're doing a little bit more on individual GPUs and les…”
Tuhin Srivastava Mar 21, 2024 ▶ 4:57
Insight
Srivastava: AI inference demands strict uptime, while training tolerates node terminations
“Resiliency and reliability matters a lot more. You know, downtime is unacceptable from an input perspective. Nodes get terminated all the time from a training perspective.”
Tuhin Srivastava Mar 21, 2024 ▶ 6:01
Assertion Not checkable as stated
Srivastava: Customer compute demands continuously escalate to A100s and H100s
“At the end of 2022, most of our customers were using T-folds and AT&Gs. You know, after that it changed to A-one hundreds. Now it's going to H-one hundreds. Like the computer needs aren't necessarily going down. They're only going up, especially as these servi…”
Tuhin Srivastava Mar 21, 2024 ▶ 8:28
Insight
Srivastava: Inference lacks abstractions, forcing engineers to rewrite low-level GPU kernels
“A lot of the optimization you're doing is pretty low level and there's no real abstraction. So you either have to learn how to use open source very well or rewrite some of these kernels by yourself.”
Tuhin Srivastava Mar 21, 2024 ▶ 10:38
Prediction Not checkable as stated
Srivastava: LLM inference performance will commoditize and increasingly run locally
“We think over time, It will get somewhat commoditized, the performance, especially, especially for language models, to be honest. I think, you know, more and more of that stuff should run locally to some degree, I think”
Tuhin Srivastava Mar 21, 2024 ▶ 11:34
Opinion
Srivastava: Shared model endpoints do not work for enterprise AI workloads
“Shared endpoints don't really work for large companies, it's not going to work. They're not definitely not going to work for enterprises.”
Tuhin Srivastava Mar 21, 2024 ▶ 17:03
Insight
Srivastava: Self-hosting AI models offers massive cost advantages to large enterprises
“Especially larger customers, they're going to have actually pretty good compute deals and they're going to have their own spent you know, their own credit system or spend commit with the marketplaces and actually running it on the infrastructure solves a lot o…”
Tuhin Srivastava Mar 21, 2024 ▶ 17:24
Insight
Srivastava: Model deployment progresses from shared endpoints to dedicated to self-hosted cloud
“I think there's like three stages to it, which is like you started sharing from standpoint providers, you go to dedicated in the cloud, But I think for some customers, that's not enough either, and you want to go into your own cloud.”
Tuhin Srivastava Mar 21, 2024 ▶ 17:48
Assertion Not checkable as stated
Srivastava: Most enterprises start AI adoption with developer copilots
“What we're seeing now is that, like, copilots, especially cogen stuff that actually has already made its way into the enterprise. Like, most enterprises we talk to, like, when you say, when you ask them how advanced are you on their AI strategy, they'll tell y…”
Tuhin Srivastava Mar 21, 2024 ▶ 18:49
Disclosure
Guo: Portfolio startups show higher headcount efficiency alongside rising compute spend
“And, you know, I, at least in, in my portfolio, we are seeing more efficiency on headcount and like a lot more compute spend.”
Sarah Guo Mar 21, 2024 ▶ 22:30
Insight
Srivastava: Efficient AI businesses can achieve healthy margins despite heavy compute spend
“And, you know, what, what's crazy about it though, is I think the most efficient businesses through markups and through software optimization can actually direct, like drive pretty healthy margins. And still have these really aggressive consumption contract.”
Tuhin Srivastava Mar 21, 2024 ▶ 23:53
Insight
Gil: Fast revenue ramps on young AI products may indicate lack of defensibility
“Here it feels like things are ramping really fast off of products that are a couple months old, which sometimes suggests that there's not defensibility.”
Elad Gil Mar 21, 2024 ▶ 24:52
Assertion Not checkable as stated
Gil: Multiple competing AI startups are reaching $10M revenue in one year
“There's a couple different markets where suddenly you see three companies all go from zero to five or zero to ten million of revenue in a year.”
Elad Gil Mar 21, 2024 ▶ 25:05
Opinion
Srivastava: Hedge funds and financial services lag in generative AI adoption
“I don't actually think they're on the cusp of it as much as, you know, you'd think, like, I, you know, if you think about a lot of this, the big data stuff, like, 10 years ago you know, the hedge funds were all over that, right? They're like, hey, there's alph…”
Tuhin Srivastava Mar 21, 2024 ▶ 27:33
Insight
Guo: AI agents cannot automate early-stage VC due to missing data
“At the early stages, a lot of the data doesn't exist, right? Like you'd have to capture real world data. You have increasingly meetings over zoom, but you would want to capture a lot of information about people. So much of it is access and the information abou…”
Sarah Guo Mar 21, 2024 ▶ 29:12
Assertion Not checkable as stated
Srivastava: Securing H100s still requires weeks of provider negotiations and escalations
“I think customers are still struggling with availability for the most premium chips. And I think, you know, whether that's eight, 108, 100, I think even when there is availability, you're oftentimes looking for like three to six weeks of negotiating with cloud…”
Tuhin Srivastava Mar 21, 2024 ▶ 32:24
Opinion
Srivastava: The ease of running CUDA workloads on AMD is significantly overstated
“I've, I personally think that it's pretty overstated how easy it is to run something that looks like CUDA or CUDA in some form on an AMD chip seems, seems like a challenge to me.”
Tuhin Srivastava Mar 21, 2024 ▶ 33:10
Disclosure
Srivastava: Baseten is avoiding non-NVIDIA hardware investments in the short term
“I think short term you know, it's really hard for me to see how we make investments beyond NVIDIA, especially when there's a customer, where there's like customer A crunch on the other side from customers who are like, hey, we need this now.”
Tuhin Srivastava Mar 21, 2024 ▶ 33:51
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.