Oct 31, 2024 · 56m · mad

Can AI Infrastructure Work Like Magic? Erik Bernhardsson, CEO, Modal

Erik Bernhardsson · 42m spoken Matt Turck · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Modal Labs CEO Erik Bernhardsson about building serverless cloud infrastructure for AI and data applications. Erik discusses Modal's custom technical stack, developer-centric product design, go-to-market evolution, engineering management philosophy, and the changing economics of the AI compute landscape.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 19.3% of the talking time here. How this is scored →

Matt as informed peer 3.6 Guest teaching 5.5 Guest disagreement 1.7 Matt pushing back 2.1
05100:0015:0030:0045:000:56–3:55 · Matt as informed peer 3/10 Modal's Angel Investors and Data Influencer Strategy Matt opens with friendly banter regarding Modal's impressive list of angel investors and jokes about 'data influencers.' Erik concurs playfully and lightheartedly recalls citing Matt's market map in Modal's launch post.3:55–8:34 · Matt as informed peer 4/10 Categorizing the AI Compute Stack: From Hyperscalers to Serverless Matt prompts Erik to categorize the AI compute space. Erik lays out a clear taxonomy spanning hyperscalers, alt-clouds, LLM API providers, and serverless code platforms, while Matt summarizes alt-clouds as real estate businesses.8:34–12:56 · Matt as informed peer 3/10 Thesis on Custom Models vs. Universal Multimodal Models Matt questions whether developers will build custom models rather than relying entirely on universal API models. Erik explains his thesis on why domain-specific models like voice cloning and video models will persist alongside general multimodal models.12:56–17:34 · Matt as informed peer 3/10 Convergence of Data, Machine Learning, and Artificial Intelligence Erik shares the deep technical origin of Modal, explaining why off-the-shelf tools like Kubernetes and Docker were insufficient for fast container cold starts. Matt asks clarifying questions about development timeline.17:34–20:28 · Matt as informed peer 3/10 Product-Market Fit with Generative AI and Multi-Cloud GPU Orchestration Erik details finding PMF during the 2022 generative AI boom with Stable Diffusion. He explains how Modal manages multi-cloud capacity across AWS, GCP, and Oracle using dynamic optimization algorithms.20:28–24:35 · Matt as informed peer 4/10 Real-World Use Cases: Suno AI and the Economics of Inference Erik uses Suno AI as an example of audio generation workloads running on Modal. He explains why inference spend will eventually eclipse training spend as models move from creation to commercial usage.24:35–26:46 · Matt as informed peer 3/10 Sandboxed Code Execution for AI Code Agents Matt asks about batch processing and sandboxed code execution features. Erik explains how sandboxing arbitrary untrusted code enables emerging LLM coding agents.26:46–29:08 · Matt as informed peer 3/10 Why Serverless Fits AI and Data Teams Better Than Traditional Web Dev Erik explains why serverless initially failed to gain massive traction in web development but fits data and AI teams perfectly due to bursty workloads and distinct developer ergonomics.29:08–31:36 · Matt as informed peer 5/10 Cloud Maximalism vs. Self-Hosting and BYOC Matt pushes on enterprise demand for self-hosting and BYOC setups. Erik forcefully rejects typical startup wisdom, arguing that customer requests for self-hosting usually mean the product isn't compelling enough to waive soft security preferences.31:36–36:25 · Matt as informed peer 6/10 Pricing Power, Unit Economics, and GPU Market Dynamics Matt directly challenges Modal's exposure to GPU commitments by asking if they risk becoming a WeWork. Erik explains how Modal maintains pricing power and software margins above raw hardware COGS.36:25–42:12 · Matt as informed peer 4/10 Designing Developer Tools for Humans: Frictionless Onboarding and Magic UX Matt cites Erik's blog post about building developer tools for humans. Erik breaks down onboarding ergonomics, documentation mistakes, and comparing developer experience magic to early Spotify.42:12–45:33 · Matt as informed peer 3/10 Simplifying Product and SDK Features Matt asks about removing unused features and transitions to go-to-market strategy. Erik discusses moving from organic social media marketing on Twitter/X to hiring dedicated enterprise sales roles.45:33–48:27 · Matt as informed peer 3/10 Autonomy and Management in Early Engineering Teams Matt highlights Erik's low-overhead management style. Erik reflects on early Spotify autonomy and argues smart engineering teams thrive when given context rather than heavy project management.48:27–52:05 · Matt as informed peer 4/10 Sequential Process for Hiring and Building Roles Matt asks about Erik's philosophy on expanding company functions. Erik explains why founders must do jobs themselves before hiring managers, and highlights building engineering teams in New York and Europe.52:05–55:45 · Matt as informed peer 3/10 Modal's Roadmap: Real-Time Compute and Batch Scaling Erik outlines Modal's roadmap around sub-150ms real-time audio/video inference and massive batch scaling, concluding with predictions on consumer AI applications.0:56–3:55 · Guest teaching 2/10 Modal's Angel Investors and Data Influencer Strategy Matt opens with friendly banter regarding Modal's impressive list of angel investors and jokes about 'data influencers.' Erik concurs playfully and lightheartedly recalls citing Matt's market map in Modal's launch post.3:55–8:34 · Guest teaching 6/10 Categorizing the AI Compute Stack: From Hyperscalers to Serverless Matt prompts Erik to categorize the AI compute space. Erik lays out a clear taxonomy spanning hyperscalers, alt-clouds, LLM API providers, and serverless code platforms, while Matt summarizes alt-clouds as real estate businesses.8:34–12:56 · Guest teaching 6/10 Thesis on Custom Models vs. Universal Multimodal Models Matt questions whether developers will build custom models rather than relying entirely on universal API models. Erik explains his thesis on why domain-specific models like voice cloning and video models will persist alongside general multimodal models.12:56–17:34 · Guest teaching 7/10 Convergence of Data, Machine Learning, and Artificial Intelligence Erik shares the deep technical origin of Modal, explaining why off-the-shelf tools like Kubernetes and Docker were insufficient for fast container cold starts. Matt asks clarifying questions about development timeline.17:34–20:28 · Guest teaching 6/10 Product-Market Fit with Generative AI and Multi-Cloud GPU Orchestration Erik details finding PMF during the 2022 generative AI boom with Stable Diffusion. He explains how Modal manages multi-cloud capacity across AWS, GCP, and Oracle using dynamic optimization algorithms.20:28–24:35 · Guest teaching 6/10 Real-World Use Cases: Suno AI and the Economics of Inference Erik uses Suno AI as an example of audio generation workloads running on Modal. He explains why inference spend will eventually eclipse training spend as models move from creation to commercial usage.24:35–26:46 · Guest teaching 5/10 Sandboxed Code Execution for AI Code Agents Matt asks about batch processing and sandboxed code execution features. Erik explains how sandboxing arbitrary untrusted code enables emerging LLM coding agents.26:46–29:08 · Guest teaching 6/10 Why Serverless Fits AI and Data Teams Better Than Traditional Web Dev Erik explains why serverless initially failed to gain massive traction in web development but fits data and AI teams perfectly due to bursty workloads and distinct developer ergonomics.29:08–31:36 · Guest teaching 7/10 Cloud Maximalism vs. Self-Hosting and BYOC Matt pushes on enterprise demand for self-hosting and BYOC setups. Erik forcefully rejects typical startup wisdom, arguing that customer requests for self-hosting usually mean the product isn't compelling enough to waive soft security preferences.31:36–36:25 · Guest teaching 6/10 Pricing Power, Unit Economics, and GPU Market Dynamics Matt directly challenges Modal's exposure to GPU commitments by asking if they risk becoming a WeWork. Erik explains how Modal maintains pricing power and software margins above raw hardware COGS.36:25–42:12 · Guest teaching 5/10 Designing Developer Tools for Humans: Frictionless Onboarding and Magic UX Matt cites Erik's blog post about building developer tools for humans. Erik breaks down onboarding ergonomics, documentation mistakes, and comparing developer experience magic to early Spotify.42:12–45:33 · Guest teaching 5/10 Simplifying Product and SDK Features Matt asks about removing unused features and transitions to go-to-market strategy. Erik discusses moving from organic social media marketing on Twitter/X to hiring dedicated enterprise sales roles.45:33–48:27 · Guest teaching 5/10 Autonomy and Management in Early Engineering Teams Matt highlights Erik's low-overhead management style. Erik reflects on early Spotify autonomy and argues smart engineering teams thrive when given context rather than heavy project management.48:27–52:05 · Guest teaching 5/10 Sequential Process for Hiring and Building Roles Matt asks about Erik's philosophy on expanding company functions. Erik explains why founders must do jobs themselves before hiring managers, and highlights building engineering teams in New York and Europe.52:05–55:45 · Guest teaching 6/10 Modal's Roadmap: Real-Time Compute and Batch Scaling Erik outlines Modal's roadmap around sub-150ms real-time audio/video inference and massive batch scaling, concluding with predictions on consumer AI applications.0:56–3:55 · Guest disagreement 1/10 Modal's Angel Investors and Data Influencer Strategy Matt opens with friendly banter regarding Modal's impressive list of angel investors and jokes about 'data influencers.' Erik concurs playfully and lightheartedly recalls citing Matt's market map in Modal's launch post.3:55–8:34 · Guest disagreement 2/10 Categorizing the AI Compute Stack: From Hyperscalers to Serverless Matt prompts Erik to categorize the AI compute space. Erik lays out a clear taxonomy spanning hyperscalers, alt-clouds, LLM API providers, and serverless code platforms, while Matt summarizes alt-clouds as real estate businesses.8:34–12:56 · Guest disagreement 2/10 Thesis on Custom Models vs. Universal Multimodal Models Matt questions whether developers will build custom models rather than relying entirely on universal API models. Erik explains his thesis on why domain-specific models like voice cloning and video models will persist alongside general multimodal models.12:56–17:34 · Guest disagreement 1/10 Convergence of Data, Machine Learning, and Artificial Intelligence Erik shares the deep technical origin of Modal, explaining why off-the-shelf tools like Kubernetes and Docker were insufficient for fast container cold starts. Matt asks clarifying questions about development timeline.17:34–20:28 · Guest disagreement 1/10 Product-Market Fit with Generative AI and Multi-Cloud GPU Orchestration Erik details finding PMF during the 2022 generative AI boom with Stable Diffusion. He explains how Modal manages multi-cloud capacity across AWS, GCP, and Oracle using dynamic optimization algorithms.20:28–24:35 · Guest disagreement 2/10 Real-World Use Cases: Suno AI and the Economics of Inference Erik uses Suno AI as an example of audio generation workloads running on Modal. He explains why inference spend will eventually eclipse training spend as models move from creation to commercial usage.24:35–26:46 · Guest disagreement 1/10 Sandboxed Code Execution for AI Code Agents Matt asks about batch processing and sandboxed code execution features. Erik explains how sandboxing arbitrary untrusted code enables emerging LLM coding agents.26:46–29:08 · Guest disagreement 2/10 Why Serverless Fits AI and Data Teams Better Than Traditional Web Dev Erik explains why serverless initially failed to gain massive traction in web development but fits data and AI teams perfectly due to bursty workloads and distinct developer ergonomics.29:08–31:36 · Guest disagreement 4/10 Cloud Maximalism vs. Self-Hosting and BYOC Matt pushes on enterprise demand for self-hosting and BYOC setups. Erik forcefully rejects typical startup wisdom, arguing that customer requests for self-hosting usually mean the product isn't compelling enough to waive soft security preferences.31:36–36:25 · Guest disagreement 3/10 Pricing Power, Unit Economics, and GPU Market Dynamics Matt directly challenges Modal's exposure to GPU commitments by asking if they risk becoming a WeWork. Erik explains how Modal maintains pricing power and software margins above raw hardware COGS.36:25–42:12 · Guest disagreement 1/10 Designing Developer Tools for Humans: Frictionless Onboarding and Magic UX Matt cites Erik's blog post about building developer tools for humans. Erik breaks down onboarding ergonomics, documentation mistakes, and comparing developer experience magic to early Spotify.42:12–45:33 · Guest disagreement 1/10 Simplifying Product and SDK Features Matt asks about removing unused features and transitions to go-to-market strategy. Erik discusses moving from organic social media marketing on Twitter/X to hiring dedicated enterprise sales roles.45:33–48:27 · Guest disagreement 2/10 Autonomy and Management in Early Engineering Teams Matt highlights Erik's low-overhead management style. Erik reflects on early Spotify autonomy and argues smart engineering teams thrive when given context rather than heavy project management.48:27–52:05 · Guest disagreement 2/10 Sequential Process for Hiring and Building Roles Matt asks about Erik's philosophy on expanding company functions. Erik explains why founders must do jobs themselves before hiring managers, and highlights building engineering teams in New York and Europe.52:05–55:45 · Guest disagreement 1/10 Modal's Roadmap: Real-Time Compute and Batch Scaling Erik outlines Modal's roadmap around sub-150ms real-time audio/video inference and massive batch scaling, concluding with predictions on consumer AI applications.0:56–3:55 · Matt pushing back 1/10 Modal's Angel Investors and Data Influencer Strategy Matt opens with friendly banter regarding Modal's impressive list of angel investors and jokes about 'data influencers.' Erik concurs playfully and lightheartedly recalls citing Matt's market map in Modal's launch post.3:55–8:34 · Matt pushing back 2/10 Categorizing the AI Compute Stack: From Hyperscalers to Serverless Matt prompts Erik to categorize the AI compute space. Erik lays out a clear taxonomy spanning hyperscalers, alt-clouds, LLM API providers, and serverless code platforms, while Matt summarizes alt-clouds as real estate businesses.8:34–12:56 · Matt pushing back 3/10 Thesis on Custom Models vs. Universal Multimodal Models Matt questions whether developers will build custom models rather than relying entirely on universal API models. Erik explains his thesis on why domain-specific models like voice cloning and video models will persist alongside general multimodal models.12:56–17:34 · Matt pushing back 1/10 Convergence of Data, Machine Learning, and Artificial Intelligence Erik shares the deep technical origin of Modal, explaining why off-the-shelf tools like Kubernetes and Docker were insufficient for fast container cold starts. Matt asks clarifying questions about development timeline.17:34–20:28 · Matt pushing back 1/10 Product-Market Fit with Generative AI and Multi-Cloud GPU Orchestration Erik details finding PMF during the 2022 generative AI boom with Stable Diffusion. He explains how Modal manages multi-cloud capacity across AWS, GCP, and Oracle using dynamic optimization algorithms.20:28–24:35 · Matt pushing back 2/10 Real-World Use Cases: Suno AI and the Economics of Inference Erik uses Suno AI as an example of audio generation workloads running on Modal. He explains why inference spend will eventually eclipse training spend as models move from creation to commercial usage.24:35–26:46 · Matt pushing back 1/10 Sandboxed Code Execution for AI Code Agents Matt asks about batch processing and sandboxed code execution features. Erik explains how sandboxing arbitrary untrusted code enables emerging LLM coding agents.26:46–29:08 · Matt pushing back 2/10 Why Serverless Fits AI and Data Teams Better Than Traditional Web Dev Erik explains why serverless initially failed to gain massive traction in web development but fits data and AI teams perfectly due to bursty workloads and distinct developer ergonomics.29:08–31:36 · Matt pushing back 5/10 Cloud Maximalism vs. Self-Hosting and BYOC Matt pushes on enterprise demand for self-hosting and BYOC setups. Erik forcefully rejects typical startup wisdom, arguing that customer requests for self-hosting usually mean the product isn't compelling enough to waive soft security preferences.31:36–36:25 · Matt pushing back 5/10 Pricing Power, Unit Economics, and GPU Market Dynamics Matt directly challenges Modal's exposure to GPU commitments by asking if they risk becoming a WeWork. Erik explains how Modal maintains pricing power and software margins above raw hardware COGS.36:25–42:12 · Matt pushing back 1/10 Designing Developer Tools for Humans: Frictionless Onboarding and Magic UX Matt cites Erik's blog post about building developer tools for humans. Erik breaks down onboarding ergonomics, documentation mistakes, and comparing developer experience magic to early Spotify.42:12–45:33 · Matt pushing back 2/10 Simplifying Product and SDK Features Matt asks about removing unused features and transitions to go-to-market strategy. Erik discusses moving from organic social media marketing on Twitter/X to hiring dedicated enterprise sales roles.45:33–48:27 · Matt pushing back 2/10 Autonomy and Management in Early Engineering Teams Matt highlights Erik's low-overhead management style. Erik reflects on early Spotify autonomy and argues smart engineering teams thrive when given context rather than heavy project management.48:27–52:05 · Matt pushing back 2/10 Sequential Process for Hiring and Building Roles Matt asks about Erik's philosophy on expanding company functions. Erik explains why founders must do jobs themselves before hiring managers, and highlights building engineering teams in New York and Europe.52:05–55:45 · Matt pushing back 1/10 Modal's Roadmap: Real-Time Compute and Batch Scaling Erik outlines Modal's roadmap around sub-150ms real-time audio/video inference and massive batch scaling, concluding with predictions on consumer AI applications.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 70.7% · guest 29.3%0:00 · Matt 70.7% · guest 29.3%3:00 · Matt 26% · guest 74%3:00 · Matt 26% · guest 74%6:00 · Matt 22.5% · guest 77.5%6:00 · Matt 22.5% · guest 77.5%9:00 · Matt 5.7% · guest 94.3%9:00 · Matt 5.7% · guest 94.3%12:00 · Matt 38% · guest 62%12:00 · Matt 38% · guest 62%15:00 · Matt 0.4% · guest 99.6%15:00 · Matt 0.4% · guest 99.6%18:00 · Matt 11.6% · guest 88.4%18:00 · Matt 11.6% · guest 88.4%21:00 · Matt 23.5% · guest 76.5%21:00 · Matt 23.5% · guest 76.5%24:00 · Matt 15.1% · guest 84.9%24:00 · Matt 15.1% · guest 84.9%27:00 · Matt 8.5% · guest 91.5%27:00 · Matt 8.5% · guest 91.5%30:00 · Matt 10.6% · guest 89.4%30:00 · Matt 10.6% · guest 89.4%33:00 · Matt 11.7% · guest 88.3%33:00 · Matt 11.7% · guest 88.3%36:00 · Matt 14.4% · guest 85.6%36:00 · Matt 14.4% · guest 85.6%39:00 · Matt 12.8% · guest 87.2%39:00 · Matt 12.8% · guest 87.2%42:00 · Matt 26.1% · guest 73.9%42:00 · Matt 26.1% · guest 73.9%45:00 · Matt 16.5% · guest 83.5%45:00 · Matt 16.5% · guest 83.5%48:00 · Matt 17.2% · guest 82.8%48:00 · Matt 17.2% · guest 82.8%51:00 · Matt 7.1% · guest 92.9%51:00 · Matt 7.1% · guest 92.9%54:00 · Matt 33.2% · guest 66.8%54:00 · Matt 33.2% · guest 66.8%
Sharpest disagreement ▶ 29:50 Dismissing self-hosting demands

Erik forcefully rejects traditional enterprise SaaS wisdom, contending that when prospects insist on self-hosting, it simply means the vendor hasn't built a product good enough for them to care.

Hardest push from Matt ▶ 33:54 The WeWork comparison challenge

Matt directly confronts Erik on unit economics and capital risk, challenging whether Modal's GPU commitment model makes them vulnerable to a WeWork-style duration mismatch.

Biggest teaching moment ▶ 16:20 Ditching Kubernetes and Docker primitives

Erik explains why standard cloud primitives fail for fast cold starts, detailing how Modal threw out Kubernetes and Docker to build custom schedulers and image distribution filesystems.

Matt holds his own ▶ 33:54 Drilling on GPU commitment unit economics

Matt demonstrates sharp financial perspective by pushing Erik hard on whether Modal takes on dangerous long-term GPU liability relative to volatile customer demand.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Modal's Angel Investors and Data Influencer Strategy 3211 Matt opens with friendly banter regarding Modal's impressive list of angel investors and jokes about 'data influencers.' Erik concurs playfully and lightheartedly recalls citing Matt's market map in Modal's launch post.
Categorizing the AI Compute Stack: From Hyperscalers to Serverless 4622 Matt prompts Erik to categorize the AI compute space. Erik lays out a clear taxonomy spanning hyperscalers, alt-clouds, LLM API providers, and serverless code platforms, while Matt summarizes alt-clouds as real estate businesses.
Thesis on Custom Models vs. Universal Multimodal Models 3623 Matt questions whether developers will build custom models rather than relying entirely on universal API models. Erik explains his thesis on why domain-specific models like voice cloning and video models will persist alongside general multimodal models.
Convergence of Data, Machine Learning, and Artificial Intelligence 3711 Erik shares the deep technical origin of Modal, explaining why off-the-shelf tools like Kubernetes and Docker were insufficient for fast container cold starts. Matt asks clarifying questions about development timeline.
Product-Market Fit with Generative AI and Multi-Cloud GPU Orchestration 3611 Erik details finding PMF during the 2022 generative AI boom with Stable Diffusion. He explains how Modal manages multi-cloud capacity across AWS, GCP, and Oracle using dynamic optimization algorithms.
Real-World Use Cases: Suno AI and the Economics of Inference 4622 Erik uses Suno AI as an example of audio generation workloads running on Modal. He explains why inference spend will eventually eclipse training spend as models move from creation to commercial usage.
Sandboxed Code Execution for AI Code Agents 3511 Matt asks about batch processing and sandboxed code execution features. Erik explains how sandboxing arbitrary untrusted code enables emerging LLM coding agents.
Why Serverless Fits AI and Data Teams Better Than Traditional Web Dev 3622 Erik explains why serverless initially failed to gain massive traction in web development but fits data and AI teams perfectly due to bursty workloads and distinct developer ergonomics.
Cloud Maximalism vs. Self-Hosting and BYOC 5745 Matt pushes on enterprise demand for self-hosting and BYOC setups. Erik forcefully rejects typical startup wisdom, arguing that customer requests for self-hosting usually mean the product isn't compelling enough to waive soft security preferences.
Pricing Power, Unit Economics, and GPU Market Dynamics 6635 Matt directly challenges Modal's exposure to GPU commitments by asking if they risk becoming a WeWork. Erik explains how Modal maintains pricing power and software margins above raw hardware COGS.
Designing Developer Tools for Humans: Frictionless Onboarding and Magic UX 4511 Matt cites Erik's blog post about building developer tools for humans. Erik breaks down onboarding ergonomics, documentation mistakes, and comparing developer experience magic to early Spotify.
Simplifying Product and SDK Features 3512 Matt asks about removing unused features and transitions to go-to-market strategy. Erik discusses moving from organic social media marketing on Twitter/X to hiring dedicated enterprise sales roles.
Autonomy and Management in Early Engineering Teams 3522 Matt highlights Erik's low-overhead management style. Erik reflects on early Spotify autonomy and argues smart engineering teams thrive when given context rather than heavy project management.
Sequential Process for Hiring and Building Roles 4522 Matt asks about Erik's philosophy on expanding company functions. Erik explains why founders must do jobs themselves before hiring managers, and highlights building engineering teams in New York and Europe.
Modal's Roadmap: Real-Time Compute and Batch Scaling 3611 Erik outlines Modal's roadmap around sub-150ms real-time audio/video inference and massive batch scaling, concluding with predictions on consumer AI applications.

Statements from this episode (25)

Disclosure
Bernhardsson: Modal containerizes Python code to scale execution to thousands of GPUs
“Modal makes it easy to build, scale and deploy applications in the data, AI, machine learning realm. So, so we basically, you can think of as like, we take, you write a little bit of Python code, and we take that code, we stick it in a container, we execute th…”
Erik Bernhardsson Oct 31, 2024 ▶ 1:51
Disclosure
Bernhardsson: Modal relies on Oracle Cloud for its infrastructure
“I actually like them a lot. We use them. Big fan.”
Erik Bernhardsson Oct 31, 2024 ▶ 4:10
Opinion
Bernhardsson: GPU alt-clouds are real estate businesses adding little stack value
“It's like sort of like almost like real estate, right? Like they own sort of, you know, real estate and they rent it out, arguably sort of, you know, we work, you could say, if you want to be, you know, a little cynical, but like not a ton of value, like highe…”
Erik Bernhardsson Oct 31, 2024 ▶ 5:29
Disclosure
Model inference represents the majority of Modal's platform usage
“Model inference is probably the majority of our use case.”
Erik Bernhardsson Oct 31, 2024 ▶ 6:54
Prediction Not checkable as stated
Bernhardsson: LLM inference compute is shifting toward specialized chips like crypto mining
“I do think it's kind of fascinating how it's like kind of going in the same direction as like crypto mining towards more like specialized hardware.”
Erik Bernhardsson Oct 31, 2024 ▶ 7:41
Prediction Not checkable as stated
Bernhardsson: People will train many custom AI models in the long run
“Like, so, I don't know, I'm a big believer that, like, you know, in the long run, like, people will train a lot of different custom models.”
Erik Bernhardsson Oct 31, 2024 ▶ 9:36
Insight
Bernhardsson: Great developer experience requires cloud execution to feel local
“If you don't have like a good, super fast feedback loop, like you can never make it like so much of developer experience comes from having like a super fast feedback loop where you can like take code and execute in the cloud in a way where it almost feels like…”
Erik Bernhardsson Oct 31, 2024 ▶ 15:09
Assertion Supported
Bernhardsson: Modal discarded Kubernetes and Docker to build a custom stack
“So we built our own, we threw out Kubernetes out the window. We threw out Docker out the window. We built our own File system in order to optimize for how container images are distributed. We built our own scheduler in order to maintain this pool of workers an…”
Erik Bernhardsson Oct 31, 2024 ▶ 16:45
Disclosure
Modal uses real-time pricing optimization to route GPU compute across 100 regions
“We actually have a system that continuously looks at cloud pricing and solves an optimization problem to sort of, you know, figure out How do we allocate capacity across like, you know, a hundred different regions in order to deliver the capacity we need to ou…”
Erik Bernhardsson Oct 31, 2024 ▶ 19:01
Disclosure
Bernhardsson: Modal can typically scale customers to 1,000 GPUs within minutes
“And then if you one day need a thousand GPUs, we can get you, we can typically get you a thousand GPUs like pretty quickly, like talking minutes.”
Erik Bernhardsson Oct 31, 2024 ▶ 20:12
Assertion Not checkable as stated
Erik Bernhardsson: Suno relies on Modal for AI music inference
“Suno which is AI generated music. And, you know, they have a big cluster. They train their own models outside of modal and then they use modal for the inference side.”
Erik Bernhardsson Oct 31, 2024 ▶ 20:41
Disclosure
Erik Bernhardsson: Very large-scale AI training is not a high priority for Modal
“It puts a lot of different, you know constraints on, on the infrastructure that we don't support today. It's something we're interested in, like, you know, somewhat like, you know, looking at down the road but it's not like a super high priority for the busine…”
Erik Bernhardsson Oct 31, 2024 ▶ 21:51
Opinion
Bernhardsson: Companies training their own LLMs makes almost no sense
“A couple of years ago, you know, a lot of companies try to train their own LMs. Like that to me, it makes almost no sense, right?”
Erik Bernhardsson Oct 31, 2024 ▶ 23:52
Prediction Held up
Bernhardsson: AI inference spend will dominate model training spend long-term
“It's obvious that like in the long run, like inference spend will dominate just for like the obvious reason that like, that's where you make the money, right? Like, you know, when you're training a big model, that's like, you know, you, you're spending a lot o…”
Erik Bernhardsson Oct 31, 2024 ▶ 24:04
Insight
Bernhardsson: Serverless architecture fits AI workloads better than web backends
“So it's sort of always felt to me like, you know, in, in hindsight, it looks like the right idea, but applied to the wrong problems. And I always like wonder, you know, now, like I'm looking at like what we built at modal. Like I I'm kind of convinced that lik…”
Erik Bernhardsson Oct 31, 2024 ▶ 28:24
Prediction Not checkable as stated
Bernhardsson: Cloud trends are shifting toward multi-tenancy and application-layer security
“I think the general wind is like blowing in the direction of like multi-tenancy and like, you know, security moving away from the sort of network layer into the app layer and a bunch of different other things.”
Erik Bernhardsson Oct 31, 2024 ▶ 30:28
Insight
Bernhardsson: Demands for self-hosting mean the product isn't compelling enough
“One trap that I think a lot of startups fall into is that, you know, you go out and talk to a lot of customers and they sort of insist on, on self-hosting. In my opinion, like, you know, sometimes that's actually like a true, like hard concerns, but in many ca…”
Erik Bernhardsson Oct 31, 2024 ▶ 30:37
Insight
Bernhardsson: Resource pooling is a rare free lunch in cloud economics
“Resource pooling is like one of the Free lunches that exists in sort of, you know, in the cloud economics, you know, you can get much faster capacity and much more liquid, you know, capacity, et cetera.”
Erik Bernhardsson Oct 31, 2024 ▶ 31:23
Assertion Supported
Bernhardsson: Nvidia H100 GPU prices are falling significantly
“For a while, like, each 100 prices were, like, almost, like, going up, and then they were kind of still for a while, and then now they're, like, actually going down quite a lot, like, you know, what I'm seeing in the market, right?”
Erik Bernhardsson Oct 31, 2024 ▶ 35:03
Assertion Not checkable as stated
Bernhardsson: AWS EC2 retains 50% to 70% profit margins
“At the end of the day, like EC two is like what, like, you know, 50, 60, 70% margins. It's actually, you know, not as much of a commodity as like people think.”
Erik Bernhardsson Oct 31, 2024 ▶ 35:22
Insight
Bernhardsson: Lower GPU prices increase value of developer AI software
“If GPU prices were to crash, like the relative value of that software actually goes up. So like, I'm not necessarily sure that like a crash in GPU would be bad for us. I think it actually could be great for us because, you know, if GPU prices go down, there's …”
Erik Bernhardsson Oct 31, 2024 ▶ 35:54
Assertion Not checkable as stated
Bernhardsson: AWS Lambda has higher profit margins than GPU serverless products
“Margins on Lambda is much higher than margins on, on GPU serverless products.”
Erik Bernhardsson Oct 31, 2024 ▶ 36:16
Assertion Not checkable as stated
Bernhardsson had no assigned manager during his first year at Spotify
“No one told me who my manager was for the last year, for the first year of Spotify.”
Erik Bernhardsson Oct 31, 2024 ▶ 46:00
Insight
Bernhardsson: Early engineering teams thrive with far less formal management
“You can actually get much further, which a lot less of that than you, than people think is my experience. And you're actually better off doing that because you sort of retain people's autonomy. You retain people's ability to sort of, you know, come up with sol…”
Erik Bernhardsson Oct 31, 2024 ▶ 46:45
Insight
Bernhardsson: Sub-200ms latency requires a decentralized control plane
“In order to get to sort of latencies below, you know, a 152 hundred milliseconds, we probably need to split that up and run like a decentralized control plane.”
Erik Bernhardsson Oct 31, 2024 ▶ 52:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.