Feb 19, 2024 · 1h 8m · latent-space

Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal

Erik Bernhardsson · 45m spoken Shawn Wang · 12m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, Modal founder and CEO Erik Bernhardsson discusses modernizing cloud infrastructure for AI and data engineering, highlighting custom container runtimes, serverless GPU economics, and the developer experience of self-provisioning Python systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.3% of the talking time here. How this is scored →

The hosts as informed peer 5.8 Guest teaching 5.8 Guest disagreement 2.0 The hosts pushing back 1.5
05100:0015:0030:0045:001:00:001:12–5:33 · The hosts as informed peer 6/10 Early Career at Spotify, Luigi, and Vector Databases Swyx demonstrates familiarity with Bernhardsson's early career at Spotify, his open-source work on Luigi and Annoy, and vector indexing algorithms like HNSW. Bernhardsson explains the early pre-hype landscape of vector databases in 2012 and how Annoy used recursive hyperspace splitting.5:33–11:44 · The hosts as informed peer 5/10 Scaling Better.com and Conceptualizing Modal's Compute Vision Erik details why existing cloud infrastructure and Kubernetes failed data teams due to burstiness, GPU management, and fragile dependency environments. Swyx interjects with domain context on the modern data stack and feedback loops.11:44–13:46 · The hosts as informed peer 4/10 Re-Engineering Container Runtimes and Custom Network File Systems Erik educates the hosts on how Modal achieved sub-second container cold starts by bypassing standard Docker image pull mechanisms, using low-level runC primitives, and inventing a custom network-based virtual file system with deduplicated page caching.13:47–21:11 · The hosts as informed peer 8/10 Self-Provisioning Runtimes and Python Developer Experience Swyx demonstrates deep domain expertise regarding self-provisioning runtimes, contrasting AWS CDK and SAM with Modal's unified application-and-infrastructure code model. Erik enthusiastically acknowledges Swyx's framing and discusses the trade-offs of embedding infrastructure directly in Python decorators.21:12–28:44 · The hosts as informed peer 6/10 Serverless AI Inference Dynamics and GPU Utilization Alessio queries how serverless handles GPU utilization and batching compared to static memory allocation in foundational training. Erik explains how fast autoscaling and CPU/GPU memory checkpointing enable cost advantages despite higher per-hour GPU margins.28:44–35:20 · The hosts as informed peer 7/10 Modal's Market Positioning: From Data Teams to Layer-Two Cloud Swyx challenges Erik's insistence on labeling Modal's target user as data teams rather than AI engineers, citing adoption data from the Vercel AI Accelerator. Swyx also compares Modal to Replicate and Modular, prompting Erik to articulate Modal's layer-two cloud positioning.35:20–38:00 · The hosts as informed peer 5/10 Overcoming the Cloud Graduation Problem and Enterprise Retention Alessio asks about the enterprise graduation problem historically faced by PaaS providers like Heroku and how it applies to AI workloads. Erik explains why training lacks software differentiation compared to inference and how superior utilization economics retain customers.38:01–44:01 · The hosts as informed peer 6/10 Language Model Workloads, Fine-Tuning, and Massive Parallelism The hosts discuss LLM workloads including fine-tuning and self-contained agents like ErikBot. Swyx brings up his experience using Modal to orchestrate concurrency limits and rate-limiting when building Small Developer.44:04–48:52 · The hosts as informed peer 6/10 Secure Code Sandboxes and Platform-for-Platforms Architecture Swyx and Erik discuss secure code execution sandboxes built on gVisor. Swyx connects this architectural evolution to platform-for-platforms patterns, referencing Cloudflare's Functions-as-a-Service model.48:54–53:27 · The hosts as informed peer 6/10 AI Inference Price Wars and Value in Custom Software Alessio highlights negative unit economics in commodity LLM inference price wars. Erik agrees that raw token serving has zero switching costs, comparing VC-subsidized inference to the late-90s fiber optics boom and defending Modal's focus on proprietary, custom pipelines.53:29–1:00:16 · The hosts as informed peer 6/10 High-IO Roadmaps, Oracle Cloud GPUs, and Hardware Philosophy Swyx brings up alternative perspectives on buying owned hardware and stretching depreciation schedules over seven years. Erik strongly rejects the premise for venture startups, citing the Waste Management accounting scandal where aggressive truck depreciation forced massive earnings restatements.1:00:17–1:08:32 · The hosts as informed peer 4/10 Competitive Programming Culture and Advice for Deep Tech Founders Alessio and Swyx explore Erik's background in competitive programming (IOI) and high engineering talent density at Modal. Erik delivers a passionate closing thesis criticizing startups that act as thin Kubernetes wrappers instead of diving deep into systems infrastructure.1:12–5:33 · Guest teaching 5/10 Early Career at Spotify, Luigi, and Vector Databases Swyx demonstrates familiarity with Bernhardsson's early career at Spotify, his open-source work on Luigi and Annoy, and vector indexing algorithms like HNSW. Bernhardsson explains the early pre-hype landscape of vector databases in 2012 and how Annoy used recursive hyperspace splitting.5:33–11:44 · Guest teaching 6/10 Scaling Better.com and Conceptualizing Modal's Compute Vision Erik details why existing cloud infrastructure and Kubernetes failed data teams due to burstiness, GPU management, and fragile dependency environments. Swyx interjects with domain context on the modern data stack and feedback loops.11:44–13:46 · Guest teaching 8/10 Re-Engineering Container Runtimes and Custom Network File Systems Erik educates the hosts on how Modal achieved sub-second container cold starts by bypassing standard Docker image pull mechanisms, using low-level runC primitives, and inventing a custom network-based virtual file system with deduplicated page caching.13:47–21:11 · Guest teaching 4/10 Self-Provisioning Runtimes and Python Developer Experience Swyx demonstrates deep domain expertise regarding self-provisioning runtimes, contrasting AWS CDK and SAM with Modal's unified application-and-infrastructure code model. Erik enthusiastically acknowledges Swyx's framing and discusses the trade-offs of embedding infrastructure directly in Python decorators.21:12–28:44 · Guest teaching 7/10 Serverless AI Inference Dynamics and GPU Utilization Alessio queries how serverless handles GPU utilization and batching compared to static memory allocation in foundational training. Erik explains how fast autoscaling and CPU/GPU memory checkpointing enable cost advantages despite higher per-hour GPU margins.28:44–35:20 · Guest teaching 5/10 Modal's Market Positioning: From Data Teams to Layer-Two Cloud Swyx challenges Erik's insistence on labeling Modal's target user as data teams rather than AI engineers, citing adoption data from the Vercel AI Accelerator. Swyx also compares Modal to Replicate and Modular, prompting Erik to articulate Modal's layer-two cloud positioning.35:20–38:00 · Guest teaching 6/10 Overcoming the Cloud Graduation Problem and Enterprise Retention Alessio asks about the enterprise graduation problem historically faced by PaaS providers like Heroku and how it applies to AI workloads. Erik explains why training lacks software differentiation compared to inference and how superior utilization economics retain customers.38:01–44:01 · Guest teaching 5/10 Language Model Workloads, Fine-Tuning, and Massive Parallelism The hosts discuss LLM workloads including fine-tuning and self-contained agents like ErikBot. Swyx brings up his experience using Modal to orchestrate concurrency limits and rate-limiting when building Small Developer.44:04–48:52 · Guest teaching 6/10 Secure Code Sandboxes and Platform-for-Platforms Architecture Swyx and Erik discuss secure code execution sandboxes built on gVisor. Swyx connects this architectural evolution to platform-for-platforms patterns, referencing Cloudflare's Functions-as-a-Service model.48:54–53:27 · Guest teaching 6/10 AI Inference Price Wars and Value in Custom Software Alessio highlights negative unit economics in commodity LLM inference price wars. Erik agrees that raw token serving has zero switching costs, comparing VC-subsidized inference to the late-90s fiber optics boom and defending Modal's focus on proprietary, custom pipelines.53:29–1:00:16 · Guest teaching 6/10 High-IO Roadmaps, Oracle Cloud GPUs, and Hardware Philosophy Swyx brings up alternative perspectives on buying owned hardware and stretching depreciation schedules over seven years. Erik strongly rejects the premise for venture startups, citing the Waste Management accounting scandal where aggressive truck depreciation forced massive earnings restatements.1:00:17–1:08:32 · Guest teaching 6/10 Competitive Programming Culture and Advice for Deep Tech Founders Alessio and Swyx explore Erik's background in competitive programming (IOI) and high engineering talent density at Modal. Erik delivers a passionate closing thesis criticizing startups that act as thin Kubernetes wrappers instead of diving deep into systems infrastructure.1:12–5:33 · Guest disagreement 1/10 Early Career at Spotify, Luigi, and Vector Databases Swyx demonstrates familiarity with Bernhardsson's early career at Spotify, his open-source work on Luigi and Annoy, and vector indexing algorithms like HNSW. Bernhardsson explains the early pre-hype landscape of vector databases in 2012 and how Annoy used recursive hyperspace splitting.5:33–11:44 · Guest disagreement 2/10 Scaling Better.com and Conceptualizing Modal's Compute Vision Erik details why existing cloud infrastructure and Kubernetes failed data teams due to burstiness, GPU management, and fragile dependency environments. Swyx interjects with domain context on the modern data stack and feedback loops.11:44–13:46 · Guest disagreement 1/10 Re-Engineering Container Runtimes and Custom Network File Systems Erik educates the hosts on how Modal achieved sub-second container cold starts by bypassing standard Docker image pull mechanisms, using low-level runC primitives, and inventing a custom network-based virtual file system with deduplicated page caching.13:47–21:11 · Guest disagreement 1/10 Self-Provisioning Runtimes and Python Developer Experience Swyx demonstrates deep domain expertise regarding self-provisioning runtimes, contrasting AWS CDK and SAM with Modal's unified application-and-infrastructure code model. Erik enthusiastically acknowledges Swyx's framing and discusses the trade-offs of embedding infrastructure directly in Python decorators.21:12–28:44 · Guest disagreement 2/10 Serverless AI Inference Dynamics and GPU Utilization Alessio queries how serverless handles GPU utilization and batching compared to static memory allocation in foundational training. Erik explains how fast autoscaling and CPU/GPU memory checkpointing enable cost advantages despite higher per-hour GPU margins.28:44–35:20 · Guest disagreement 3/10 Modal's Market Positioning: From Data Teams to Layer-Two Cloud Swyx challenges Erik's insistence on labeling Modal's target user as data teams rather than AI engineers, citing adoption data from the Vercel AI Accelerator. Swyx also compares Modal to Replicate and Modular, prompting Erik to articulate Modal's layer-two cloud positioning.35:20–38:00 · Guest disagreement 2/10 Overcoming the Cloud Graduation Problem and Enterprise Retention Alessio asks about the enterprise graduation problem historically faced by PaaS providers like Heroku and how it applies to AI workloads. Erik explains why training lacks software differentiation compared to inference and how superior utilization economics retain customers.38:01–44:01 · Guest disagreement 1/10 Language Model Workloads, Fine-Tuning, and Massive Parallelism The hosts discuss LLM workloads including fine-tuning and self-contained agents like ErikBot. Swyx brings up his experience using Modal to orchestrate concurrency limits and rate-limiting when building Small Developer.44:04–48:52 · Guest disagreement 1/10 Secure Code Sandboxes and Platform-for-Platforms Architecture Swyx and Erik discuss secure code execution sandboxes built on gVisor. Swyx connects this architectural evolution to platform-for-platforms patterns, referencing Cloudflare's Functions-as-a-Service model.48:54–53:27 · Guest disagreement 3/10 AI Inference Price Wars and Value in Custom Software Alessio highlights negative unit economics in commodity LLM inference price wars. Erik agrees that raw token serving has zero switching costs, comparing VC-subsidized inference to the late-90s fiber optics boom and defending Modal's focus on proprietary, custom pipelines.53:29–1:00:16 · Guest disagreement 4/10 High-IO Roadmaps, Oracle Cloud GPUs, and Hardware Philosophy Swyx brings up alternative perspectives on buying owned hardware and stretching depreciation schedules over seven years. Erik strongly rejects the premise for venture startups, citing the Waste Management accounting scandal where aggressive truck depreciation forced massive earnings restatements.1:00:17–1:08:32 · Guest disagreement 3/10 Competitive Programming Culture and Advice for Deep Tech Founders Alessio and Swyx explore Erik's background in competitive programming (IOI) and high engineering talent density at Modal. Erik delivers a passionate closing thesis criticizing startups that act as thin Kubernetes wrappers instead of diving deep into systems infrastructure.1:12–5:33 · The hosts pushing back 1/10 Early Career at Spotify, Luigi, and Vector Databases Swyx demonstrates familiarity with Bernhardsson's early career at Spotify, his open-source work on Luigi and Annoy, and vector indexing algorithms like HNSW. Bernhardsson explains the early pre-hype landscape of vector databases in 2012 and how Annoy used recursive hyperspace splitting.5:33–11:44 · The hosts pushing back 1/10 Scaling Better.com and Conceptualizing Modal's Compute Vision Erik details why existing cloud infrastructure and Kubernetes failed data teams due to burstiness, GPU management, and fragile dependency environments. Swyx interjects with domain context on the modern data stack and feedback loops.11:44–13:46 · The hosts pushing back 0/10 Re-Engineering Container Runtimes and Custom Network File Systems Erik educates the hosts on how Modal achieved sub-second container cold starts by bypassing standard Docker image pull mechanisms, using low-level runC primitives, and inventing a custom network-based virtual file system with deduplicated page caching.13:47–21:11 · The hosts pushing back 2/10 Self-Provisioning Runtimes and Python Developer Experience Swyx demonstrates deep domain expertise regarding self-provisioning runtimes, contrasting AWS CDK and SAM with Modal's unified application-and-infrastructure code model. Erik enthusiastically acknowledges Swyx's framing and discusses the trade-offs of embedding infrastructure directly in Python decorators.21:12–28:44 · The hosts pushing back 1/10 Serverless AI Inference Dynamics and GPU Utilization Alessio queries how serverless handles GPU utilization and batching compared to static memory allocation in foundational training. Erik explains how fast autoscaling and CPU/GPU memory checkpointing enable cost advantages despite higher per-hour GPU margins.28:44–35:20 · The hosts pushing back 5/10 Modal's Market Positioning: From Data Teams to Layer-Two Cloud Swyx challenges Erik's insistence on labeling Modal's target user as data teams rather than AI engineers, citing adoption data from the Vercel AI Accelerator. Swyx also compares Modal to Replicate and Modular, prompting Erik to articulate Modal's layer-two cloud positioning.35:20–38:00 · The hosts pushing back 1/10 Overcoming the Cloud Graduation Problem and Enterprise Retention Alessio asks about the enterprise graduation problem historically faced by PaaS providers like Heroku and how it applies to AI workloads. Erik explains why training lacks software differentiation compared to inference and how superior utilization economics retain customers.38:01–44:01 · The hosts pushing back 1/10 Language Model Workloads, Fine-Tuning, and Massive Parallelism The hosts discuss LLM workloads including fine-tuning and self-contained agents like ErikBot. Swyx brings up his experience using Modal to orchestrate concurrency limits and rate-limiting when building Small Developer.44:04–48:52 · The hosts pushing back 1/10 Secure Code Sandboxes and Platform-for-Platforms Architecture Swyx and Erik discuss secure code execution sandboxes built on gVisor. Swyx connects this architectural evolution to platform-for-platforms patterns, referencing Cloudflare's Functions-as-a-Service model.48:54–53:27 · The hosts pushing back 2/10 AI Inference Price Wars and Value in Custom Software Alessio highlights negative unit economics in commodity LLM inference price wars. Erik agrees that raw token serving has zero switching costs, comparing VC-subsidized inference to the late-90s fiber optics boom and defending Modal's focus on proprietary, custom pipelines.53:29–1:00:16 · The hosts pushing back 3/10 High-IO Roadmaps, Oracle Cloud GPUs, and Hardware Philosophy Swyx brings up alternative perspectives on buying owned hardware and stretching depreciation schedules over seven years. Erik strongly rejects the premise for venture startups, citing the Waste Management accounting scandal where aggressive truck depreciation forced massive earnings restatements.1:00:17–1:08:32 · The hosts pushing back 0/10 Competitive Programming Culture and Advice for Deep Tech Founders Alessio and Swyx explore Erik's background in competitive programming (IOI) and high engineering talent density at Modal. Erik delivers a passionate closing thesis criticizing startups that act as thin Kubernetes wrappers instead of diving deep into systems infrastructure.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.5% · guest 64.5%0:00 · the hosts 35.5% · guest 64.5%3:00 · the hosts 15% · guest 85%3:00 · the hosts 15% · guest 85%6:00 · the hosts 20% · guest 80%6:00 · the hosts 20% · guest 80%9:00 · the hosts 6.1% · guest 93.9%9:00 · the hosts 6.1% · guest 93.9%12:00 · the hosts 16.1% · guest 83.9%12:00 · the hosts 16.1% · guest 83.9%15:00 · the hosts 29.5% · guest 70.5%15:00 · the hosts 29.5% · guest 70.5%18:00 · the hosts 44.1% · guest 55.9%18:00 · the hosts 44.1% · guest 55.9%21:00 · the hosts 25.8% · guest 74.2%21:00 · the hosts 25.8% · guest 74.2%24:00 · the hosts 25.2% · guest 74.8%24:00 · the hosts 25.2% · guest 74.8%27:00 · the hosts 29.8% · guest 70.2%27:00 · the hosts 29.8% · guest 70.2%30:00 · the hosts 32.1% · guest 67.9%30:00 · the hosts 32.1% · guest 67.9%33:00 · the hosts 42% · guest 58%33:00 · the hosts 42% · guest 58%36:00 · the hosts 15.3% · guest 84.7%36:00 · the hosts 15.3% · guest 84.7%39:00 · the hosts 31.2% · guest 68.8%39:00 · the hosts 31.2% · guest 68.8%42:00 · the hosts 57.4% · guest 42.6%42:00 · the hosts 57.4% · guest 42.6%45:00 · the hosts 36.3% · guest 63.7%45:00 · the hosts 36.3% · guest 63.7%48:00 · the hosts 33.8% · guest 66.2%48:00 · the hosts 33.8% · guest 66.2%51:00 · the hosts 21.8% · guest 78.2%51:00 · the hosts 21.8% · guest 78.2%54:00 · the hosts 12.6% · guest 87.4%54:00 · the hosts 12.6% · guest 87.4%57:00 · the hosts 32.8% · guest 67.2%57:00 · the hosts 32.8% · guest 67.2%1:00:00 · the hosts 24.5% · guest 75.5%1:00:00 · the hosts 24.5% · guest 75.5%1:03:00 · the hosts 37.8% · guest 62.2%1:03:00 · the hosts 37.8% · guest 62.2%1:06:00 · the hosts 26.8% · guest 73.2%1:06:00 · the hosts 26.8% · guest 73.2%
Sharpest disagreement ▶ 59:42 Erik mocks extended hardware depreciation schedules

When Swyx brings up CTOs who claim they can make GPUs last seven years to justify buying hardware, Erik aggressively refutes the framing by comparing it to the Waste Management accounting fraud scandal where overextended truck depreciation led to restatements.

Hardest push from the hosts ▶ 28:44 Swyx rejects Modal's data teams positioning

Swyx openly challenges Erik's persistent marketing of Modal as built for data teams, arguing that AI engineers represent the actual user base and citing empirical proof from the Vercel AI Accelerator.

Biggest teaching moment ▶ 11:54 Erik explains low-level container runtime bypasses

Erik explains step-by-step why Docker image pulls are wasteful and details how building a custom network virtual file system hooked directly to runC allows instant container cold starts without downloading multi-gigabyte layers.

The host holds their own ▶ 16:47 Swyx explains the necessity of self-provisioning runtimes

Swyx lays out his thesis on self-provisioning runtimes, contrasting the awkward compile steps of AWS CDK and CloudFormation with true infrastructure-application convergence and programming language ergonomics.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Early Career at Spotify, Luigi, and Vector Databases 6511 Swyx demonstrates familiarity with Bernhardsson's early career at Spotify, his open-source work on Luigi and Annoy, and vector indexing algorithms like HNSW. Bernhardsson explains the early pre-hype landscape of vector databases in 2012 and how Annoy used recursive hyperspace splitting.
Scaling Better.com and Conceptualizing Modal's Compute Vision 5621 Erik details why existing cloud infrastructure and Kubernetes failed data teams due to burstiness, GPU management, and fragile dependency environments. Swyx interjects with domain context on the modern data stack and feedback loops.
Re-Engineering Container Runtimes and Custom Network File Systems 4810 Erik educates the hosts on how Modal achieved sub-second container cold starts by bypassing standard Docker image pull mechanisms, using low-level runC primitives, and inventing a custom network-based virtual file system with deduplicated page caching.
Self-Provisioning Runtimes and Python Developer Experience 8412 Swyx demonstrates deep domain expertise regarding self-provisioning runtimes, contrasting AWS CDK and SAM with Modal's unified application-and-infrastructure code model. Erik enthusiastically acknowledges Swyx's framing and discusses the trade-offs of embedding infrastructure directly in Python decorators.
Serverless AI Inference Dynamics and GPU Utilization 6721 Alessio queries how serverless handles GPU utilization and batching compared to static memory allocation in foundational training. Erik explains how fast autoscaling and CPU/GPU memory checkpointing enable cost advantages despite higher per-hour GPU margins.
Modal's Market Positioning: From Data Teams to Layer-Two Cloud 7535 Swyx challenges Erik's insistence on labeling Modal's target user as data teams rather than AI engineers, citing adoption data from the Vercel AI Accelerator. Swyx also compares Modal to Replicate and Modular, prompting Erik to articulate Modal's layer-two cloud positioning.
Overcoming the Cloud Graduation Problem and Enterprise Retention 5621 Alessio asks about the enterprise graduation problem historically faced by PaaS providers like Heroku and how it applies to AI workloads. Erik explains why training lacks software differentiation compared to inference and how superior utilization economics retain customers.
Language Model Workloads, Fine-Tuning, and Massive Parallelism 6511 The hosts discuss LLM workloads including fine-tuning and self-contained agents like ErikBot. Swyx brings up his experience using Modal to orchestrate concurrency limits and rate-limiting when building Small Developer.
Secure Code Sandboxes and Platform-for-Platforms Architecture 6611 Swyx and Erik discuss secure code execution sandboxes built on gVisor. Swyx connects this architectural evolution to platform-for-platforms patterns, referencing Cloudflare's Functions-as-a-Service model.
AI Inference Price Wars and Value in Custom Software 6632 Alessio highlights negative unit economics in commodity LLM inference price wars. Erik agrees that raw token serving has zero switching costs, comparing VC-subsidized inference to the late-90s fiber optics boom and defending Modal's focus on proprietary, custom pipelines.
High-IO Roadmaps, Oracle Cloud GPUs, and Hardware Philosophy 6643 Swyx brings up alternative perspectives on buying owned hardware and stretching depreciation schedules over seven years. Erik strongly rejects the premise for venture startups, citing the Waste Management accounting scandal where aggressive truck depreciation forced massive earnings restatements.
Competitive Programming Culture and Advice for Deep Tech Founders 4630 Alessio and Swyx explore Erik's background in competitive programming (IOI) and high engineering talent density at Modal. Erik delivers a passionate closing thesis criticizing startups that act as thin Kubernetes wrappers instead of diving deep into systems infrastructure.

Statements from this episode (31)

Opinion
Daniel Ek Is a Generational CEO Among Swedish Startups
“Overall, I mean, I think he was a great CEO, like, definitely, you know, up there, like, generational CEO, at least for, like, Swedish startups.”
Erik Bernhardsson Feb 19, 2024 ▶ 5:17
Assertion Supported
Better.com Grew to 10,000 Employees Before Shrinking Back to 1,000
“Yeah, I mean, the company, like, grew from, like, 10 people when I joined to 10,000, now it's back to a thousand, and, but yeah, they actually went public a few months ago, kind of crazy.”
Erik Bernhardsson Feb 19, 2024 ▶ 6:01
Insight
Developer Productivity Is Best Measured by Iteration Feedback Speed
“I think the best way to measure developer productivity is, like, in terms of the feedback loops. Like, it, like, in, like, How quickly when you iterate, like when you write code, like how quickly can you get feedback? And at the innermost loop, it's like runni…”
Erik Bernhardsson Feb 19, 2024 ▶ 9:22
Opinion
Kubernetes Works for Backend Teams but Fails Data Teams
“Like in particular, like Kubernetes, I feel like it's like kind of worked okay for backend teams, but not so well for data teams.”
Erik Bernhardsson Feb 19, 2024 ▶ 10:03
Insight
99% of Data in Multi-Gigabyte Docker Images Is Never Read
“It turns out, like, when you start a Docker container, like, first of all, like, most Docker images are, like, several gigabytes, and, like, 99% of that is never going to be consumed.”
Erik Bernhardsson Feb 19, 2024 ▶ 13:04
Assertion Supported
Modal's Custom Container File System Enables Sub-Second Cloud Launches
“So that was, like, the first, like, sort of stuff we started working on was, like, let's build this, like, container file system, And, you know, coupled with like, you know, just using RunC directly. And that actually enabled us to like get to this point of li…”
Erik Bernhardsson Feb 19, 2024 ▶ 13:25
Opinion
AWS Has Always Struggled With Developer Experience
“AWS has always like struggled with developer experience”
Erik Bernhardsson Feb 19, 2024 ▶ 17:35
Prediction Not checkable as stated
Developer Productivity Will Continue Growing 10x Per Decade
“You look at the developers are, I've been getting like probably 10 X more productive every decade for the last four, four decades or something that was kind of crazy. Like on an exponential scale, we're talking about 10 X or is there a 10,000 X like, you know,…”
Erik Bernhardsson Feb 19, 2024 ▶ 20:39
Insight
Serverless Architecture Fits Generative AI Far Better Than Backend Services
“Serverless always, to me, felt like a little bit of, like, a solution looking for a problem. I don't know, I don't actually, like, don't think, like, backend is, like, the problem that needs serverless, or, like, not as much, but I look at data, and in particu…”
Erik Bernhardsson Feb 19, 2024 ▶ 22:27
Disclosure
Modal Had Little Traction Until the Release of Stable Diffusion
“That's, like, you know, a year and a half in, like, we barely had any users or any revenue, and, like, we were, like, well, maybe we should look at, like, some use case, trying to think of use cases, and that was around the time, same, same time Stable Edition…”
Erik Bernhardsson Feb 19, 2024 ▶ 23:10
Assertion Not checkable as stated
Modal Users Save Money on Inference Despite Higher GPU Rates
“In many cases, like users who do inference on those platforms or those clouds even though we charge a slightly higher price per, per GPU hour, a lot of users like moving their large scale inference use cases to model, like end up saving a lot of money because …”
Erik Bernhardsson Feb 19, 2024 ▶ 25:47
Disclosure
Modal Implemented Full CPU Memory Checkpointing for Container Restores
“We've implemented recently CPU memory checkpointing, so we can take running containers and snapshot the entire CPU, like including registers and everything, and restore it from that point, which means we can restore it from a like initialized state.”
Erik Bernhardsson Feb 19, 2024 ▶ 26:26
Disclosure
Modal Is Not Designed for Training Large Foundation Models
“When you get to like training, like very large foundational models, that that's a use case we don't support super well because that's very high IO, you know, you need to have like infinite band and all these things. And those are things we haven't supported ye…”
Erik Bernhardsson Feb 19, 2024 ▶ 27:44
Assertion Not checkable as stated
Modal Was the Top Tool Used in Vercel's AI Accelerator
“And in the Versailles AI accelerator, a bunch of startups gave like free credits and like signups and talks and all that stuff. The only ones that stuck are the people, are the ones that actually appealed to engineers and the top usage, the top tool used by fa…”
Shawn Wang Feb 19, 2024 ▶ 29:10
Assertion Supported
Suno Uses Modal for Its AI Music Production Infrastructure
“Yeah, so, I mean, they're using model for, like, production infrastructure, like, they have their own, like, custom model, like, custom code and custom weights, you know, for AI-generated music, Suno.ai.”
Erik Bernhardsson Feb 19, 2024 ▶ 32:22
Prediction Not checkable as stated
VCs Will Eventually Force Modular to Offer a Hosted Cloud Service
“I'm sure their VCs at some point are gonna force them to reconsider.”
Erik Bernhardsson Feb 19, 2024 ▶ 33:58
Insight
Snowflake Proved Layer-Two Clouds Can Win on Developer Experience
“And I think Snowflake, you know, to me won because like, I mean, in the end, like AWS makes all the money anyway, like, and like Snowflake just had the ability to like focus on like developer experience or like, you know, or user experience. And to me, like re…”
Erik Bernhardsson Feb 19, 2024 ▶ 34:45
Insight
Cloud Infra Startups Must Capture the Entire Hobbyist-to-Enterprise Spectrum
“The only way to do, to build infrastructure to build a successful infrastructure company in the long run in, in, in the cloud today is, You have to appeal to the entire spectrum, right? Or at least like the enterprise, like you have to capture the enterprise m…”
Erik Bernhardsson Feb 19, 2024 ▶ 36:06
Opinion
Buying Dedicated GPU Clusters Offers Better Economics for Model Training
“In training, I think, you know, there's less software differentiation. So in training, I think there's certainly, like, better economics of, like, buying big clusters.”
Erik Bernhardsson Feb 19, 2024 ▶ 37:25
Prediction Not checkable as stated
Only Tech Giants Will Build In-House AI Infrastructure Long-Term
“I think a lot of these companies over in the long run, like, you know, they're accepted maybe super big ones, like, you know, the Facebook and Google, they're always gonna build their own ones, but like everyone else, like some extent, you know, I think they'r…”
Erik Bernhardsson Feb 19, 2024 ▶ 37:42
Disclosure
Modal Will Not Try to Compete Directly With OpenAI's API
“So many people get started with APIs and that's just, you know, they're just dominating a space in particular, open AI. Right. And that's not necessarily like a place where we aim to compete. I mean, maybe at some point, but like, it's just not like a core foc…”
Erik Bernhardsson Feb 19, 2024 ▶ 41:02
Assertion Supported
Modal Can Fan Out Workloads to Thousands of GPUs in Minutes
“It's like pretty easy in Moto, like, fan out to like, you know, at least like a hundred GPUs, like in a few seconds, and you know, if you give it like a couple of minutes, like we can, you know, you can fan out to like thousands of GPUs.”
Erik Bernhardsson Feb 19, 2024 ▶ 42:23
Disclosure
Modal's Sandbox Feature Is Evolving Into a Platform for Platforms
“The core, like, container infrastructure we offer could actually be, like, you know, unbundled from, like, the client SDK and offered to, like, other, you know, like, we're talking to a couple of, like other companies that want to run, you know, through their,…”
Erik Bernhardsson Feb 19, 2024 ▶ 44:50
Insight
Switching Costs for Open-Source LLM Inference Providers Are Zero
“The LLM space is, like, the opposite. Like, the switching cost of LLMs is zero, right? Like, if all you're doing is, like, straight up, like, at least, like, open source, right? Like, if all you're doing is, like, you know, using some, you know, inference endp…”
Erik Bernhardsson Feb 19, 2024 ▶ 49:55
Prediction Not checkable as stated
VCs Subsidizing AI Inference Will Not See Their Expected Returns
“In the end, like, I don't think VCs will have the return they expected. Like, you know, in these things, but guess who's going to benefit? Like, you know, it's the consumers, right? Like someone's like reaping that the value of this. And that's, I think an ama…”
Erik Bernhardsson Feb 19, 2024 ▶ 50:15
Prediction Not checkable as stated
AI Inference Margins Will Rise While Consumer Prices Stay Low
“Margins are gonna go up for sure, but I don't know if prices will go up, because like, GPU prices have to drop eventually, right? So like, you know, like in the long run, I still think like prices may not go up that much but certainly margins will go up.”
Erik Bernhardsson Feb 19, 2024 ▶ 52:56
Opinion
Oracle Cloud GPUs Actually Offer Great Value for Money
“I love Oracle's GPUs. I mean, I don't, like, I don't know why, you know, what the economics looks like for Oracle, but like, I think they're great value for money. Like, we run a bunch of stuff in Oracle and They have bare metal machines with, like, two teraby…”
Erik Bernhardsson Feb 19, 2024 ▶ 55:03
Insight
Buying Physical Hardware Is a Poor Use of Venture Capital
“My feeling is that when you're a venture-funded startup, like, buying physical hardware is maybe not the best use of the money.”
Erik Bernhardsson Feb 19, 2024 ▶ 57:50
Insight
AI Productivity Gains Will Ultimately Increase Demand for Software Engineers
“Like, I think we can, you know, 10 x the amount of developers, and still, you know, have a lot of people making a lot of money, you know, building amazing software, and also being, while at the same time being more productive. Like, I never understood this, li…”
Erik Bernhardsson Feb 19, 2024 ▶ 1:01:23
Assertion Not checkable as stated
Modal Uses Mixed Integer Programming to Optimize Cloud Resource Allocation
“You know, the resource allocation, like turns out like that actually, like you can phrase that as a mixed integer programming problem. Like we now have that running in production, like constantly optimizing how we allocate cloud research.”
Erik Bernhardsson Feb 19, 2024 ▶ 1:03:41
Insight
Startups Must Build Deep Infrastructure Rather Than Thin Kubernetes Wrappers
“And so one of my frustrations has been like so many startups are like, in my opinion, like Kubernetes wrappers and like, you know, like, and not very like thick wrappers, like fairly thin wrappers. And I think, you know, every startup is a wrapper to some exte…”
Erik Bernhardsson Feb 19, 2024 ▶ 1:07:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.