Jul 8, 2026 · 59m · latent-space

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO

Akshat Bubna · 30m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this interview, Modal CTO Akshat Bubna explains how Modal engineered a high-performance, serverless cloud platform tailored for elastic AI inference, fine-tuning, and autonomous agent sandboxes. He details the transition from Kubernetes-based developer workflows to agent-optimized infrastructure, highlighting innovations in GPU snapshotting, speculative decoding, and low-level kernel security.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.6 Guest teaching 4.1 Guest disagreement 1.3 The hosts pushing back 2.1
05100:0015:0030:0045:001:18–4:28 · The hosts as informed peer 6/10 Modal's Origin Story and Escaping Kubernetes Limitations The host recalls Modal's early serverless container stack presentations from Data Council and questions whether data pipelines genuinely needed that level of infrastructure. Akshat explains the evolution from serverless primitives to bursty ML inference and computer vision workloads.4:29–7:49 · The hosts as informed peer 6/10 Self-Provisioning Infrastructure and Decorators Over YAML The host brings up his own past post on self-provisioning infrastructure and playfully notes the vendor lock-in pushback around custom DSLs. Akshat reframes the value around code ownership and how self-provisioning runtime benefits both human DX and agent AX.7:50–12:07 · The hosts as informed peer 6/10 Defining Modal's Purpose: Cloud Primitives for AI Applications Hosts ask Akshat to define Modal's core identity beyond general compute, citing their early direct experience running Small Developer. Akshat outlines Modal as scratch-built cloud primitives specifically optimized for AI applications rather than static web servers.12:07–15:52 · The hosts as informed peer 5/10 Handling Massive Scale and Elastic RL Training Runs Vibhu probes into rapid RL training run patterns like Cursor Composer scaling up and down hourly. Akshat clarifies the distinct workload profiles between diurnal custom inference, batch pre-training encoders, and massive 100k sandbox RL rollouts.15:52–21:13 · The hosts as informed peer 6/10 Continual Learning and Upstreaming Open-Source LLM Advancements The host and guest discuss speculative decoding, with the host asking for technical clarification on draft models versus compute overhead. Akshat breaks down the mathematics of memory bandwidth bounds versus multi-token accept length multiplicative speedups in dFlash.21:13–23:15 · The hosts as informed peer 6/10 The Modal Infrastructure Delta and Production-Grade Reliability Vibhu asks what technical delta Modal offers over running open-source engines like SGLang directly on raw compute. Akshat explains that Modal actively upstreams performance work to SGLang while offering zero-scaling elasticity and production-grade tail latency management.23:16–27:29 · The hosts as informed peer 6/10 Supporting the Complete Model Life Cycle and Ramp Inspect The host notes the shifting CPU-to-GPU workload ratio from 8:1 to 1:1 in agentic workloads and asks about multi-cloud arbitrage. Akshat explains Modal's software-defined reliability layer aggregating over 17 distinct neocloud providers without owning physical data centers.27:30–31:02 · The hosts as informed peer 5/10 Advanced Agent Networking, Sidecars, and Private IPv6 Overlays The host queries Modal's emerging networking capabilities, drawing comparisons to Docker Compose. Akshat reveals unpublished features like multi-container sidecars and their ISXPN private IPv6 overlay network originally built for distributed training.31:02–34:55 · The hosts as informed peer 5/10 Inter-GPU Memory Management, RDMA, and eBPF Network Security The host mixes up WireGuard encryption with Modal's kernel-level TCP overlay and questions whether cross-node training networking is fast enough. Akshat politely disentangles eBPF access control and clarifies their internal 3 Tbps RDMA interconnect for medium post-training workloads.34:56–37:38 · The hosts as informed peer 5/10 Autonomous Research, Auto Inference, and Parameter Sweeps The host characterizes auto-research as mostly science-fair hype and asks if it is merely hyperparameter sweeps. Akshat reveals Modal's internal 'auto inference' harness that autonomously runs NVIDIA Nsight profilers and tunes configurations to replace manual forward-deployed engineering.37:38–41:49 · The hosts as informed peer 5/10 Evaluating LLM-Generated Modal Code and Building Modal Bench The host inquires about LLM code generation reliability for Modal and hardware supply shortages. Akshat discusses developing Modal Bench and explains compute strategy roles that model fungible multi-year reservations and capacity pricing tiers.41:50–45:13 · The hosts as informed peer 6/10 Future Product Roadmap: Real-Time Media and Agent Ecosystems The host pushes a provocative hypothesis that an LLM should act directly as the OS kernel with adaptive permissions. Akshat firmly pushes back, maintaining that security at the sandbox level strictly requires deterministic hard boundaries rather than purely LLM-mediated guardrails.45:15–47:34 · The hosts as informed peer 5/10 Managed Agent Ecosystem vs Dedicated Sandbox Infrastructure The host asks whether managed agent suites from frontier labs threaten Modal's position. Akshat explains that while labs provide great starting points, production enterprise agents like Ramp require granular control over persistent storage, GPU sandboxes, and snapshotting.47:34–50:40 · The hosts as informed peer 6/10 Code-First Infrastructure vs Commodity Model APIs The host brings up Replicate and asks for a postmortem on why model API marketplaces struggle. Akshat points out that raw model APIs cater to transient hobbyist traffic, whereas Modal provides programmatic, customizable code-level infrastructure for enterprise products.50:40–53:00 · The hosts as informed peer 6/10 Differentiated AI Products and Video Orchestration Agents The host cites predictions from xAI team members about video generation moving toward multi-model orchestration agents. Akshat validates this with real customer behavior where agents execute complex FFmpeg workflows inside GPU sandboxes.53:01–55:40 · The hosts as informed peer 6/10 The Evolution of CI/CD for Coding Agents and Runtime Sandboxes The host discusses Gitpod/Ona joining OpenAI and queries the architecture of agent CI/CD. Akshat details how memory snapshot-and-restore primitives can eliminate dependency preparation bottlenecks and contrasts build-time with runtime sandboxes.55:41–57:02 · The hosts as informed peer 5/10 Multi-Language SDKs and Programming Paradigms in the AI Era The host asks if DX and AX are nearly identical paradigms. Akshat explains that while conceptually aligned, AX requires concrete product changes like porting UI observability metrics into machine-parseable CLI commands and benchmark-driven feature additions.1:18–4:28 · Guest teaching 4/10 Modal's Origin Story and Escaping Kubernetes Limitations The host recalls Modal's early serverless container stack presentations from Data Council and questions whether data pipelines genuinely needed that level of infrastructure. Akshat explains the evolution from serverless primitives to bursty ML inference and computer vision workloads.4:29–7:49 · Guest teaching 3/10 Self-Provisioning Infrastructure and Decorators Over YAML The host brings up his own past post on self-provisioning infrastructure and playfully notes the vendor lock-in pushback around custom DSLs. Akshat reframes the value around code ownership and how self-provisioning runtime benefits both human DX and agent AX.7:50–12:07 · Guest teaching 3/10 Defining Modal's Purpose: Cloud Primitives for AI Applications Hosts ask Akshat to define Modal's core identity beyond general compute, citing their early direct experience running Small Developer. Akshat outlines Modal as scratch-built cloud primitives specifically optimized for AI applications rather than static web servers.12:07–15:52 · Guest teaching 5/10 Handling Massive Scale and Elastic RL Training Runs Vibhu probes into rapid RL training run patterns like Cursor Composer scaling up and down hourly. Akshat clarifies the distinct workload profiles between diurnal custom inference, batch pre-training encoders, and massive 100k sandbox RL rollouts.15:52–21:13 · Guest teaching 5/10 Continual Learning and Upstreaming Open-Source LLM Advancements The host and guest discuss speculative decoding, with the host asking for technical clarification on draft models versus compute overhead. Akshat breaks down the mathematics of memory bandwidth bounds versus multi-token accept length multiplicative speedups in dFlash.21:13–23:15 · Guest teaching 4/10 The Modal Infrastructure Delta and Production-Grade Reliability Vibhu asks what technical delta Modal offers over running open-source engines like SGLang directly on raw compute. Akshat explains that Modal actively upstreams performance work to SGLang while offering zero-scaling elasticity and production-grade tail latency management.23:16–27:29 · Guest teaching 4/10 Supporting the Complete Model Life Cycle and Ramp Inspect The host notes the shifting CPU-to-GPU workload ratio from 8:1 to 1:1 in agentic workloads and asks about multi-cloud arbitrage. Akshat explains Modal's software-defined reliability layer aggregating over 17 distinct neocloud providers without owning physical data centers.27:30–31:02 · Guest teaching 5/10 Advanced Agent Networking, Sidecars, and Private IPv6 Overlays The host queries Modal's emerging networking capabilities, drawing comparisons to Docker Compose. Akshat reveals unpublished features like multi-container sidecars and their ISXPN private IPv6 overlay network originally built for distributed training.31:02–34:55 · Guest teaching 6/10 Inter-GPU Memory Management, RDMA, and eBPF Network Security The host mixes up WireGuard encryption with Modal's kernel-level TCP overlay and questions whether cross-node training networking is fast enough. Akshat politely disentangles eBPF access control and clarifies their internal 3 Tbps RDMA interconnect for medium post-training workloads.34:56–37:38 · Guest teaching 4/10 Autonomous Research, Auto Inference, and Parameter Sweeps The host characterizes auto-research as mostly science-fair hype and asks if it is merely hyperparameter sweeps. Akshat reveals Modal's internal 'auto inference' harness that autonomously runs NVIDIA Nsight profilers and tunes configurations to replace manual forward-deployed engineering.37:38–41:49 · Guest teaching 4/10 Evaluating LLM-Generated Modal Code and Building Modal Bench The host inquires about LLM code generation reliability for Modal and hardware supply shortages. Akshat discusses developing Modal Bench and explains compute strategy roles that model fungible multi-year reservations and capacity pricing tiers.41:50–45:13 · Guest teaching 4/10 Future Product Roadmap: Real-Time Media and Agent Ecosystems The host pushes a provocative hypothesis that an LLM should act directly as the OS kernel with adaptive permissions. Akshat firmly pushes back, maintaining that security at the sandbox level strictly requires deterministic hard boundaries rather than purely LLM-mediated guardrails.45:15–47:34 · Guest teaching 4/10 Managed Agent Ecosystem vs Dedicated Sandbox Infrastructure The host asks whether managed agent suites from frontier labs threaten Modal's position. Akshat explains that while labs provide great starting points, production enterprise agents like Ramp require granular control over persistent storage, GPU sandboxes, and snapshotting.47:34–50:40 · Guest teaching 4/10 Code-First Infrastructure vs Commodity Model APIs The host brings up Replicate and asks for a postmortem on why model API marketplaces struggle. Akshat points out that raw model APIs cater to transient hobbyist traffic, whereas Modal provides programmatic, customizable code-level infrastructure for enterprise products.50:40–53:00 · Guest teaching 3/10 Differentiated AI Products and Video Orchestration Agents The host cites predictions from xAI team members about video generation moving toward multi-model orchestration agents. Akshat validates this with real customer behavior where agents execute complex FFmpeg workflows inside GPU sandboxes.53:01–55:40 · Guest teaching 4/10 The Evolution of CI/CD for Coding Agents and Runtime Sandboxes The host discusses Gitpod/Ona joining OpenAI and queries the architecture of agent CI/CD. Akshat details how memory snapshot-and-restore primitives can eliminate dependency preparation bottlenecks and contrasts build-time with runtime sandboxes.55:41–57:02 · Guest teaching 4/10 Multi-Language SDKs and Programming Paradigms in the AI Era The host asks if DX and AX are nearly identical paradigms. Akshat explains that while conceptually aligned, AX requires concrete product changes like porting UI observability metrics into machine-parseable CLI commands and benchmark-driven feature additions.1:18–4:28 · Guest disagreement 1/10 Modal's Origin Story and Escaping Kubernetes Limitations The host recalls Modal's early serverless container stack presentations from Data Council and questions whether data pipelines genuinely needed that level of infrastructure. Akshat explains the evolution from serverless primitives to bursty ML inference and computer vision workloads.4:29–7:49 · Guest disagreement 2/10 Self-Provisioning Infrastructure and Decorators Over YAML The host brings up his own past post on self-provisioning infrastructure and playfully notes the vendor lock-in pushback around custom DSLs. Akshat reframes the value around code ownership and how self-provisioning runtime benefits both human DX and agent AX.7:50–12:07 · Guest disagreement 1/10 Defining Modal's Purpose: Cloud Primitives for AI Applications Hosts ask Akshat to define Modal's core identity beyond general compute, citing their early direct experience running Small Developer. Akshat outlines Modal as scratch-built cloud primitives specifically optimized for AI applications rather than static web servers.12:07–15:52 · Guest disagreement 2/10 Handling Massive Scale and Elastic RL Training Runs Vibhu probes into rapid RL training run patterns like Cursor Composer scaling up and down hourly. Akshat clarifies the distinct workload profiles between diurnal custom inference, batch pre-training encoders, and massive 100k sandbox RL rollouts.15:52–21:13 · Guest disagreement 1/10 Continual Learning and Upstreaming Open-Source LLM Advancements The host and guest discuss speculative decoding, with the host asking for technical clarification on draft models versus compute overhead. Akshat breaks down the mathematics of memory bandwidth bounds versus multi-token accept length multiplicative speedups in dFlash.21:13–23:15 · Guest disagreement 1/10 The Modal Infrastructure Delta and Production-Grade Reliability Vibhu asks what technical delta Modal offers over running open-source engines like SGLang directly on raw compute. Akshat explains that Modal actively upstreams performance work to SGLang while offering zero-scaling elasticity and production-grade tail latency management.23:16–27:29 · Guest disagreement 1/10 Supporting the Complete Model Life Cycle and Ramp Inspect The host notes the shifting CPU-to-GPU workload ratio from 8:1 to 1:1 in agentic workloads and asks about multi-cloud arbitrage. Akshat explains Modal's software-defined reliability layer aggregating over 17 distinct neocloud providers without owning physical data centers.27:30–31:02 · Guest disagreement 1/10 Advanced Agent Networking, Sidecars, and Private IPv6 Overlays The host queries Modal's emerging networking capabilities, drawing comparisons to Docker Compose. Akshat reveals unpublished features like multi-container sidecars and their ISXPN private IPv6 overlay network originally built for distributed training.31:02–34:55 · Guest disagreement 2/10 Inter-GPU Memory Management, RDMA, and eBPF Network Security The host mixes up WireGuard encryption with Modal's kernel-level TCP overlay and questions whether cross-node training networking is fast enough. Akshat politely disentangles eBPF access control and clarifies their internal 3 Tbps RDMA interconnect for medium post-training workloads.34:56–37:38 · Guest disagreement 1/10 Autonomous Research, Auto Inference, and Parameter Sweeps The host characterizes auto-research as mostly science-fair hype and asks if it is merely hyperparameter sweeps. Akshat reveals Modal's internal 'auto inference' harness that autonomously runs NVIDIA Nsight profilers and tunes configurations to replace manual forward-deployed engineering.37:38–41:49 · Guest disagreement 1/10 Evaluating LLM-Generated Modal Code and Building Modal Bench The host inquires about LLM code generation reliability for Modal and hardware supply shortages. Akshat discusses developing Modal Bench and explains compute strategy roles that model fungible multi-year reservations and capacity pricing tiers.41:50–45:13 · Guest disagreement 2/10 Future Product Roadmap: Real-Time Media and Agent Ecosystems The host pushes a provocative hypothesis that an LLM should act directly as the OS kernel with adaptive permissions. Akshat firmly pushes back, maintaining that security at the sandbox level strictly requires deterministic hard boundaries rather than purely LLM-mediated guardrails.45:15–47:34 · Guest disagreement 1/10 Managed Agent Ecosystem vs Dedicated Sandbox Infrastructure The host asks whether managed agent suites from frontier labs threaten Modal's position. Akshat explains that while labs provide great starting points, production enterprise agents like Ramp require granular control over persistent storage, GPU sandboxes, and snapshotting.47:34–50:40 · Guest disagreement 2/10 Code-First Infrastructure vs Commodity Model APIs The host brings up Replicate and asks for a postmortem on why model API marketplaces struggle. Akshat points out that raw model APIs cater to transient hobbyist traffic, whereas Modal provides programmatic, customizable code-level infrastructure for enterprise products.50:40–53:00 · Guest disagreement 1/10 Differentiated AI Products and Video Orchestration Agents The host cites predictions from xAI team members about video generation moving toward multi-model orchestration agents. Akshat validates this with real customer behavior where agents execute complex FFmpeg workflows inside GPU sandboxes.53:01–55:40 · Guest disagreement 1/10 The Evolution of CI/CD for Coding Agents and Runtime Sandboxes The host discusses Gitpod/Ona joining OpenAI and queries the architecture of agent CI/CD. Akshat details how memory snapshot-and-restore primitives can eliminate dependency preparation bottlenecks and contrasts build-time with runtime sandboxes.55:41–57:02 · Guest disagreement 1/10 Multi-Language SDKs and Programming Paradigms in the AI Era The host asks if DX and AX are nearly identical paradigms. Akshat explains that while conceptually aligned, AX requires concrete product changes like porting UI observability metrics into machine-parseable CLI commands and benchmark-driven feature additions.1:18–4:28 · The hosts pushing back 2/10 Modal's Origin Story and Escaping Kubernetes Limitations The host recalls Modal's early serverless container stack presentations from Data Council and questions whether data pipelines genuinely needed that level of infrastructure. Akshat explains the evolution from serverless primitives to bursty ML inference and computer vision workloads.4:29–7:49 · The hosts pushing back 3/10 Self-Provisioning Infrastructure and Decorators Over YAML The host brings up his own past post on self-provisioning infrastructure and playfully notes the vendor lock-in pushback around custom DSLs. Akshat reframes the value around code ownership and how self-provisioning runtime benefits both human DX and agent AX.7:50–12:07 · The hosts pushing back 2/10 Defining Modal's Purpose: Cloud Primitives for AI Applications Hosts ask Akshat to define Modal's core identity beyond general compute, citing their early direct experience running Small Developer. Akshat outlines Modal as scratch-built cloud primitives specifically optimized for AI applications rather than static web servers.12:07–15:52 · The hosts pushing back 2/10 Handling Massive Scale and Elastic RL Training Runs Vibhu probes into rapid RL training run patterns like Cursor Composer scaling up and down hourly. Akshat clarifies the distinct workload profiles between diurnal custom inference, batch pre-training encoders, and massive 100k sandbox RL rollouts.15:52–21:13 · The hosts pushing back 2/10 Continual Learning and Upstreaming Open-Source LLM Advancements The host and guest discuss speculative decoding, with the host asking for technical clarification on draft models versus compute overhead. Akshat breaks down the mathematics of memory bandwidth bounds versus multi-token accept length multiplicative speedups in dFlash.21:13–23:15 · The hosts pushing back 2/10 The Modal Infrastructure Delta and Production-Grade Reliability Vibhu asks what technical delta Modal offers over running open-source engines like SGLang directly on raw compute. Akshat explains that Modal actively upstreams performance work to SGLang while offering zero-scaling elasticity and production-grade tail latency management.23:16–27:29 · The hosts pushing back 2/10 Supporting the Complete Model Life Cycle and Ramp Inspect The host notes the shifting CPU-to-GPU workload ratio from 8:1 to 1:1 in agentic workloads and asks about multi-cloud arbitrage. Akshat explains Modal's software-defined reliability layer aggregating over 17 distinct neocloud providers without owning physical data centers.27:30–31:02 · The hosts pushing back 1/10 Advanced Agent Networking, Sidecars, and Private IPv6 Overlays The host queries Modal's emerging networking capabilities, drawing comparisons to Docker Compose. Akshat reveals unpublished features like multi-container sidecars and their ISXPN private IPv6 overlay network originally built for distributed training.31:02–34:55 · The hosts pushing back 3/10 Inter-GPU Memory Management, RDMA, and eBPF Network Security The host mixes up WireGuard encryption with Modal's kernel-level TCP overlay and questions whether cross-node training networking is fast enough. Akshat politely disentangles eBPF access control and clarifies their internal 3 Tbps RDMA interconnect for medium post-training workloads.34:56–37:38 · The hosts pushing back 2/10 Autonomous Research, Auto Inference, and Parameter Sweeps The host characterizes auto-research as mostly science-fair hype and asks if it is merely hyperparameter sweeps. Akshat reveals Modal's internal 'auto inference' harness that autonomously runs NVIDIA Nsight profilers and tunes configurations to replace manual forward-deployed engineering.37:38–41:49 · The hosts pushing back 2/10 Evaluating LLM-Generated Modal Code and Building Modal Bench The host inquires about LLM code generation reliability for Modal and hardware supply shortages. Akshat discusses developing Modal Bench and explains compute strategy roles that model fungible multi-year reservations and capacity pricing tiers.41:50–45:13 · The hosts pushing back 3/10 Future Product Roadmap: Real-Time Media and Agent Ecosystems The host pushes a provocative hypothesis that an LLM should act directly as the OS kernel with adaptive permissions. Akshat firmly pushes back, maintaining that security at the sandbox level strictly requires deterministic hard boundaries rather than purely LLM-mediated guardrails.45:15–47:34 · The hosts pushing back 2/10 Managed Agent Ecosystem vs Dedicated Sandbox Infrastructure The host asks whether managed agent suites from frontier labs threaten Modal's position. Akshat explains that while labs provide great starting points, production enterprise agents like Ramp require granular control over persistent storage, GPU sandboxes, and snapshotting.47:34–50:40 · The hosts pushing back 2/10 Code-First Infrastructure vs Commodity Model APIs The host brings up Replicate and asks for a postmortem on why model API marketplaces struggle. Akshat points out that raw model APIs cater to transient hobbyist traffic, whereas Modal provides programmatic, customizable code-level infrastructure for enterprise products.50:40–53:00 · The hosts pushing back 1/10 Differentiated AI Products and Video Orchestration Agents The host cites predictions from xAI team members about video generation moving toward multi-model orchestration agents. Akshat validates this with real customer behavior where agents execute complex FFmpeg workflows inside GPU sandboxes.53:01–55:40 · The hosts pushing back 2/10 The Evolution of CI/CD for Coding Agents and Runtime Sandboxes The host discusses Gitpod/Ona joining OpenAI and queries the architecture of agent CI/CD. Akshat details how memory snapshot-and-restore primitives can eliminate dependency preparation bottlenecks and contrasts build-time with runtime sandboxes.55:41–57:02 · The hosts pushing back 2/10 Multi-Language SDKs and Programming Paradigms in the AI Era The host asks if DX and AX are nearly identical paradigms. Akshat explains that while conceptually aligned, AX requires concrete product changes like porting UI observability metrics into machine-parseable CLI commands and benchmark-driven feature additions.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 44:20 Rejecting LLM-mediated security kernel concept

Akshat directly dismisses the host's speculative hypothesis about an 'LLM OS' kernel, arguing firmly that sandbox security requires uncompromised deterministic boundaries.

Hardest push from the hosts ▶ 44:40 Host pushes non-consensus LLM-as-kernel thesis

The host challenges conventional security paradigms by suggesting standard sandbox approaches are dinosaur thinking compared to a hypothetical pure LLM kernel.

Biggest teaching moment ▶ 32:04 Clarifying eBPF packet filtering versus WireGuard VPNs

Akshat methodically clears up the host's confusion between user-space encrypted tunnels and in-kernel eBPF TCP authorization for inter-container communication.

The host holds their own ▶ 24:22 Host analyzes the 8:1 to 1:1 GPU-CPU inference inflection

The host synthesizes insights from Jensen Huang's GTC keynote to articulate how emerging agent workloads shift computational bottlenecks between CPU and GPU.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Modal's Origin Story and Escaping Kubernetes Limitations 6412 The host recalls Modal's early serverless container stack presentations from Data Council and questions whether data pipelines genuinely needed that level of infrastructure. Akshat explains the evolution from serverless primitives to bursty ML inference and computer vision workloads.
Self-Provisioning Infrastructure and Decorators Over YAML 6323 The host brings up his own past post on self-provisioning infrastructure and playfully notes the vendor lock-in pushback around custom DSLs. Akshat reframes the value around code ownership and how self-provisioning runtime benefits both human DX and agent AX.
Defining Modal's Purpose: Cloud Primitives for AI Applications 6312 Hosts ask Akshat to define Modal's core identity beyond general compute, citing their early direct experience running Small Developer. Akshat outlines Modal as scratch-built cloud primitives specifically optimized for AI applications rather than static web servers.
Handling Massive Scale and Elastic RL Training Runs 5522 Vibhu probes into rapid RL training run patterns like Cursor Composer scaling up and down hourly. Akshat clarifies the distinct workload profiles between diurnal custom inference, batch pre-training encoders, and massive 100k sandbox RL rollouts.
Continual Learning and Upstreaming Open-Source LLM Advancements 6512 The host and guest discuss speculative decoding, with the host asking for technical clarification on draft models versus compute overhead. Akshat breaks down the mathematics of memory bandwidth bounds versus multi-token accept length multiplicative speedups in dFlash.
The Modal Infrastructure Delta and Production-Grade Reliability 6412 Vibhu asks what technical delta Modal offers over running open-source engines like SGLang directly on raw compute. Akshat explains that Modal actively upstreams performance work to SGLang while offering zero-scaling elasticity and production-grade tail latency management.
Supporting the Complete Model Life Cycle and Ramp Inspect 6412 The host notes the shifting CPU-to-GPU workload ratio from 8:1 to 1:1 in agentic workloads and asks about multi-cloud arbitrage. Akshat explains Modal's software-defined reliability layer aggregating over 17 distinct neocloud providers without owning physical data centers.
Advanced Agent Networking, Sidecars, and Private IPv6 Overlays 5511 The host queries Modal's emerging networking capabilities, drawing comparisons to Docker Compose. Akshat reveals unpublished features like multi-container sidecars and their ISXPN private IPv6 overlay network originally built for distributed training.
Inter-GPU Memory Management, RDMA, and eBPF Network Security 5623 The host mixes up WireGuard encryption with Modal's kernel-level TCP overlay and questions whether cross-node training networking is fast enough. Akshat politely disentangles eBPF access control and clarifies their internal 3 Tbps RDMA interconnect for medium post-training workloads.
Autonomous Research, Auto Inference, and Parameter Sweeps 5412 The host characterizes auto-research as mostly science-fair hype and asks if it is merely hyperparameter sweeps. Akshat reveals Modal's internal 'auto inference' harness that autonomously runs NVIDIA Nsight profilers and tunes configurations to replace manual forward-deployed engineering.
Evaluating LLM-Generated Modal Code and Building Modal Bench 5412 The host inquires about LLM code generation reliability for Modal and hardware supply shortages. Akshat discusses developing Modal Bench and explains compute strategy roles that model fungible multi-year reservations and capacity pricing tiers.
Future Product Roadmap: Real-Time Media and Agent Ecosystems 6423 The host pushes a provocative hypothesis that an LLM should act directly as the OS kernel with adaptive permissions. Akshat firmly pushes back, maintaining that security at the sandbox level strictly requires deterministic hard boundaries rather than purely LLM-mediated guardrails.
Managed Agent Ecosystem vs Dedicated Sandbox Infrastructure 5412 The host asks whether managed agent suites from frontier labs threaten Modal's position. Akshat explains that while labs provide great starting points, production enterprise agents like Ramp require granular control over persistent storage, GPU sandboxes, and snapshotting.
Code-First Infrastructure vs Commodity Model APIs 6422 The host brings up Replicate and asks for a postmortem on why model API marketplaces struggle. Akshat points out that raw model APIs cater to transient hobbyist traffic, whereas Modal provides programmatic, customizable code-level infrastructure for enterprise products.
Differentiated AI Products and Video Orchestration Agents 6311 The host cites predictions from xAI team members about video generation moving toward multi-model orchestration agents. Akshat validates this with real customer behavior where agents execute complex FFmpeg workflows inside GPU sandboxes.
The Evolution of CI/CD for Coding Agents and Runtime Sandboxes 6412 The host discusses Gitpod/Ona joining OpenAI and queries the architecture of agent CI/CD. Akshat details how memory snapshot-and-restore primitives can eliminate dependency preparation bottlenecks and contrasts build-time with runtime sandboxes.
Multi-Language SDKs and Programming Paradigms in the AI Era 5412 The host asks if DX and AX are nearly identical paradigms. Akshat explains that while conceptually aligned, AX requires concrete product changes like porting UI observability metrics into machine-parseable CLI commands and benchmark-driven feature additions.

Statements from this episode (40)

Opinion
Bubna: Kubernetes lacks burstiness support and has terrible developer experience
“Kubernetes is hard to manage. It's not built for burstiness and custom images and has a terrible developer experience.”
Akshat Bubna Jul 8, 2026 ▶ 2:12
Disclosure
Bubna: Modal built decorators to eliminate YAML and make infra dynamic
“That, that, that was really important because we really didn't want people to spend so much time writing YAML, and it seemed like you could really condense the surface area of what you're doing put it in code so you can actually operate on it, just like you ca…”
Akshat Bubna Jul 8, 2026 ▶ 4:55
Disclosure
Bubna: Modal Refocused Its SDK Team on Agent Experience
“We've actually changed our SDK team to think about agent experience, sort of developer experience”
Akshat Bubna Jul 8, 2026 ▶ 6:09
Insight
Bubna: Agent Experience Benefits From Typed Decorators Over Kubernetes YAML
“We think that the same benefits that apply for DX also actually apply for AX, which is why would you have an agent read through hundreds of Kubernetes files and, like, write YAML that's not even typed when it can basically make a couple of changes in a decorat…”
Akshat Bubna Jul 8, 2026 ▶ 6:09
Opinion
Bubna: System Observability Is Becoming More Important Than Reading Code
“You still need humans to go interpret what's going on and you know, make judgment calls and whatnot. And that's, I feel like maybe more important now than looking at the code itself.”
Akshat Bubna Jul 8, 2026 ▶ 7:07
Disclosure
Bubna: Modal embedded an engineer at Cognition to reduce communication latency
“We sent him over because the latency of communication was too high otherwise.”
Akshat Bubna Jul 8, 2026 ▶ 9:20
Disclosure
Bubna: Modal built agent sandboxes and published smol-developer recursive loops early
“We built sandboxes in May of the year before anyone knew this was going to be a thing. And the first example we published was we took a small developer and put in a loop so the agent can iterate on itself.”
Akshat Bubna Jul 8, 2026 ▶ 10:34
Assertion Not checkable as stated
Modal CTO: Production scale requires elastically scaling 1,000 to 1,500 GPUs quickly
“There it's not about scaling from zero to one, but it's how do we scale really elastically from, like, thousand to 1500 GPUs very quickly in, in a given region.”
Akshat Bubna Jul 8, 2026 ▶ 12:44
Disclosure
Bubna: Modal's Biggest Use Case and Initial PMF Was Custom Non-LLM Inference
“Our biggest use case actually is elastic inference. And the thing we first found product market fit with was inference for custom models. So we kind of stayed away from the LM space and we were serving companies like Suno for audio, Runway for video, robotics …”
Akshat Bubna Jul 8, 2026 ▶ 13:33
Assertion Supported
Bubna: Modal Uses GPU Snapshotting to Speed Up Cold Starts
“We've incorporated GPU snapshotting to the product so we can actually take the GPU state, like your Torch compiler model, snapshot it, and the next call starts way faster.”
Akshat Bubna Jul 8, 2026 ▶ 14:59
Insight
Bubna: RL Rollouts Are Extremely Bursty and Can Require 100,000 Sandboxes
“RL is insanely bursty. Like when you're doing rollouts you sometimes need a 100,000 sandboxes.”
Akshat Bubna Jul 8, 2026 ▶ 15:42
Assertion Supported
Bubna: Open-source DFlash matches proprietary speculative decoding performance
“Recently we shared our work on dflash, which is a block-based speculator, and we've open sourced all of it, so you can get, by using open source dflash, you can get the same performance as you would with one of the proprietary providers.”
Akshat Bubna Jul 8, 2026 ▶ 17:13
Insight
Bubna: Speculative decoding accept length delivers multiplicative speedups over kernel tuning
“People talk a lot about, we made these kernels faster and whatnot, but improving kernel only give you like a few percentage points of improvement and increasing except length literally is a multiplicative decrease.”
Akshat Bubna Jul 8, 2026 ▶ 18:48
Assertion Supported
Bubna: Speculative decoding has zero impact on model output quality
“So there's no drop in quality performance, because you're always, you're never accepting a token that's a big model.”
Akshat Bubna Jul 8, 2026 ▶ 19:14
Insight
Modal CTO: Production AI inference is difficult due to tail latency
“Running production grade inference is a hard and fair problem. Even if you subtract out the auto scaling, it's controlling things like tail latency and making sure every request is delivered at least once and whatnot.”
Akshat Bubna Jul 8, 2026 ▶ 22:58
Assertion Not checkable as stated
Bubna: Ramp Inspect Succeeded Using Modal Snapshotting and Fast Scaling
“Ramp Inspect was a great example of a background agent that was really successful because they were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.”
Akshat Bubna Jul 8, 2026 ▶ 23:59
Assertion Not checkable as stated
Bubna: Modal's capacity pool spans 17 cloud providers
“We've built this capacity pool that spans 17 cloud providers. So we're very good at running on various kinds of cloud capacity across the world.”
Akshat Bubna Jul 8, 2026 ▶ 25:23
Disclosure
Bubna: Modal operates zero data centers, running entirely on NeoClouds
“We don't have our own data centers. We just run across a lot of Neo clouds.”
Akshat Bubna Jul 8, 2026 ▶ 25:40
Assertion Not checkable as stated
Bubna: Modal's custom reliability layer insulates users from GPU hardware drops
“That's why it's something we've invested a lot of time in is actually building our own reliability layer on top. So if the GPU falls off the bus or something happens, we user workloads are not affected.”
Akshat Bubna Jul 8, 2026 ▶ 26:27
Disclosure
Modal sandboxes support multi-container sidecar pods
“So actually, if you want Docker Compose our sandboxes now support this thing called Sidecarves. So you can, a sandbox is actually a pod of containers, and you can run multiple containers in a sandbox.”
Akshat Bubna Jul 8, 2026 ▶ 28:19
Disclosure
Modal built an IPv6 overlay network for private container addressing
“We have this thing called I-SXPN, which we haven't talked about which is this, like, overlay network using IPv six addresses so if modal containers within the same workspace when this is enabled, can actually address each other using this private IPv six addre…”
Akshat Bubna Jul 8, 2026 ▶ 29:18
Insight
Bubna: Transferring RL weights is fundamentally an OS memory problem
“Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is, there's a lot of degrees of freedom, and it is basically a systems problem of Moving me…”
Akshat Bubna Jul 8, 2026 ▶ 31:40
Disclosure
Bubna: Modal uses eBPF-filtered TCP rather than VPN encryption for internal networking
“This is TCP, and we have eBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.”
Akshat Bubna Jul 8, 2026 ▶ 32:41
Assertion Supported
Bubna: Modal provides approximately 3 Tbps internal networking bandwidth
“And we have I think like three terabit per second internal networking, which is the standard that's needed.”
Akshat Bubna Jul 8, 2026 ▶ 33:49
Disclosure
Bubna: Modal multi-node training targets post-training, not large-scale pre-training
“And we're not going for obviously like large scale pre-training runs. The thing that we've built multi-handle training for is we see a lot of smaller scale post-training like people are post-training like medium-sized fun models so they can get higher quality …”
Akshat Bubna Jul 8, 2026 ▶ 34:14
Disclosure
Bubna: Modal uses autonomous agents to run internal inference optimization sweeps
“Internal both training and inference teams actually use this sort of the general shape of this quite a bit. Like we have this one internal repo called auto inference, which essentially we've automated our own FDE efforts using this harness, which is the agent …”
Akshat Bubna Jul 8, 2026 ▶ 35:28
Insight
Bubna: Auto-research is hyperparameter tuning guided by model intuition, not architecture changes
“I, so the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's basically a high proprietor sweep that's guided by some sort of model intuition. So it's like much more efficient than whatever oth…”
Akshat Bubna Jul 8, 2026 ▶ 37:10
Insight
Bubna: AI agents struggle to reason through logs and observability
“I think the things that sometimes agents struggle with without right guidance and a skill is how to use the rest of our observability. Like, how to, something is failing, like, How do you look at the logs and then update the right thing? It's sort of reasoning…”
Akshat Bubna Jul 8, 2026 ▶ 38:15
Disclosure
Bubna: Modal is building a cheaper 24-hour batch compute tier
“One of the things we're building now is, like, a way for customers to get if they don't care about latency like, get much cheaper pricing and they'll get results back in, like, next 24 hours or something. Like a batch tier, essentially.”
Akshat Bubna Jul 8, 2026 ▶ 40:48
Assertion Not checkable as stated
Bubna: Batch compute demand comes mainly from non-LLM workloads like computational biology
“The demand that we see for something like that is actually not for LLMs. Although sometimes people want to run evals and do synthetic data prep and there it makes sense. But it's from a lot of non LLM companies like people who are doing computational bio, like…”
Akshat Bubna Jul 8, 2026 ▶ 41:15
Prediction Not checkable as stated
Bubna: Thousands more companies will post-train and deploy open-source models
“I think for example, LM inference, thousands more companies are going to post train their own models and deploy open source models for inference.”
Akshat Bubna Jul 8, 2026 ▶ 42:17
Opinion
Bubna: Sandboxes Require Hard Boundaries, Not LLM-Mediated Permissions
“I'm skeptical of LLM-mediated permissions for stuff that is At the sandbox level, because you do want hard boundaries. Otherwise, obviously someone can exfiltrate stuff.”
Akshat Bubna Jul 8, 2026 ▶ 44:21
Disclosure
Bubna: Ramp runs its external-facing accounting agent on Modal
“Ramp also runs their accounting agent on us, so their external facing agent.”
Akshat Bubna Jul 8, 2026 ▶ 45:57
Insight
Bubna: Production AI agents require specialized sandboxes for compute and networking control
“You need a lot more control over your compute primitive on things like what sort of, how do you persist different files that the agent has access to, and how do you snapshot and restore? How do you control the networking? Maybe you want GPUs. When you get to t…”
Akshat Bubna Jul 8, 2026 ▶ 46:05
Insight
Bubna: Model APIs primarily serve a less sticky hobbyist market
“This is one thing we've kind of stayed away from is providing an API for models, because I think providing Model APIs is, some of it ends up serving like a really hobbyist market, which is much less sticky.”
Akshat Bubna Jul 8, 2026 ▶ 49:49
Assertion Open · timeframe Jul 2026
Bubna: Ramp trained custom tokenizers to swap into LLaMA
“Ramp actually early in the day was training their own tokenizer and, like, Swapping out the tokenizer in Lama and whatnot.”
Akshat Bubna Jul 8, 2026 ▶ 51:11
Disclosure
Modal CTO: Suno runs 100% of inference on Modal
“They use modal for all their inference, and that's because they have like a custom, they have completely custom model architecture, and that means that they have to be at the code level and tweak things that are not Yeah, it's an API.”
Akshat Bubna Jul 8, 2026 ▶ 51:42
Prediction Not checkable as stated
Bubna: Modal is bullish on CI as coding agents run more tests
“We're very bullish and modal on the CI market as well because it has, There's more agents coding agents they're gonna run a lot more CI, and the preventives there can be much better.”
Akshat Bubna Jul 8, 2026 ▶ 53:25
Insight
Bubna: AI Agent Developers Prefer TypeScript Over Python
“And the interesting thing with, like, the agent stuff is people use their TypeScript SDK a lot more because they're not actually doing anything that needs ML.”
Akshat Bubna Jul 8, 2026 ▶ 56:19
Insight
Bubna: AI agent hallucinations are actionable product feedback
“Sometimes it makes sense, like, if they're reaching for this thing, it's product feedback, like, give it to them.”
Akshat Bubna Jul 8, 2026 ▶ 58:26
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.