Jun 24, 2026 · 1h 10m · latent-space

The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin

Matei Zaharia · 29m spoken Reynold Xin · 21m spoken Shawn Wang · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space interview, Databricks co-founders Matei Zaharia and Reynold Xin explore how the convergence of open-source agent orchestration (Omnigent), unified lakehouse storage (LTAP), and ML-driven database engines (Raiden) is shifting enterprise software from traditional codebases to data-centric AI agent platforms.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.1 Guest teaching 3.3 Guest disagreement 1.1 The hosts pushing back 1.2
05100:0015:0030:0045:001:00:000:28–2:48 · The hosts as informed peer 1/10 Host Welcome and Channel Subscription Announcement The segment begins with a solo host channel announcement followed by warm introductory banter regarding the growth of the Databricks Summit from a small 50-person Berkeley meetup.2:48–5:20 · The hosts as informed peer 4/10 Introducing Omnigent: The Meta-Harness for Agent Orchestration Swyx introduces Omnigent and queries why Databricks pursued a meta-harness. Matei explains how internal coding workflows and research agent developments converged on the need for shared sessions and model portability.5:20–9:52 · The hosts as informed peer 5/10 Architectural Parallels: Protocols, Interoperability, and Open Sharing When Swyx suggests an operating system framing, Matei reframes the architecture to network protocols and open data sharing like real-time supplier tables. Reynold shares an anecdote about driving while coding on a laptop, prompting the cloud sandbox concept.9:52–13:50 · The hosts as informed peer 5/10 The Open Source Architecture of the Agent Cloud Swyx asks about open-sourcing philosophy, drawing parallels to Spark. Matei and Reynold elaborate on network effects, community pull requests, and the boundary between open protocols and proprietary managed infrastructure.13:50–16:35 · The hosts as informed peer 5/10 Deconstructing the Modern Data Stack vs Modern AI Stack Swyx pronounces the Modern Data Stack dead and asks if a Modern AI Stack replaces it. Reynold defends the core abstractions while explaining customer pressure toward unification, and Matei highlights the common harness API.16:36–19:13 · The hosts as informed peer 6/10 Compute Sandboxing, Operational Scale, and Database Virtualization Swyx brings up Neon and observes that database companies are fundamentally compute providers. Reynold shares massive scale metrics, detailing 15-16 million VMs and exabytes of daily processing.19:14–25:24 · The hosts as informed peer 5/10 Agent Security: Contextual Policies and Financial Guardrails Matei details Omnigent's contextual and stateful security policies, balancing developer convenience against supply-chain injection attacks and token spend caps.25:25–28:48 · The hosts as informed peer 6/10 Extensibility, Developer Tools, and Emerging Startup Opportunities Swyx inquires about startup opportunities around coding agents, citing examples like Git AI and Artificial Analysis. Matei emphasizes management planes and quality attribution tools.28:48–35:18 · The hosts as informed peer 6/10 The LTAP Architecture: Unifying OLTP and OLAP for Agents Reynold unpacks the LTAP architecture, comparing it to historical HTAP and lampooning change data capture (CDC) as brittle 'continuous data corruption'.35:18–38:00 · The hosts as informed peer 5/10 The Technical Breakthrough: Storage-Layer Parquet Transcoding Reynold describes the technical breakthrough of LTAP: utilizing idle CPUs in the storage layer to transcode Postgres row pages into column-oriented Parquet files without formal bureaucracy.38:01–42:43 · The hosts as informed peer 4/10 Innovation Culture: Rapid Prototyping Over Bureaucratic Specs The conversation shifts to execution culture. Matei and Reynold advocate rapid incremental prototyping anchored to specific customer design partners rather than boiling the ocean.42:43–45:02 · The hosts as informed peer 5/10 Navigating Enterprise Needs vs Tech Startup Mentalities The founders delineate the stark operational contrast between agile tech companies prone to DIY and regulated enterprise customers prioritizing security, compliance, and reliability.45:03–51:00 · The hosts as informed peer 6/10 The Dream Engine: ML-Driven Database Architecture (Raiden) Reynold explains 'Project Raiden', a clean-sheet engine rewrite avoiding second-system syndrome by using ML models trained on quadrillions of historical execution traces to optimize algorithms dynamically.51:00–53:45 · The hosts as informed peer 5/10 Incremental Rollouts and the Future of Specialized Databases Reynold dismisses vector databases as an unnecessary separate product category, arguing LTAP unifies storage while query engines remain specialized and LLM agents seamlessly generate whatever dialect is required.53:45–58:57 · The hosts as informed peer 7/10 Strategic Differentiation: Databricks vs Snowflake Swyx asks a direct question on why Databricks outpaced Snowflake. Reynold and Matei point to their early bet on open table formats, native ML workloads, and starting upstream in large-scale data ingestion.58:58–1:04:25 · The hosts as informed peer 7/10 Model Strategy: Domain-Specific AI and the Mosaic Vision Swyx presses on Databricks' post-Mosaic model strategy. Matei explains pivoting away from general frontier model pretraining toward task-specialized sub-agents, vision document parsers, and automated RL loops.1:04:25–1:08:23 · The hosts as informed peer 5/10 Context as IP: Enterprise Reinforcement Learning and AI Runtimes Swyx references Satya Nadella's essay on frontier context as IP. Matei and Reynold agree that enterprise proprietary data powering agentic reasoning will reinvent vertical software stacks.0:28–2:48 · Guest teaching 0/10 Host Welcome and Channel Subscription Announcement The segment begins with a solo host channel announcement followed by warm introductory banter regarding the growth of the Databricks Summit from a small 50-person Berkeley meetup.2:48–5:20 · Guest teaching 3/10 Introducing Omnigent: The Meta-Harness for Agent Orchestration Swyx introduces Omnigent and queries why Databricks pursued a meta-harness. Matei explains how internal coding workflows and research agent developments converged on the need for shared sessions and model portability.5:20–9:52 · Guest teaching 4/10 Architectural Parallels: Protocols, Interoperability, and Open Sharing When Swyx suggests an operating system framing, Matei reframes the architecture to network protocols and open data sharing like real-time supplier tables. Reynold shares an anecdote about driving while coding on a laptop, prompting the cloud sandbox concept.9:52–13:50 · Guest teaching 3/10 The Open Source Architecture of the Agent Cloud Swyx asks about open-sourcing philosophy, drawing parallels to Spark. Matei and Reynold elaborate on network effects, community pull requests, and the boundary between open protocols and proprietary managed infrastructure.13:50–16:35 · Guest teaching 4/10 Deconstructing the Modern Data Stack vs Modern AI Stack Swyx pronounces the Modern Data Stack dead and asks if a Modern AI Stack replaces it. Reynold defends the core abstractions while explaining customer pressure toward unification, and Matei highlights the common harness API.16:36–19:13 · Guest teaching 3/10 Compute Sandboxing, Operational Scale, and Database Virtualization Swyx brings up Neon and observes that database companies are fundamentally compute providers. Reynold shares massive scale metrics, detailing 15-16 million VMs and exabytes of daily processing.19:14–25:24 · Guest teaching 4/10 Agent Security: Contextual Policies and Financial Guardrails Matei details Omnigent's contextual and stateful security policies, balancing developer convenience against supply-chain injection attacks and token spend caps.25:25–28:48 · Guest teaching 2/10 Extensibility, Developer Tools, and Emerging Startup Opportunities Swyx inquires about startup opportunities around coding agents, citing examples like Git AI and Artificial Analysis. Matei emphasizes management planes and quality attribution tools.28:48–35:18 · Guest teaching 4/10 The LTAP Architecture: Unifying OLTP and OLAP for Agents Reynold unpacks the LTAP architecture, comparing it to historical HTAP and lampooning change data capture (CDC) as brittle 'continuous data corruption'.35:18–38:00 · Guest teaching 3/10 The Technical Breakthrough: Storage-Layer Parquet Transcoding Reynold describes the technical breakthrough of LTAP: utilizing idle CPUs in the storage layer to transcode Postgres row pages into column-oriented Parquet files without formal bureaucracy.38:01–42:43 · Guest teaching 3/10 Innovation Culture: Rapid Prototyping Over Bureaucratic Specs The conversation shifts to execution culture. Matei and Reynold advocate rapid incremental prototyping anchored to specific customer design partners rather than boiling the ocean.42:43–45:02 · Guest teaching 4/10 Navigating Enterprise Needs vs Tech Startup Mentalities The founders delineate the stark operational contrast between agile tech companies prone to DIY and regulated enterprise customers prioritizing security, compliance, and reliability.45:03–51:00 · Guest teaching 4/10 The Dream Engine: ML-Driven Database Architecture (Raiden) Reynold explains 'Project Raiden', a clean-sheet engine rewrite avoiding second-system syndrome by using ML models trained on quadrillions of historical execution traces to optimize algorithms dynamically.51:00–53:45 · Guest teaching 4/10 Incremental Rollouts and the Future of Specialized Databases Reynold dismisses vector databases as an unnecessary separate product category, arguing LTAP unifies storage while query engines remain specialized and LLM agents seamlessly generate whatever dialect is required.53:45–58:57 · Guest teaching 4/10 Strategic Differentiation: Databricks vs Snowflake Swyx asks a direct question on why Databricks outpaced Snowflake. Reynold and Matei point to their early bet on open table formats, native ML workloads, and starting upstream in large-scale data ingestion.58:58–1:04:25 · Guest teaching 4/10 Model Strategy: Domain-Specific AI and the Mosaic Vision Swyx presses on Databricks' post-Mosaic model strategy. Matei explains pivoting away from general frontier model pretraining toward task-specialized sub-agents, vision document parsers, and automated RL loops.1:04:25–1:08:23 · Guest teaching 3/10 Context as IP: Enterprise Reinforcement Learning and AI Runtimes Swyx references Satya Nadella's essay on frontier context as IP. Matei and Reynold agree that enterprise proprietary data powering agentic reasoning will reinvent vertical software stacks.0:28–2:48 · Guest disagreement 0/10 Host Welcome and Channel Subscription Announcement The segment begins with a solo host channel announcement followed by warm introductory banter regarding the growth of the Databricks Summit from a small 50-person Berkeley meetup.2:48–5:20 · Guest disagreement 1/10 Introducing Omnigent: The Meta-Harness for Agent Orchestration Swyx introduces Omnigent and queries why Databricks pursued a meta-harness. Matei explains how internal coding workflows and research agent developments converged on the need for shared sessions and model portability.5:20–9:52 · Guest disagreement 1/10 Architectural Parallels: Protocols, Interoperability, and Open Sharing When Swyx suggests an operating system framing, Matei reframes the architecture to network protocols and open data sharing like real-time supplier tables. Reynold shares an anecdote about driving while coding on a laptop, prompting the cloud sandbox concept.9:52–13:50 · Guest disagreement 0/10 The Open Source Architecture of the Agent Cloud Swyx asks about open-sourcing philosophy, drawing parallels to Spark. Matei and Reynold elaborate on network effects, community pull requests, and the boundary between open protocols and proprietary managed infrastructure.13:50–16:35 · Guest disagreement 2/10 Deconstructing the Modern Data Stack vs Modern AI Stack Swyx pronounces the Modern Data Stack dead and asks if a Modern AI Stack replaces it. Reynold defends the core abstractions while explaining customer pressure toward unification, and Matei highlights the common harness API.16:36–19:13 · Guest disagreement 1/10 Compute Sandboxing, Operational Scale, and Database Virtualization Swyx brings up Neon and observes that database companies are fundamentally compute providers. Reynold shares massive scale metrics, detailing 15-16 million VMs and exabytes of daily processing.19:14–25:24 · Guest disagreement 0/10 Agent Security: Contextual Policies and Financial Guardrails Matei details Omnigent's contextual and stateful security policies, balancing developer convenience against supply-chain injection attacks and token spend caps.25:25–28:48 · Guest disagreement 0/10 Extensibility, Developer Tools, and Emerging Startup Opportunities Swyx inquires about startup opportunities around coding agents, citing examples like Git AI and Artificial Analysis. Matei emphasizes management planes and quality attribution tools.28:48–35:18 · Guest disagreement 2/10 The LTAP Architecture: Unifying OLTP and OLAP for Agents Reynold unpacks the LTAP architecture, comparing it to historical HTAP and lampooning change data capture (CDC) as brittle 'continuous data corruption'.35:18–38:00 · Guest disagreement 1/10 The Technical Breakthrough: Storage-Layer Parquet Transcoding Reynold describes the technical breakthrough of LTAP: utilizing idle CPUs in the storage layer to transcode Postgres row pages into column-oriented Parquet files without formal bureaucracy.38:01–42:43 · Guest disagreement 1/10 Innovation Culture: Rapid Prototyping Over Bureaucratic Specs The conversation shifts to execution culture. Matei and Reynold advocate rapid incremental prototyping anchored to specific customer design partners rather than boiling the ocean.42:43–45:02 · Guest disagreement 1/10 Navigating Enterprise Needs vs Tech Startup Mentalities The founders delineate the stark operational contrast between agile tech companies prone to DIY and regulated enterprise customers prioritizing security, compliance, and reliability.45:03–51:00 · Guest disagreement 1/10 The Dream Engine: ML-Driven Database Architecture (Raiden) Reynold explains 'Project Raiden', a clean-sheet engine rewrite avoiding second-system syndrome by using ML models trained on quadrillions of historical execution traces to optimize algorithms dynamically.51:00–53:45 · Guest disagreement 4/10 Incremental Rollouts and the Future of Specialized Databases Reynold dismisses vector databases as an unnecessary separate product category, arguing LTAP unifies storage while query engines remain specialized and LLM agents seamlessly generate whatever dialect is required.53:45–58:57 · Guest disagreement 2/10 Strategic Differentiation: Databricks vs Snowflake Swyx asks a direct question on why Databricks outpaced Snowflake. Reynold and Matei point to their early bet on open table formats, native ML workloads, and starting upstream in large-scale data ingestion.58:58–1:04:25 · Guest disagreement 1/10 Model Strategy: Domain-Specific AI and the Mosaic Vision Swyx presses on Databricks' post-Mosaic model strategy. Matei explains pivoting away from general frontier model pretraining toward task-specialized sub-agents, vision document parsers, and automated RL loops.1:04:25–1:08:23 · Guest disagreement 0/10 Context as IP: Enterprise Reinforcement Learning and AI Runtimes Swyx references Satya Nadella's essay on frontier context as IP. Matei and Reynold agree that enterprise proprietary data powering agentic reasoning will reinvent vertical software stacks.0:28–2:48 · The hosts pushing back 0/10 Host Welcome and Channel Subscription Announcement The segment begins with a solo host channel announcement followed by warm introductory banter regarding the growth of the Databricks Summit from a small 50-person Berkeley meetup.2:48–5:20 · The hosts pushing back 1/10 Introducing Omnigent: The Meta-Harness for Agent Orchestration Swyx introduces Omnigent and queries why Databricks pursued a meta-harness. Matei explains how internal coding workflows and research agent developments converged on the need for shared sessions and model portability.5:20–9:52 · The hosts pushing back 1/10 Architectural Parallels: Protocols, Interoperability, and Open Sharing When Swyx suggests an operating system framing, Matei reframes the architecture to network protocols and open data sharing like real-time supplier tables. Reynold shares an anecdote about driving while coding on a laptop, prompting the cloud sandbox concept.9:52–13:50 · The hosts pushing back 1/10 The Open Source Architecture of the Agent Cloud Swyx asks about open-sourcing philosophy, drawing parallels to Spark. Matei and Reynold elaborate on network effects, community pull requests, and the boundary between open protocols and proprietary managed infrastructure.13:50–16:35 · The hosts pushing back 2/10 Deconstructing the Modern Data Stack vs Modern AI Stack Swyx pronounces the Modern Data Stack dead and asks if a Modern AI Stack replaces it. Reynold defends the core abstractions while explaining customer pressure toward unification, and Matei highlights the common harness API.16:36–19:13 · The hosts pushing back 1/10 Compute Sandboxing, Operational Scale, and Database Virtualization Swyx brings up Neon and observes that database companies are fundamentally compute providers. Reynold shares massive scale metrics, detailing 15-16 million VMs and exabytes of daily processing.19:14–25:24 · The hosts pushing back 1/10 Agent Security: Contextual Policies and Financial Guardrails Matei details Omnigent's contextual and stateful security policies, balancing developer convenience against supply-chain injection attacks and token spend caps.25:25–28:48 · The hosts pushing back 1/10 Extensibility, Developer Tools, and Emerging Startup Opportunities Swyx inquires about startup opportunities around coding agents, citing examples like Git AI and Artificial Analysis. Matei emphasizes management planes and quality attribution tools.28:48–35:18 · The hosts pushing back 1/10 The LTAP Architecture: Unifying OLTP and OLAP for Agents Reynold unpacks the LTAP architecture, comparing it to historical HTAP and lampooning change data capture (CDC) as brittle 'continuous data corruption'.35:18–38:00 · The hosts pushing back 1/10 The Technical Breakthrough: Storage-Layer Parquet Transcoding Reynold describes the technical breakthrough of LTAP: utilizing idle CPUs in the storage layer to transcode Postgres row pages into column-oriented Parquet files without formal bureaucracy.38:01–42:43 · The hosts pushing back 0/10 Innovation Culture: Rapid Prototyping Over Bureaucratic Specs The conversation shifts to execution culture. Matei and Reynold advocate rapid incremental prototyping anchored to specific customer design partners rather than boiling the ocean.42:43–45:02 · The hosts pushing back 1/10 Navigating Enterprise Needs vs Tech Startup Mentalities The founders delineate the stark operational contrast between agile tech companies prone to DIY and regulated enterprise customers prioritizing security, compliance, and reliability.45:03–51:00 · The hosts pushing back 2/10 The Dream Engine: ML-Driven Database Architecture (Raiden) Reynold explains 'Project Raiden', a clean-sheet engine rewrite avoiding second-system syndrome by using ML models trained on quadrillions of historical execution traces to optimize algorithms dynamically.51:00–53:45 · The hosts pushing back 1/10 Incremental Rollouts and the Future of Specialized Databases Reynold dismisses vector databases as an unnecessary separate product category, arguing LTAP unifies storage while query engines remain specialized and LLM agents seamlessly generate whatever dialect is required.53:45–58:57 · The hosts pushing back 3/10 Strategic Differentiation: Databricks vs Snowflake Swyx asks a direct question on why Databricks outpaced Snowflake. Reynold and Matei point to their early bet on open table formats, native ML workloads, and starting upstream in large-scale data ingestion.58:58–1:04:25 · The hosts pushing back 2/10 Model Strategy: Domain-Specific AI and the Mosaic Vision Swyx presses on Databricks' post-Mosaic model strategy. Matei explains pivoting away from general frontier model pretraining toward task-specialized sub-agents, vision document parsers, and automated RL loops.1:04:25–1:08:23 · The hosts pushing back 1/10 Context as IP: Enterprise Reinforcement Learning and AI Runtimes Swyx references Satya Nadella's essay on frontier context as IP. Matei and Reynold agree that enterprise proprietary data powering agentic reasoning will reinvent vertical software stacks.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 52:32 Vector databases dismissal

Reynold bluntly dismisses an entire VC category by asserting that vector databases should never have existed as an independent architectural segment.

Hardest push from the hosts ▶ 53:45 Pushing the Snowflake comparison

Swyx directly challenges the guests to explain objectively why they succeeded and outpaced Snowflake where Snowflake failed.

Biggest teaching moment ▶ 5:35 Reframing agent architecture to network protocols

Matei politely corrects the host's operating system metaphor by demonstrating why agent coordination mirrors open data sharing and network protocols.

The host holds their own ▶ 16:36 Host connects database architectures to compute sandboxes

Swyx demonstrates deep domain knowledge by citing Neon's compute-storage separation to explain why every modern database is secretly a compute orchestrator.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Host Welcome and Channel Subscription Announcement 1000 The segment begins with a solo host channel announcement followed by warm introductory banter regarding the growth of the Databricks Summit from a small 50-person Berkeley meetup.
Introducing Omnigent: The Meta-Harness for Agent Orchestration 4311 Swyx introduces Omnigent and queries why Databricks pursued a meta-harness. Matei explains how internal coding workflows and research agent developments converged on the need for shared sessions and model portability.
Architectural Parallels: Protocols, Interoperability, and Open Sharing 5411 When Swyx suggests an operating system framing, Matei reframes the architecture to network protocols and open data sharing like real-time supplier tables. Reynold shares an anecdote about driving while coding on a laptop, prompting the cloud sandbox concept.
The Open Source Architecture of the Agent Cloud 5301 Swyx asks about open-sourcing philosophy, drawing parallels to Spark. Matei and Reynold elaborate on network effects, community pull requests, and the boundary between open protocols and proprietary managed infrastructure.
Deconstructing the Modern Data Stack vs Modern AI Stack 5422 Swyx pronounces the Modern Data Stack dead and asks if a Modern AI Stack replaces it. Reynold defends the core abstractions while explaining customer pressure toward unification, and Matei highlights the common harness API.
Compute Sandboxing, Operational Scale, and Database Virtualization 6311 Swyx brings up Neon and observes that database companies are fundamentally compute providers. Reynold shares massive scale metrics, detailing 15-16 million VMs and exabytes of daily processing.
Agent Security: Contextual Policies and Financial Guardrails 5401 Matei details Omnigent's contextual and stateful security policies, balancing developer convenience against supply-chain injection attacks and token spend caps.
Extensibility, Developer Tools, and Emerging Startup Opportunities 6201 Swyx inquires about startup opportunities around coding agents, citing examples like Git AI and Artificial Analysis. Matei emphasizes management planes and quality attribution tools.
The LTAP Architecture: Unifying OLTP and OLAP for Agents 6421 Reynold unpacks the LTAP architecture, comparing it to historical HTAP and lampooning change data capture (CDC) as brittle 'continuous data corruption'.
The Technical Breakthrough: Storage-Layer Parquet Transcoding 5311 Reynold describes the technical breakthrough of LTAP: utilizing idle CPUs in the storage layer to transcode Postgres row pages into column-oriented Parquet files without formal bureaucracy.
Innovation Culture: Rapid Prototyping Over Bureaucratic Specs 4310 The conversation shifts to execution culture. Matei and Reynold advocate rapid incremental prototyping anchored to specific customer design partners rather than boiling the ocean.
Navigating Enterprise Needs vs Tech Startup Mentalities 5411 The founders delineate the stark operational contrast between agile tech companies prone to DIY and regulated enterprise customers prioritizing security, compliance, and reliability.
The Dream Engine: ML-Driven Database Architecture (Raiden) 6412 Reynold explains 'Project Raiden', a clean-sheet engine rewrite avoiding second-system syndrome by using ML models trained on quadrillions of historical execution traces to optimize algorithms dynamically.
Incremental Rollouts and the Future of Specialized Databases 5441 Reynold dismisses vector databases as an unnecessary separate product category, arguing LTAP unifies storage while query engines remain specialized and LLM agents seamlessly generate whatever dialect is required.
Strategic Differentiation: Databricks vs Snowflake 7423 Swyx asks a direct question on why Databricks outpaced Snowflake. Reynold and Matei point to their early bet on open table formats, native ML workloads, and starting upstream in large-scale data ingestion.
Model Strategy: Domain-Specific AI and the Mosaic Vision 7412 Swyx presses on Databricks' post-Mosaic model strategy. Matei explains pivoting away from general frontier model pretraining toward task-specialized sub-agents, vision document parsers, and automated RL loops.
Context as IP: Enterprise Reinforcement Learning and AI Runtimes 5301 Swyx references Satya Nadella's essay on frontier context as IP. Matei and Reynold agree that enterprise proprietary data powering agentic reasoning will reinvent vertical software stacks.

Statements from this episode (31)

Disclosure
Zaharia: Databricks built Isaac as an internal wrapper for Claude Code and Codex
“We have a really great dev info team. They built something called Isaac that's basically like a wrapper on cloud code and codex and let's you use them either on the web and like, Sandboxes or just on your dev machine or on your laptop or whatever.”
Matei Zaharia Jun 24, 2026 ▶ 3:47
Insight
Zaharia: AI agents are useless without a collaboration and history layer
“Plus the agent is like completely useless if you can't share sessions with someone and have history and have search and all this like layer on top of it for collaboration.”
Matei Zaharia Jun 24, 2026 ▶ 4:41
Insight
Zaharia: Coding agents and custom agents share identical underlying technical problems
“They're like, why are you doing coding agents and custom agents in the same thing, but I said it's basically the same problems.”
Matei Zaharia Jun 24, 2026 ▶ 4:56
Insight
Zaharia: Protocol design remains essential for multi-party interoperability despite fast coding
“For this type of interoperability where multiple parties that are moving at different speeds are building stuff and you still want some layer on top to coordinate you do want to design it and build it. So it reminds me of that, like agents talking to each othe…”
Matei Zaharia Jun 24, 2026 ▶ 6:37
Disclosure
Zaharia: Omnigent was built to allow teams to securely share custom agent setups
“We had a lot of engineers building, you know, their own Vibe coding setup. But then the other thing they all said is like, Hey, I built something that's amazing for me. But like no one else on the team can use it because I don't have a server to collaborate. A…”
Matei Zaharia Jun 24, 2026 ▶ 9:12
Insight
Zaharia: Prototyped AI agents stall when security teams block internal data access
“That's where we've seen a lot of other agents like hit things like people think they prototyped an awesome agent, but you know, it's not allowed to connect to like some really important data or whatever because of the, Security team.”
Matei Zaharia Jun 24, 2026 ▶ 9:39
Assertion Not checkable as stated
Xin: Databricks teams internally built five or six redundant agent frameworks
“I think we had like five or six different agentic frameworks built by every different team. They do all do more or less the same thing.”
Reynold Xin Jun 24, 2026 ▶ 10:59
Insight
Zaharia: Software with integration network effects should be open source
“One, so, I mean, one of the reasons to open source something is if you think it's a layer that will actually, there'll be some network effect. It'll benefit from many people collaborating on it.”
Matei Zaharia Jun 24, 2026 ▶ 11:20
Prediction Not checkable as stated
Zaharia: Open agent hosting layers will win over proprietary alternatives
“Another way to think about it is like, imagine, you know we, our thing wasn't open. We had some kind of agent hosting thing, but it's not open. And then there is an open one. If you're, which one's gonna win in the long run? So like here, because there is this…”
Matei Zaharia Jun 24, 2026 ▶ 12:09
Insight
Xin: Customer friction with vendor sprawl forced Modern Data Stack consolidation
“What people eventually run into, It's kind of a question of a unification and consolidation is, hey, do you really need to chop all of this into different pieces and work with so many different vendors and platforms in order to get like a very simple visualiza…”
Reynold Xin Jun 24, 2026 ▶ 14:40
Prediction Not checkable as stated
Xin: AI agent tooling is consolidating like Modern Data Stack did
“I think honestly something like this is probably happening in how many different frameworks do you want to hook up together in order to produce, like do a very simple agent?”
Reynold Xin Jun 24, 2026 ▶ 15:10
Insight
Swyx: Every database company is also fundamentally a compute company
“Every database company is also a compute company.”
Shawn Wang Jun 24, 2026 ▶ 16:59
Assertion Not checkable as stated
Xin: Databricks launches 15 to 16 million VMs per day across clouds
“So we, on the analytics side, I think we launched Maybe 50 or sixteen million virtual machines a day across all three clouds, so which are one of the biggest compute orchestrators out there. Stuff for sure for CPU compute.”
Reynold Xin Jun 24, 2026 ▶ 18:16
Assertion Not checkable as stated
Xin: Neon is launching 13 million databases per day
“And on Neon, it's actually pretty interesting too. It's launching, I think, thirteen million databases a day now.”
Reynold Xin Jun 24, 2026 ▶ 18:42
Insight
Matei Zaharia: AI agents need stateful contextual security policies, not binary permissions
“A lot of coding agents today have very basic things like you can tell me which tool patterns I'll allow or disallow or whatever. It's like yes or no, but that puts you in a very tough spot. So just as an example, like, should my agent be able to read, you know…”
Matei Zaharia Jun 24, 2026 ▶ 19:46
Disclosure
Zaharia: Databricks gives internal developers unlimited AI token spend
“It's unlimited, but we do you know, we use our own product to like analyze the traces and stuff, and we have a team that's, you know, looking to optimize and to see if anyone's doing something weird.”
Matei Zaharia Jun 24, 2026 ▶ 24:22
Prediction Not checkable as stated
Swyx: Coding agent management will move from consultants to software
“I think this is like the domain of consultants first, but then people will actually build software that has like the management plane for coding agents.”
Shawn Wang Jun 24, 2026 ▶ 28:35
Insight
Xin: Single HTAP database engines compromise ecosystem compatibility and performance
“This is sort of the holy grail of database engineering is, why not build a single system that can do both of this? But it ends up just being a lot of compromises. And one, I think one of the first issue is that, hey, each, they say Postgres has a massive ecosy…”
Reynold Xin Jun 24, 2026 ▶ 32:30
Insight
Xin: Unifying storage delivers 99% of HTAP database benefits
“HTAP wants to build a single engine for both. We think you can get 99% of what you need by unifying the storage and just have a single storage layer.”
Reynold Xin Jun 24, 2026 ▶ 33:16
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Reynold Xin Jun 24, 2026 ▶ 36:50
Insight
Xin: Overfitting to Initial Customers Has Far Lower Downside Than Boiling the Ocean
“I think the industry has a sense of, hey, maybe if you overfit to like one or two customers, it's going to be really bad for you. But I think the downside overfitting is much smaller than the upside itself. And if you sort of try to be too ambitious and boil t…”
Reynold Xin Jun 24, 2026 ▶ 41:55
Insight
Xin: Optimizing Solely for Tech Companies Impedes Traditional Enterprise Scaling
“One of the challenges I think we probably see, and maybe many newer generation companies are seeing is, so tech companies are very, very different from non-tech companies or traditional enterprises. And if you optimize everything just for tech companies, you m…”
Reynold Xin Jun 24, 2026 ▶ 42:25
Assertion Supported
Xin: Every major analytics database engine in traction is a decade old
“Actually, every single database engine out there, especially on the analytics side, are kind of a decade old. Pretty much everything that had reasonable traction are about a decade old.”
Reynold Xin Jun 24, 2026 ▶ 45:08
Disclosure
Xin: Databricks uses ML models to optimize database algorithms at runtime
“They use that to build a model. Like a machine learning model, not an L, a machine learning model. Machine learning model basically can very, very quickly tell us how any algorithm and how any implementation will perform for any specific type of queries with v…”
Reynold Xin Jun 24, 2026 ▶ 48:03
Opinion
Xin: Vector databases should never have been a separate category
“Vector database should have never been a separate category.”
Reynold Xin Jun 24, 2026 ▶ 52:32
Opinion
Xin: Multiple database query languages are not an issue for AI agents
“Instead of worrying about PostgreSQL and maybe Spark SQL, why not just one? But I don't think that's an issue for agents. Agents are very eloquent in PostgreSQL or Spark SQL. It's never going to get confused. As long as the data is there and it's accessible ag…”
Reynold Xin Jun 24, 2026 ▶ 53:10
Insight
Zaharia: Scaling down from bulk ingestion to serving is easier than scaling up
“It turned out that, you know, it's easier to go from that bod thing that's really good at the scale and ingesting and super low cost and create versions in it that have the speed and features of the, you know, super easy to use, like smaller data for business …”
Matei Zaharia Jun 24, 2026 ▶ 55:55
Disclosure
Zaharia: Databricks abandons frontier AI models to focus on agent systems
“Even though we did launch open source model DBRX, and, you know, we went up to, like, sort of above the LAMA-R III scale, we decided that we really want to focus on, there'll be so many people releasing models, and instead of doing the general model where, lik…”
Matei Zaharia Jun 24, 2026 ▶ 1:00:03
Assertion Not checkable as stated
Zaharia: Databricks' document vision model is ~100x cheaper than frontier models
“Our team built this document sort of vision model that takes a page and gives you back a nice JSON with all the components. And it's very competitive. It's like probably like a hundred X cheaper than those frontier models and still better.”
Matei Zaharia Jun 24, 2026 ▶ 1:01:56
Prediction Not checkable as stated
Zaharia: Customizing AI models will get significantly easier over time
“My feeling is, like customizing models is actually going to get way easier over time. That's what we're finding, because The base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces, and the…”
Matei Zaharia Jun 24, 2026 ▶ 1:03:14
Prediction Not checkable as stated
Xin: Much of traditional software will be rewritten with data and agents
“Actually, I think many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there. And then they slap some agent on top.”
Reynold Xin Jun 24, 2026 ▶ 1:07:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.