May 24, 2021 · 46m · mad

Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)

Ali Ghodsi · 33m spoken Matt Turck · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat hosted by Matt Turck, Databricks Founder and CEO Ali Ghodsi discusses the journey of founding Databricks from UC Berkeley's AMP Lab to scaling it into a global enterprise data and AI leader. He shares insights into the technical evolution of the Lakehouse architecture, open-source commercialization strategies, organizational scaling, and the future role of AI in enterprise software.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 19% of the talking time here. How this is scored →

Matt as informed peer 4.1 Guest teaching 5.0 Guest disagreement 1.1 Matt pushing back 1.3
05100:0015:0030:0045:000:09–5:01 · Matt as informed peer 4/10 The Origin of Databricks and UC Berkeley's AMP Lab Matt opens with well-informed context about UC Berkeley's AMP Lab, Spark, and academic culture. Ali educates the host on how 1970s ML algorithms achieved modern breakthroughs simply by scaling data volumes on distributed systems.5:01–10:51 · Matt as informed peer 3/10 Managing a Large Team of Seven Co-Founders Matt probes into the rare dynamic of having seven co-founders and asks about scaling from 0 to 10M ARR. Ali candidly explains their early GTM misstep of relying on product-led credit-card growth rather than executive enterprise sales.10:51–15:42 · Matt as informed peer 6/10 Evolution of Databricks into a Multi-Product Platform Matt demonstrates high expertise by referencing his 2015 interview with Databricks co-founder Ion Stoica and detailing the product evolution from Spark to MLflow and Lakehouse. Ali explains their AWS-inspired open-source monetization playbook.15:42–20:36 · Matt as informed peer 5/10 The Lakehouse Paradigm vs. Traditional Warehouses Matt asks technical questions contrasting data lakes with data warehouses and specifically inquires about Apache Iceberg and trade-offs. Ali details how Delta Lake provided transactional structure over data lakes.20:36–25:19 · Matt as informed peer 5/10 Lakehouse Dominance vs. Coexistence with Snowflake Matt raises industry chatter regarding the competition between Snowflake and Databricks. Ali reframes the zero-sum narrative by drawing parallels to major cloud provider coexistence while holding that Lakehouse architecture will win long-term.25:19–28:33 · Matt as informed peer 4/10 Organizational Structure for Multi-Product R&D Matt questions how Databricks organizes R&D to avoid being a one-product company. Ali outlines their organizational separation of disruptive zero-to-one teams from enterprise maintenance teams based on 'Zone to Win'.28:33–31:44 · Matt as informed peer 4/10 Open Source Business Model and the Future of AI Matt asks if Ali would start another company as open-source first. Ali argues open source hosted in the cloud is evolutionarily superior to proprietary software models and predicts AI will consume traditional software.31:44–34:42 · Matt as informed peer 4/10 Hybrid Go-To-Market Strategy and Developer Advocacy Matt asks how open source fits into sales motions without internal channel conflict. Ali explains how community edition usage generates half their enterprise sales pipeline and bypasses long proof-of-concept cycles.34:42–39:33 · Matt as informed peer 4/10 Scaling as CEO and Executive Leadership Recruitment Matt asks how Ali scaled personally as CEO and handles internal promotion versus external executive recruiting. Ali describes evaluating leaders on whether they can build a car from scratch rather than just drive it.39:33–42:27 · Matt as informed peer 3/10 Q&A: Data Mesh Architecture Matt asks an audience question about Data Mesh architecture. Ali explains Data Mesh as an organizational decentralization response to centralized bottlenecked data teams, supported by Lakehouse governance.42:27–44:20 · Matt as informed peer 3/10 Q&A: Managed Cloud Containers and Kubernetes Matt selects a technical question on cloud containers and Kubernetes. Ali compares Kubernetes to a universal hardware standard like USB, while stressing that data platforms require higher-level governance layers.0:09–5:01 · Guest teaching 5/10 The Origin of Databricks and UC Berkeley's AMP Lab Matt opens with well-informed context about UC Berkeley's AMP Lab, Spark, and academic culture. Ali educates the host on how 1970s ML algorithms achieved modern breakthroughs simply by scaling data volumes on distributed systems.5:01–10:51 · Guest teaching 5/10 Managing a Large Team of Seven Co-Founders Matt probes into the rare dynamic of having seven co-founders and asks about scaling from 0 to 10M ARR. Ali candidly explains their early GTM misstep of relying on product-led credit-card growth rather than executive enterprise sales.10:51–15:42 · Guest teaching 5/10 Evolution of Databricks into a Multi-Product Platform Matt demonstrates high expertise by referencing his 2015 interview with Databricks co-founder Ion Stoica and detailing the product evolution from Spark to MLflow and Lakehouse. Ali explains their AWS-inspired open-source monetization playbook.15:42–20:36 · Guest teaching 5/10 The Lakehouse Paradigm vs. Traditional Warehouses Matt asks technical questions contrasting data lakes with data warehouses and specifically inquires about Apache Iceberg and trade-offs. Ali details how Delta Lake provided transactional structure over data lakes.20:36–25:19 · Guest teaching 5/10 Lakehouse Dominance vs. Coexistence with Snowflake Matt raises industry chatter regarding the competition between Snowflake and Databricks. Ali reframes the zero-sum narrative by drawing parallels to major cloud provider coexistence while holding that Lakehouse architecture will win long-term.25:19–28:33 · Guest teaching 5/10 Organizational Structure for Multi-Product R&D Matt questions how Databricks organizes R&D to avoid being a one-product company. Ali outlines their organizational separation of disruptive zero-to-one teams from enterprise maintenance teams based on 'Zone to Win'.28:33–31:44 · Guest teaching 5/10 Open Source Business Model and the Future of AI Matt asks if Ali would start another company as open-source first. Ali argues open source hosted in the cloud is evolutionarily superior to proprietary software models and predicts AI will consume traditional software.31:44–34:42 · Guest teaching 5/10 Hybrid Go-To-Market Strategy and Developer Advocacy Matt asks how open source fits into sales motions without internal channel conflict. Ali explains how community edition usage generates half their enterprise sales pipeline and bypasses long proof-of-concept cycles.34:42–39:33 · Guest teaching 5/10 Scaling as CEO and Executive Leadership Recruitment Matt asks how Ali scaled personally as CEO and handles internal promotion versus external executive recruiting. Ali describes evaluating leaders on whether they can build a car from scratch rather than just drive it.39:33–42:27 · Guest teaching 5/10 Q&A: Data Mesh Architecture Matt asks an audience question about Data Mesh architecture. Ali explains Data Mesh as an organizational decentralization response to centralized bottlenecked data teams, supported by Lakehouse governance.42:27–44:20 · Guest teaching 5/10 Q&A: Managed Cloud Containers and Kubernetes Matt selects a technical question on cloud containers and Kubernetes. Ali compares Kubernetes to a universal hardware standard like USB, while stressing that data platforms require higher-level governance layers.0:09–5:01 · Guest disagreement 1/10 The Origin of Databricks and UC Berkeley's AMP Lab Matt opens with well-informed context about UC Berkeley's AMP Lab, Spark, and academic culture. Ali educates the host on how 1970s ML algorithms achieved modern breakthroughs simply by scaling data volumes on distributed systems.5:01–10:51 · Guest disagreement 1/10 Managing a Large Team of Seven Co-Founders Matt probes into the rare dynamic of having seven co-founders and asks about scaling from 0 to 10M ARR. Ali candidly explains their early GTM misstep of relying on product-led credit-card growth rather than executive enterprise sales.10:51–15:42 · Guest disagreement 1/10 Evolution of Databricks into a Multi-Product Platform Matt demonstrates high expertise by referencing his 2015 interview with Databricks co-founder Ion Stoica and detailing the product evolution from Spark to MLflow and Lakehouse. Ali explains their AWS-inspired open-source monetization playbook.15:42–20:36 · Guest disagreement 1/10 The Lakehouse Paradigm vs. Traditional Warehouses Matt asks technical questions contrasting data lakes with data warehouses and specifically inquires about Apache Iceberg and trade-offs. Ali details how Delta Lake provided transactional structure over data lakes.20:36–25:19 · Guest disagreement 2/10 Lakehouse Dominance vs. Coexistence with Snowflake Matt raises industry chatter regarding the competition between Snowflake and Databricks. Ali reframes the zero-sum narrative by drawing parallels to major cloud provider coexistence while holding that Lakehouse architecture will win long-term.25:19–28:33 · Guest disagreement 1/10 Organizational Structure for Multi-Product R&D Matt questions how Databricks organizes R&D to avoid being a one-product company. Ali outlines their organizational separation of disruptive zero-to-one teams from enterprise maintenance teams based on 'Zone to Win'.28:33–31:44 · Guest disagreement 1/10 Open Source Business Model and the Future of AI Matt asks if Ali would start another company as open-source first. Ali argues open source hosted in the cloud is evolutionarily superior to proprietary software models and predicts AI will consume traditional software.31:44–34:42 · Guest disagreement 1/10 Hybrid Go-To-Market Strategy and Developer Advocacy Matt asks how open source fits into sales motions without internal channel conflict. Ali explains how community edition usage generates half their enterprise sales pipeline and bypasses long proof-of-concept cycles.34:42–39:33 · Guest disagreement 1/10 Scaling as CEO and Executive Leadership Recruitment Matt asks how Ali scaled personally as CEO and handles internal promotion versus external executive recruiting. Ali describes evaluating leaders on whether they can build a car from scratch rather than just drive it.39:33–42:27 · Guest disagreement 1/10 Q&A: Data Mesh Architecture Matt asks an audience question about Data Mesh architecture. Ali explains Data Mesh as an organizational decentralization response to centralized bottlenecked data teams, supported by Lakehouse governance.42:27–44:20 · Guest disagreement 1/10 Q&A: Managed Cloud Containers and Kubernetes Matt selects a technical question on cloud containers and Kubernetes. Ali compares Kubernetes to a universal hardware standard like USB, while stressing that data platforms require higher-level governance layers.0:09–5:01 · Matt pushing back 1/10 The Origin of Databricks and UC Berkeley's AMP Lab Matt opens with well-informed context about UC Berkeley's AMP Lab, Spark, and academic culture. Ali educates the host on how 1970s ML algorithms achieved modern breakthroughs simply by scaling data volumes on distributed systems.5:01–10:51 · Matt pushing back 1/10 Managing a Large Team of Seven Co-Founders Matt probes into the rare dynamic of having seven co-founders and asks about scaling from 0 to 10M ARR. Ali candidly explains their early GTM misstep of relying on product-led credit-card growth rather than executive enterprise sales.10:51–15:42 · Matt pushing back 1/10 Evolution of Databricks into a Multi-Product Platform Matt demonstrates high expertise by referencing his 2015 interview with Databricks co-founder Ion Stoica and detailing the product evolution from Spark to MLflow and Lakehouse. Ali explains their AWS-inspired open-source monetization playbook.15:42–20:36 · Matt pushing back 2/10 The Lakehouse Paradigm vs. Traditional Warehouses Matt asks technical questions contrasting data lakes with data warehouses and specifically inquires about Apache Iceberg and trade-offs. Ali details how Delta Lake provided transactional structure over data lakes.20:36–25:19 · Matt pushing back 3/10 Lakehouse Dominance vs. Coexistence with Snowflake Matt raises industry chatter regarding the competition between Snowflake and Databricks. Ali reframes the zero-sum narrative by drawing parallels to major cloud provider coexistence while holding that Lakehouse architecture will win long-term.25:19–28:33 · Matt pushing back 1/10 Organizational Structure for Multi-Product R&D Matt questions how Databricks organizes R&D to avoid being a one-product company. Ali outlines their organizational separation of disruptive zero-to-one teams from enterprise maintenance teams based on 'Zone to Win'.28:33–31:44 · Matt pushing back 1/10 Open Source Business Model and the Future of AI Matt asks if Ali would start another company as open-source first. Ali argues open source hosted in the cloud is evolutionarily superior to proprietary software models and predicts AI will consume traditional software.31:44–34:42 · Matt pushing back 1/10 Hybrid Go-To-Market Strategy and Developer Advocacy Matt asks how open source fits into sales motions without internal channel conflict. Ali explains how community edition usage generates half their enterprise sales pipeline and bypasses long proof-of-concept cycles.34:42–39:33 · Matt pushing back 1/10 Scaling as CEO and Executive Leadership Recruitment Matt asks how Ali scaled personally as CEO and handles internal promotion versus external executive recruiting. Ali describes evaluating leaders on whether they can build a car from scratch rather than just drive it.39:33–42:27 · Matt pushing back 1/10 Q&A: Data Mesh Architecture Matt asks an audience question about Data Mesh architecture. Ali explains Data Mesh as an organizational decentralization response to centralized bottlenecked data teams, supported by Lakehouse governance.42:27–44:20 · Matt pushing back 1/10 Q&A: Managed Cloud Containers and Kubernetes Matt selects a technical question on cloud containers and Kubernetes. Ali compares Kubernetes to a universal hardware standard like USB, while stressing that data platforms require higher-level governance layers.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 15.3% · guest 84.7%0:00 · Matt 15.3% · guest 84.7%3:00 · Matt 22.7% · guest 77.3%3:00 · Matt 22.7% · guest 77.3%6:00 · Matt 24% · guest 76%6:00 · Matt 24% · guest 76%9:00 · Matt 39.2% · guest 60.8%9:00 · Matt 39.2% · guest 60.8%12:00 · Matt 0.6% · guest 99.4%12:00 · Matt 0.6% · guest 99.4%15:00 · Matt 15.9% · guest 84.1%15:00 · Matt 15.9% · guest 84.1%18:00 · Matt 16.1% · guest 83.9%18:00 · Matt 16.1% · guest 83.9%21:00 · Matt 12.5% · guest 87.5%21:00 · Matt 12.5% · guest 87.5%24:00 · Matt 22.5% · guest 77.5%24:00 · Matt 22.5% · guest 77.5%27:00 · Matt 13.1% · guest 86.9%27:00 · Matt 13.1% · guest 86.9%30:00 · Matt 18.9% · guest 81.1%30:00 · Matt 18.9% · guest 81.1%33:00 · Matt 26.7% · guest 73.3%33:00 · Matt 26.7% · guest 73.3%36:00 · Matt 8.7% · guest 91.3%36:00 · Matt 8.7% · guest 91.3%39:00 · Matt 9.3% · guest 90.7%39:00 · Matt 9.3% · guest 90.7%42:00 · Matt 28% · guest 72%42:00 · Matt 28% · guest 72%45:00 · Matt 41.4% · guest 58.6%45:00 · Matt 41.4% · guest 58.6%
Sharpest disagreement ▶ 21:05 Rejection of zero-sum Snowflake rivalry

Ali directly reframes the host's framing of a winner-take-all clash with Snowflake, comparing the dynamic to coexisting cloud providers like AWS, Azure, and Google Cloud.

Hardest push from Matt ▶ 19:18 Host probing on Lakehouse trade-offs

Matt refuses to accept the Lakehouse model as purely superior without scrutiny, directly asking whether there are technical trade-offs to combining data lakes and warehouses.

Biggest teaching moment ▶ 8:50 Reframing PLG vs Enterprise Go-To-Market

Ali educates the host on product-market-channel fit, explaining why relying purely on product-led credit-card growth was a strategic mistake for strategic enterprise software.

Matt holds his own ▶ 11:16 Recalling 2015 interview to trace product expansion

Matt demonstrates deep background knowledge by recalling his 2015 interview with co-founder Ion Stoica and accurately mapping Databricks' transition from MapReduce replacement to full ML/Lakehouse platform.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Origin of Databricks and UC Berkeley's AMP Lab 4511 Matt opens with well-informed context about UC Berkeley's AMP Lab, Spark, and academic culture. Ali educates the host on how 1970s ML algorithms achieved modern breakthroughs simply by scaling data volumes on distributed systems.
Managing a Large Team of Seven Co-Founders 3511 Matt probes into the rare dynamic of having seven co-founders and asks about scaling from 0 to 10M ARR. Ali candidly explains their early GTM misstep of relying on product-led credit-card growth rather than executive enterprise sales.
Evolution of Databricks into a Multi-Product Platform 6511 Matt demonstrates high expertise by referencing his 2015 interview with Databricks co-founder Ion Stoica and detailing the product evolution from Spark to MLflow and Lakehouse. Ali explains their AWS-inspired open-source monetization playbook.
The Lakehouse Paradigm vs. Traditional Warehouses 5512 Matt asks technical questions contrasting data lakes with data warehouses and specifically inquires about Apache Iceberg and trade-offs. Ali details how Delta Lake provided transactional structure over data lakes.
Lakehouse Dominance vs. Coexistence with Snowflake 5523 Matt raises industry chatter regarding the competition between Snowflake and Databricks. Ali reframes the zero-sum narrative by drawing parallels to major cloud provider coexistence while holding that Lakehouse architecture will win long-term.
Organizational Structure for Multi-Product R&D 4511 Matt questions how Databricks organizes R&D to avoid being a one-product company. Ali outlines their organizational separation of disruptive zero-to-one teams from enterprise maintenance teams based on 'Zone to Win'.
Open Source Business Model and the Future of AI 4511 Matt asks if Ali would start another company as open-source first. Ali argues open source hosted in the cloud is evolutionarily superior to proprietary software models and predicts AI will consume traditional software.
Hybrid Go-To-Market Strategy and Developer Advocacy 4511 Matt asks how open source fits into sales motions without internal channel conflict. Ali explains how community edition usage generates half their enterprise sales pipeline and bypasses long proof-of-concept cycles.
Scaling as CEO and Executive Leadership Recruitment 4511 Matt asks how Ali scaled personally as CEO and handles internal promotion versus external executive recruiting. Ali describes evaluating leaders on whether they can build a car from scratch rather than just drive it.
Q&A: Data Mesh Architecture 3511 Matt asks an audience question about Data Mesh architecture. Ali explains Data Mesh as an organizational decentralization response to centralized bottlenecked data teams, supported by Lakehouse governance.
Q&A: Managed Cloud Containers and Kubernetes 3511 Matt selects a technical question on cloud containers and Kubernetes. Ali compares Kubernetes to a universal hardware standard like USB, while stressing that data platforms require higher-level governance layers.

Statements from this episode (18)

Assertion Not checkable as stated
Ghodsi: Early Big Tech achieved AI breakthroughs using 1970s algorithms with massive data
“What they were doing is they were taking those algorithms from the seventies that do not work, but they were applying orders of magnitude, more data to it. So a lot of data on modern hardware, and they were getting superhuman results.”
Ali Ghodsi May 24, 2021 ▶ 1:06
Opinion
Ghodsi: Hadoop was terrible for machine learning tasks
“The people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.”
Ali Ghodsi May 24, 2021 ▶ 2:40
Disclosure
Ghodsi: Databricks' early reliance on product-led growth was a mistake
“On the channel side, the mistake with it is we really early on, we're really big believers in this product led growth. We said, you know, we're going to build this beautiful simplified product that we now have. We put it online and it's going to be cloud based…”
Ali Ghodsi May 24, 2021 ▶ 9:20
Insight
Ghodsi: Founders cannot pick their sales channel; product-market fit dictates it
“You don't get to pick your channel. You can, you don't get to say, oh, I want my ASP to be 50 K or 60. That's not your choice. You have a product. You have a market. If it has fit, you have to find the right channel to connect those two.”
Ali Ghodsi May 24, 2021 ▶ 9:56
Disclosure
Ghodsi: Databricks open-sources products only after achieving product-market fit
“So our secret sauce is look at an enterprise problem, figure out what that is, understand it deeply by being really customer obsessed, bring the problem back, have the innovators, the PhDs that know how to solve these problems. Solve the problem. Iterate quick…”
Ali Ghodsi May 24, 2021 ▶ 14:35
Insight
Ghodsi: Creating open-source software gives cloud hosts a competitive advantage
“We're going to host open source software in the cloud, but the difference is we'll create open source software. That way we get competitive advantage with respect to anyone else who would want to do the same thing, right? Otherwise anyone can pick up any open …”
Ali Ghodsi May 24, 2021 ▶ 15:23
Assertion Partly supported
Ghodsi: Four major technical breakthroughs enabled the Lakehouse around 2016-2017
“Yeah, actually, the four technological breakthroughs that kind of happened at the same time, 2016, 17, at the same time, the one we contributed was Delta Lake, there was Hootie, there was Hive acid, and there was icebergs.”
Ali Ghodsi May 24, 2021 ▶ 17:39
Prediction Not checkable as stated
Ghodsi: Open-source Lakehouses will eventually render proprietary data warehouses obsolete
“So slowly what's happening is that the open source realm An ecosystem is emerging where you can do all of your analytics in this lake house paradigm, and you don't, eventually it will be the case that you will not need all these other, you know, proprietary ol…”
Ali Ghodsi May 24, 2021 ▶ 20:16
Assertion Not checkable as stated
Ghodsi: Snowflake coexists with Databricks in about 70% of accounts
“It's certainly going to coexist, and it already coexists with Databricks in probably 70% of the accounts we're in.”
Ali Ghodsi May 24, 2021 ▶ 21:41
Disclosure
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Ali Ghodsi May 24, 2021 ▶ 23:12
Disclosure
Ghodsi: Databricks will open-source lower-stack layers and build proprietary software above
“We're definitely going to continue to move up the stack and then commoditize the stuff that's below by open sourcing it and just releasing it to the market and making it the standard and then moving up the stack with innovations.”
Ali Ghodsi May 24, 2021 ▶ 25:02
Disclosure
Databricks splits R&D into enterprise stability and new innovation teams
“All of engineering and product is separated into two different pieces. One that focuses on the things that enterprises need, large enterprises, encryption, security, authentication, stability, and so on. And another piece that focuses on these innovations.”
Ali Ghodsi May 24, 2021 ▶ 27:21
Insight
Ghodsi: Enterprise demand starves new innovation unless R&D teams are separated
“So you actually, or chart wise should separate those out because otherwise what happens when you're successful is that the former gets all of the resources because the big enterprises have infinite demand for your for the things that you're doing.”
Ali Ghodsi May 24, 2021 ▶ 27:34
Insight
Ghodsi: Any proprietary software company is ripe for open-source disruption
“Any proprietary software company out there is ripe for disruption by an open source competitor.”
Ali Ghodsi May 24, 2021 ▶ 29:59
Disclosure
Ghodsi: Half of Databricks' leads come from free Community Edition
“Half of our leads comes from that.”
Ali Ghodsi May 24, 2021 ▶ 33:14
Disclosure
Ghodsi: Databricks hires executive leaders who build, not just maintain
“What are the things we look for? We look for people who have seen build. So the joke I say is, do you have a driver's license? And people will say, I have it. Are you good at driving your car? Yes, I'm very good at it. Why are you asking? Can you build a car w…”
Ali Ghodsi May 24, 2021 ▶ 37:35
Insight
Ghodsi: Centralized data teams inevitably become operational bottlenecks
“As you can imagine, it would never scale in a large organization. That team would become bottlenecked. They would not understand how to prioritize the different projects because they don't understand the asks of marketing, sales, customer success and everythin…”
Ali Ghodsi May 24, 2021 ▶ 40:21
Disclosure
Ali Ghodsi: Databricks will be IPO ready in 2021
“We're going to be IPO ready this year and have marched towards that and are pretty far along in terms of sort of the readiness of the business everywhere.”
Ali Ghodsi May 24, 2021 ▶ 44:38
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.