Nov 16, 2021 · 29m · mad

Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)

Mark Grover · 22m spoken Matt Turck · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat hosted by FirstMark's Matt Turck, Stemma Co-Founder and CEO Mark Grover discusses the evolution of automated data catalogs, the operational transition from open-source Amundsen to Stemma, and strategic approaches to modern enterprise data governance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.1% of the talking time here. How this is scored →

Matt as informed peer 3.8 Guest teaching 4.9 Guest disagreement 0.8 Matt pushing back 1.4
05100:0010:0020:000:12–3:12 · Matt as informed peer 1/10 Defining Data Catalogs and the Data Explosion Problem Matt asks a broad introductory question defining data catalogs. Mark takes a step back to deliver an extensive monologue mapping out the decade-long evolution of data lakes, warehouses, ETL tools, and BI platforms.3:12–7:53 · Matt as informed peer 3/10 Data Governance, Productivity, and Compliance Matt briefly prompts on governance and interjects to clarify manual wiki cataloging. Mark educates him on the core failure modes of curated catalogs versus automated catalogs.7:53–10:48 · Matt as informed peer 2/10 How Automated Metadata Integration Works Matt asks how metadata automation works technically. Mark provides a comprehensive breakdown of integrations across warehouses, BI tools, HR hierarchies, and Slack.10:48–13:11 · Matt as informed peer 4/10 Access Control Strategy and Data Permissions Matt poses a concrete scenario regarding GDPR access control. Mark gently reframes the premise, distinguishing between dataset discovery versus underlying permission enforcement.13:11–15:39 · Matt as informed peer 6/10 Navigating the Unbundled Modern Data Stack Matt cites Zhamak Dehghani and data mesh principles, challenging Mark on how Stemma avoids getting overwhelmed by limitless integrations in an unbundled ecosystem.15:39–18:48 · Matt as informed peer 2/10 The Evolution from Lyft's Amundsen to Founding Stemma Matt asks about the transition from Lyft's internal project Amundsen to commercial startup Stemma. Mark describes internal CSAT scores and open-source growth.18:48–23:10 · Matt as informed peer 7/10 Market Landscape and Competitive Positioning Matt demonstrates deep ecosystem expertise by naming five specific market competitors across different startup generations. Mark responds with a two-axis positioning framework.23:10–25:54 · Matt as informed peer 5/10 Open Source Commercial Strategy and Business Model Matt prompts a strategic discussion on open source monetization. Mark leverages his Cloudera background to explicitly reject support-and-services models as unviable.0:12–3:12 · Guest teaching 6/10 Defining Data Catalogs and the Data Explosion Problem Matt asks a broad introductory question defining data catalogs. Mark takes a step back to deliver an extensive monologue mapping out the decade-long evolution of data lakes, warehouses, ETL tools, and BI platforms.3:12–7:53 · Guest teaching 6/10 Data Governance, Productivity, and Compliance Matt briefly prompts on governance and interjects to clarify manual wiki cataloging. Mark educates him on the core failure modes of curated catalogs versus automated catalogs.7:53–10:48 · Guest teaching 6/10 How Automated Metadata Integration Works Matt asks how metadata automation works technically. Mark provides a comprehensive breakdown of integrations across warehouses, BI tools, HR hierarchies, and Slack.10:48–13:11 · Guest teaching 5/10 Access Control Strategy and Data Permissions Matt poses a concrete scenario regarding GDPR access control. Mark gently reframes the premise, distinguishing between dataset discovery versus underlying permission enforcement.13:11–15:39 · Guest teaching 4/10 Navigating the Unbundled Modern Data Stack Matt cites Zhamak Dehghani and data mesh principles, challenging Mark on how Stemma avoids getting overwhelmed by limitless integrations in an unbundled ecosystem.15:39–18:48 · Guest teaching 3/10 The Evolution from Lyft's Amundsen to Founding Stemma Matt asks about the transition from Lyft's internal project Amundsen to commercial startup Stemma. Mark describes internal CSAT scores and open-source growth.18:48–23:10 · Guest teaching 4/10 Market Landscape and Competitive Positioning Matt demonstrates deep ecosystem expertise by naming five specific market competitors across different startup generations. Mark responds with a two-axis positioning framework.23:10–25:54 · Guest teaching 5/10 Open Source Commercial Strategy and Business Model Matt prompts a strategic discussion on open source monetization. Mark leverages his Cloudera background to explicitly reject support-and-services models as unviable.0:12–3:12 · Guest disagreement 0/10 Defining Data Catalogs and the Data Explosion Problem Matt asks a broad introductory question defining data catalogs. Mark takes a step back to deliver an extensive monologue mapping out the decade-long evolution of data lakes, warehouses, ETL tools, and BI platforms.3:12–7:53 · Guest disagreement 1/10 Data Governance, Productivity, and Compliance Matt briefly prompts on governance and interjects to clarify manual wiki cataloging. Mark educates him on the core failure modes of curated catalogs versus automated catalogs.7:53–10:48 · Guest disagreement 0/10 How Automated Metadata Integration Works Matt asks how metadata automation works technically. Mark provides a comprehensive breakdown of integrations across warehouses, BI tools, HR hierarchies, and Slack.10:48–13:11 · Guest disagreement 1/10 Access Control Strategy and Data Permissions Matt poses a concrete scenario regarding GDPR access control. Mark gently reframes the premise, distinguishing between dataset discovery versus underlying permission enforcement.13:11–15:39 · Guest disagreement 1/10 Navigating the Unbundled Modern Data Stack Matt cites Zhamak Dehghani and data mesh principles, challenging Mark on how Stemma avoids getting overwhelmed by limitless integrations in an unbundled ecosystem.15:39–18:48 · Guest disagreement 0/10 The Evolution from Lyft's Amundsen to Founding Stemma Matt asks about the transition from Lyft's internal project Amundsen to commercial startup Stemma. Mark describes internal CSAT scores and open-source growth.18:48–23:10 · Guest disagreement 1/10 Market Landscape and Competitive Positioning Matt demonstrates deep ecosystem expertise by naming five specific market competitors across different startup generations. Mark responds with a two-axis positioning framework.23:10–25:54 · Guest disagreement 2/10 Open Source Commercial Strategy and Business Model Matt prompts a strategic discussion on open source monetization. Mark leverages his Cloudera background to explicitly reject support-and-services models as unviable.0:12–3:12 · Matt pushing back 0/10 Defining Data Catalogs and the Data Explosion Problem Matt asks a broad introductory question defining data catalogs. Mark takes a step back to deliver an extensive monologue mapping out the decade-long evolution of data lakes, warehouses, ETL tools, and BI platforms.3:12–7:53 · Matt pushing back 1/10 Data Governance, Productivity, and Compliance Matt briefly prompts on governance and interjects to clarify manual wiki cataloging. Mark educates him on the core failure modes of curated catalogs versus automated catalogs.7:53–10:48 · Matt pushing back 0/10 How Automated Metadata Integration Works Matt asks how metadata automation works technically. Mark provides a comprehensive breakdown of integrations across warehouses, BI tools, HR hierarchies, and Slack.10:48–13:11 · Matt pushing back 2/10 Access Control Strategy and Data Permissions Matt poses a concrete scenario regarding GDPR access control. Mark gently reframes the premise, distinguishing between dataset discovery versus underlying permission enforcement.13:11–15:39 · Matt pushing back 3/10 Navigating the Unbundled Modern Data Stack Matt cites Zhamak Dehghani and data mesh principles, challenging Mark on how Stemma avoids getting overwhelmed by limitless integrations in an unbundled ecosystem.15:39–18:48 · Matt pushing back 0/10 The Evolution from Lyft's Amundsen to Founding Stemma Matt asks about the transition from Lyft's internal project Amundsen to commercial startup Stemma. Mark describes internal CSAT scores and open-source growth.18:48–23:10 · Matt pushing back 3/10 Market Landscape and Competitive Positioning Matt demonstrates deep ecosystem expertise by naming five specific market competitors across different startup generations. Mark responds with a two-axis positioning framework.23:10–25:54 · Matt pushing back 2/10 Open Source Commercial Strategy and Business Model Matt prompts a strategic discussion on open source monetization. Mark leverages his Cloudera background to explicitly reject support-and-services models as unviable.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 7.3% · guest 92.7%0:00 · Matt 7.3% · guest 92.7%3:00 · Matt 12.9% · guest 87.1%3:00 · Matt 12.9% · guest 87.1%6:00 · Matt 5.6% · guest 94.4%6:00 · Matt 5.6% · guest 94.4%9:00 · Matt 18.5% · guest 81.5%9:00 · Matt 18.5% · guest 81.5%12:00 · Matt 22.8% · guest 77.2%12:00 · Matt 22.8% · guest 77.2%15:00 · Matt 12% · guest 88%15:00 · Matt 12% · guest 88%18:00 · Matt 32.1% · guest 67.9%18:00 · Matt 32.1% · guest 67.9%21:00 · Matt 19.1% · guest 80.9%21:00 · Matt 19.1% · guest 80.9%24:00 · Matt 16.1% · guest 83.9%24:00 · Matt 16.1% · guest 83.9%27:00 · Matt 25.5% · guest 74.5%27:00 · Matt 25.5% · guest 74.5%
Sharpest disagreement ▶ 24:23 Rejection of support business models

Mark forcefully dismisses the commercial viability of pure open-source support and service models based on his experience at Cloudera.

Hardest push from Matt ▶ 13:11 Challenge on infinite integrations

Matt presses Mark on how Stemma can scale without endlessly running around connecting to limitless unbundled data stack tools.

Biggest teaching moment ▶ 5:00 Flaws of curated catalogs

Mark dismantles the traditional wiki-style catalog model, explaining why manual curation fails as organizations scale.

Matt holds his own ▶ 18:48 Ecosystem competitive analysis

Matt demonstrates high domain familiarity by citing established players like Collibra and Alation alongside newer entrants like Castor, Atlan, and Metaphor.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Data Catalogs and the Data Explosion Problem 1600 Matt asks a broad introductory question defining data catalogs. Mark takes a step back to deliver an extensive monologue mapping out the decade-long evolution of data lakes, warehouses, ETL tools, and BI platforms.
Data Governance, Productivity, and Compliance 3611 Matt briefly prompts on governance and interjects to clarify manual wiki cataloging. Mark educates him on the core failure modes of curated catalogs versus automated catalogs.
How Automated Metadata Integration Works 2600 Matt asks how metadata automation works technically. Mark provides a comprehensive breakdown of integrations across warehouses, BI tools, HR hierarchies, and Slack.
Access Control Strategy and Data Permissions 4512 Matt poses a concrete scenario regarding GDPR access control. Mark gently reframes the premise, distinguishing between dataset discovery versus underlying permission enforcement.
Navigating the Unbundled Modern Data Stack 6413 Matt cites Zhamak Dehghani and data mesh principles, challenging Mark on how Stemma avoids getting overwhelmed by limitless integrations in an unbundled ecosystem.
The Evolution from Lyft's Amundsen to Founding Stemma 2300 Matt asks about the transition from Lyft's internal project Amundsen to commercial startup Stemma. Mark describes internal CSAT scores and open-source growth.
Market Landscape and Competitive Positioning 7413 Matt demonstrates deep ecosystem expertise by naming five specific market competitors across different startup generations. Mark responds with a two-axis positioning framework.
Open Source Commercial Strategy and Business Model 5522 Matt prompts a strategic discussion on open source monetization. Mark leverages his Cloudera background to explicitly reject support-and-services models as unviable.

Statements from this episode (12)

Insight
Curated data catalogs take up to three years and quickly become outdated
“The problem with this approach is that a, it takes a very long time to value. It takes You depending on the size of your organization, anywhere from like a year to three years to actually get all this metadata in to then hand it to your users or your complianc…”
Mark Grover Nov 16, 2021 ▶ 5:55
Insight
Manual data curation fails for fast-growing, data-intensive companies
“When you have one of one or more of these two criteria met, that system breaks. You cannot rely on curation as the source of discovery, understanding, context about a catalog, and therefore there's a need for what I now call automated data catalogs.”
Mark Grover Nov 16, 2021 ▶ 6:48
Insight
Automated data catalogs narrow scope to make manual curation possible
“An automated data catalog cannot guarantee you that this is a single source of truth, but it can tell you that out of these 200 data sets related to, I don't know, pricing, These 180 are no good for you and your use case because they haven't been updated all t…”
Mark Grover Nov 16, 2021 ▶ 7:06
Insight
Analysts spend most of their time answering Slack questions, not modeling
“The expectation is like, oh, they spent all their time and like modeling and that modeling can be analytical or algorithmic. The reality is I feel like they spent all their time on Slack and they're like, Hey, what's the source of truth for this data?”
Mark Grover Nov 16, 2021 ▶ 9:36
Insight
Data catalog discovery should be public while underlying access remains gated
“One default stance that we have is that within a certain deployment and sometimes the deployment is for the entire company, certain, sometimes the deployment is for a certain line of business within the company, within a certain deployment, discovery of the da…”
Mark Grover Nov 16, 2021 ▶ 12:11
Opinion
Unbundled data tools create management problems that require automated data catalogs
“I, in my opinion, I feel strongly that's the right thing to do and you get the best of breed products, but it creates problems around management and governance that are new and need to be solved in a new way. And that's where like something like a data catalog…”
Mark Grover Nov 16, 2021 ▶ 14:20
Assertion Not checkable as stated
Amundsen was Lyft's highest-rated internal product for three years
“From that day, which is probably mid 2018 till today, this product is the single highest CSAT scoring internal product, right?”
Mark Grover Nov 16, 2021 ▶ 16:41
Assertion Supported
Over 35 enterprise companies use the open-source data catalog Amundsen
“Over 35 companies use it in open source, you know, ING, Instacart, Brexasana.”
Mark Grover Nov 16, 2021 ▶ 17:50
Disclosure
Stemma struggles to serve command-and-control organizations that broadly restrict data access
“STEMM very clearly is in the latter category, right? We do not do well in serving the command and control style organizations.”
Mark Grover Nov 16, 2021 ▶ 21:14
Disclosure
Stemma omits messaging and BI features to force Slack and Looker integration
“But when you grow, you have to integrate with the ecosystem the organization is in, and we, for example, we have no feature to have a conversation in the data catalog, right? And that's intentional because we want to integrate with Slack, and that's why there'…”
Mark Grover Nov 16, 2021 ▶ 22:45
Opinion
Selling support and services for open-source products is a bad business model
“I find that selling support and services on an open source product is not a great business model.”
Mark Grover Nov 16, 2021 ▶ 24:24
Insight
Central data teams should provide enablement tools, not act as bottlenecks
“It is the central team's responsibility to provide that tool. And it's the owner's responsibility to use that tool and then communicate directly to their consumers. So hard for me to say is governance should be managed at domain level. I'm not sure if I'm answ…”
Mark Grover Nov 16, 2021 ▶ 28:17
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.