“An automated data catalog cannot guarantee you that this is a single source of truth, but it can tell you that out of these 200 data sets related to, I don't know, pricing, These 180 are no good for you and your use case because they haven't been updated all that while. Nobody else in your team uses it. There are no dashboards built on top of it, all that kind of stuff. And then these are the 20 that you should dig into. And it won't tell you like, Hey, this was a guaranteed source of truth, but like, it will really help you focus down on the ones that make sense. And then you bring in the rest of the curation to get the last 80, you know, last mile benefit here.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mark Grover
Insight
Curated data catalogs take up to three years and quickly become outdated
“The problem with this approach is that a, it takes a very long time to value. It takes You depending on the size of your organization, anywhere from like a year to three years to actually get all this metadata in to then hand it to your users or your complianc…”
Mark GroverNov 16, 2021▶ 5:55Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Insight
Manual data curation fails for fast-growing, data-intensive companies
“When you have one of one or more of these two criteria met, that system breaks. You cannot rely on curation as the source of discovery, understanding, context about a catalog, and therefore there's a need for what I now call automated data catalogs.”
Mark GroverNov 16, 2021▶ 6:48Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Opinion
Selling support and services for open-source products is a bad business model
“I find that selling support and services on an open source product is not a great business model.”
Mark GroverNov 16, 2021▶ 24:24Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Insight
Analysts spend most of their time answering Slack questions, not modeling
“The expectation is like, oh, they spent all their time and like modeling and that modeling can be analytical or algorithmic. The reality is I feel like they spent all their time on Slack and they're like, Hey, what's the source of truth for this data?”
Mark GroverNov 16, 2021▶ 9:36Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Insight
Data catalog discovery should be public while underlying access remains gated
“One default stance that we have is that within a certain deployment and sometimes the deployment is for the entire company, certain, sometimes the deployment is for a certain line of business within the company, within a certain deployment, discovery of the da…”
Mark GroverNov 16, 2021▶ 12:11Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Opinion
Unbundled data tools create management problems that require automated data catalogs
“I, in my opinion, I feel strongly that's the right thing to do and you get the best of breed products, but it creates problems around management and governance that are new and need to be solved in a new way. And that's where like something like a data catalog…”
Mark GroverNov 16, 2021▶ 14:20Fireside Chat: Mark Grover (Co-Founder & CEO, Stemma) with Matt Turck (Partner, FirstMark)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.