The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mark Grover no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And maybe double clicking on the, on the latter part. So governance, like why does governance matter?

A Yeah, totally. I mean, governance matters, uh, in, in those two ways, right? The first one is around productivity. So your analysts and your product managers and anybody who's going to make use of data during their job is effective at using that data. They may have skills like writing, being able to write a SQL query or interpret a dashboard that are generic, but they don't have the organizational context. And on governance, on the other side, which is compliance with regulations, Those are actually, um, the domain dependent. So there's the regular, uh, regulations here, like GDPR and CCPA that apply to almost all organizations. And then depending on the domain, if you're a financial company, like there are many in New York, you may have certain other regulations that you have to fulfill. So these could be around, I'm reporting this data to a certain auditor, and I want to be able to prove that there's no Manipulation that's happening outside of what I already know during the process, or this particular system is regulated and should not have any sensitive data outside of these bounds, right? So understanding what your data is in that system and, uh, what are the bounds there so you can alert, um, our other style of compliance and governance requirements that are pretty top of mind for users.

AI assessment note: “governance matters, uh, in, in those two ways, right? The first one is around productivity”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q It's literally a manual notes, right? They're just like, okay.

A Yeah. And so the problem with this approach is that a, it takes a very long time to value. It takes You, uh, depending on the size of your organization, anywhere from like a year to three years to actually get all this metadata in to then hand it to your users or, uh, or your compliance folks and be like, let's use it, right? The second one is the moment you write it in any company that's evolving, which is every company, this gets out of date. It's just a matter of whether it gets out of date tonight or if it's, or if it gets out of date three months from now, right? And so this had been the status quo of solving this problem. And what has happened is There is a newer creed of companies that are all cloud, fast-growing, product-led, and it didn't start yesterday. It didn't start in 2017. It actually started before that. These companies are growing at such a pace, and there's growth in two dimensions. One is the amount of data they have, and the second is the number of people that are going to use this data, right? And in these two companies, when you have one of one or more of these two criteria met, that system breaks. You cannot rely on curation as the source of discovery, understanding, context about a catalog, and therefore there's a need for what I now call automated data catalogs. An automated data catalog cannot guarantee you that this is a single source of truth, but i…

AI assessment note: “Yeah. And so the problem with this approach is that”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yeah. So how does the automation work? Is that, uh, so the metadata gets automatically captured from the various sources and like sort of like pre-populated? Is that the idea?

A Yeah. And so let's dig in, um, the main systems that you need an automated data catalog to integrate with are your data warehouse. And this could be a data lake, data warehouse, things of that sort. And from there, you get information about what data sets are present or what tables are present. I use those two terms interchangeably as well as how they are being used. So our data warehouse usually has access logs that you can take out and parse them to generate like, oh, it's often joined with this other thing. And, You know, Jack often uses this data, so on and so forth. The second system we integrate with, uh, that an automated data catalog needs to integrate with is the BI tool. And so it has all your information about what dashboards are viewed, what dashboards are built, being built from what data sets, who are the people who view those dashboards. And, and the third system is usually an HR or a team hierarchy system. So you can figure out that these people are on the same team and therefore be able to suggest interesting data or metadata for them to use based on what your peers are doing. And then there are transformation systems. You know how often or when a data set is usually transformed into another. You can understand the lineage between one data set and another lineage between one data set and a dashboard. And lastly, this category, the last ones are collaboration sy…

AI assessment note: “Yeah. And so let's dig in, um, the main systems that you need”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q like this, uh, you know, um, uh, Uh, Atlan, uh, you know, like a number of like different players, which are all, you know, in our data landscape, uh, for those that want to zoom in on those, uh, on those categories. Uh, how do you, how do you think about, about, uh, those, how do you position, um, and, uh, how do you, you know, eventually win that market?

A Yeah, totally. Um, so a few things here. One is that, uh, there are kinds of catalogs that are more Command and control catalog. And this is a product, but it's actually a manifestation of the culture of the company. So if you are in a company where data is very tightly restricted and you don't want to democratize access to certain kinds of data for others to use and make decisions with these catalogs work really well. And so these are you, you curate them, but you also end up creating these heavy workflows that requires like going to a university to understand how to manage and orchestrate these workflows. But also a team of people who are on the other side of these workflows sort of approving anything as small as like updating the description of a column to as big as access requesting, granting the request to access, right? Some of these you would have in any company, but some honestly like only exist in companies that have a very sort of command and control culture around data. Then there are companies that want to actually evangelize, democratize the right data for decision-making to everybody in the company, right? And the Sort of the first big difference that I, I, um, share in the offerings is the command and control catalogs, which often curated and workflow based while, uh, the other ones being more demo democratized and like, um, automated in their, uh, both the metad…

AI assessment note: “STEMM very clearly is in the latter category, right? We do not do well in serving”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q sort of limitless, right? And especially, uh, you know, we had, uh, Jamac Um, uh, the last event we talked about the data mesh, right? Like this is only going to get more complex and decentralized if, if we believe that the world is going towards the data mesh concept. Like how, how do you ensure that you're not like constantly running around trying to connect to the next source?

A Yeah, totally. I, I think it's a great, a great point. Um, I would say actually that's one of the things is that's the primary reason why like you need a product like this. First of all, that the modern data stack is not a bundled one. It's an unbundled one. So you have your data storage system, you have your ETL system, you have your transformation system, you have your BI tool, and none of these tools are packaged by like IBM and you're buying one behemoth IBM platform, right? You're, you're getting these products That are unbundled and that actually leads, I, in my opinion, I feel strongly that's the right thing to do and you get the best of breed products, but it creates problems around management and governance that are new and, uh, and need to be solved in a new way. And that's where like something like a data catalog can really help. Um, but the, as, as, uh, a business owner running this business and having to integrate with all these systems, uh, it is also something that We pay the cost for, and, uh, thankfully the cost hasn't been huge. Uh, I do think it's the right thing to do for us to evolve our integration when new systems emerge and our guiding light is our customers. And while we can choose at, to some degree, what kinds of customer and what verticals and what industries and what markets we focus on, um, we find that most of the customers that want a modern data…

AI assessment note: “the modern stack has actually not as exhaustive of an option”

Partly raw tape D 3 · C 5 · P 4 · Cm 3 3.85

Q Then coming back to the Munson story. So, uh, you and, uh, others started The, the open source project at Lyft, like, so how long did that last? And then maybe walk us through the birth of, uh, Stemma as a separate company.

A Yeah, totally. So continuing with the story, you know, like this is a big problem that left too much data, too many people wanted to query the data. No, one's got any clue what's going on. And, uh, this idea of a data cow, an automated data catalog stuck with me and we started building this product internally. With the target persona being an analyst and a data scientist to just really make them effective. Um, we did a hackathon, which was like a very quick throwaway version just to see if we can de-risk the project and then build the real project over the next few months. We launched an alpha, uh, off this project to 10 users and they were like, this is the best thing we've seen. We quickly got their feedback further, changed it to launch it to beta more publicly to everybody at Lyft. And at Lyft, From that day, which is probably mid 2018 till today, this product is the single highest CSAT scoring internal product, right? And so, has 750 users every week. Lyft only has, um, 250 or so data analysts, data scientists, right? So there's like all these other people who have started using this product because the barrier to entry to them to using data has been further lowered, right? And so you have these 750 users using it every week.

AI assessment note: “So continuing with the story, you know, like this is a big problem”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.