Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So to, um, help continue bringing this home, what does an implementation of it look like sort of practically?
A This is my sigh happens. Um, I have to be really honest with you. Um, at this point in time, it looks like a Frankenstein creation because, uh, we have to stitch together a lot of technologies that exist and they weren't designed for this model of reconfiguring, you know, being reconfigured in this way. So what the foundational technologies that I have used so far are more or less the same. So for example, for your, Uh, inputs, what I call input ports is where the data actually is coming from that gets transformed. You still have your, uh, ingestion mechanisms that exist today, whether you are hooking up to some upstream event stream, or you're, you know, hooking up to some API to get the data. You're, you're doing CDC against some sort of a, um, legacy system. So those ingestion mechanisms remain the same. Your transformation code that's, it's encapsulated by this, Data quantum and does the transformation. Again, those are, um, flow, flow based programming model that you have. So a lot of people still use a spark or beam or, you know, whatever. And if you have very simplistic, maybe transformation, you're running just a federated query and the output of that query is simply your transformation. So you'll have your usual suspects, uh, transformation and, um, code, uh, work orchestration. And then on the, on the output side, you are providing, uh, kind of a, a little bit of a hi…
AI assessment note: “at this point in time, it looks like a Frankenstein creation”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q what is it? What is the data mesh? You know, it's like, I'm, I'm super excited to hear it from the horses master to speak because, uh, you know, again, like on Twitter, it's like, what does it, what does it mean? And like, somebody says something, somebody else says something, um, which I, I, I've loved the conversation, but I, I'm, uh, very excited to hear it from you.
A Sure. I think, uh, it's like an onion. We've got to kind of peel the layers, but, um, when I started, I guess, talking about it, I wanted to be very cognizant of the, uh, I guess, maturity of the idea. Right. Uh, so I started with a set of principles and I think I've talked about those principles, uh, in many forms and we can go over them very quickly. So I guess at the, at the, at the high level, it is a kind of a socio-technical approach in how we Share and manage data for analytical use cases and sociotechnical, because we cannot just talk about technology without talking about people who use the technology are subjected to the technology. So it's both an organizational structure, operational model within the organization, as well as the technology in order that we can get value from data at scale for analytical use cases. So that's kind of a tagline. And then if we come layer the onion a little bit, the most fundamental on the pinning principle around it is This idea of you can get to data, get value from it, connected with other kind of data sources, no matter where the data lives and no matter who owns it. So it's counter to bring the data to one place under one team, under one model to get value from it. So underpinning that is this idea of decentralization of ownership and control. Uh, around the access of, uh, business domains, that's the access within which an organiz…
AI assessment note: “it is a kind of a socio-technical approach in how we Share and manage data”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q And then just for the, the step. So, um, data discovery and access. So is there still a concept of like somewhere that's a catalog or is that too much centralization? If you say you have a catalog that, that knows where all the data is distributed.
A Yeah. So really good question. I think all of the models, so there is a concept of A discovery portal of some sort, right? I need to be able to see, browse, search, right? Look for the data that I need. But, uh, when it comes to the next step of implementation of these affordances, if we think about it as, you know, data is this, um, kind of piece of information without agency, without any computational ability, which is how we thought about data so far. The design we come up with is that There is a central catalog that will go and look for data in different places and add some metadata to it based on who access it and who use it, and it will create a catalog, and it will constantly sync by searching this mess of a landscape that is, um, and once we, while we need a discovery portal of some sort with search and browse, what Data Mesh Introduces is this concept of a data of quantum actually rethinking data as a unit that not only constitutes the data, but it also constitute all this computational affordances that gives that data agency and intention. What do I mean by that? So let's follow with that discovery example as a, as a data product developer that I'm creating this new logical construct. I am in fact intentionally providing discoverability abilities in, in that unit. So I am providing a set of APIs that any search or browse utility can hook up to and get discoverability …
AI assessment note: “there is a concept of A discovery portal of some sort, right?”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Maybe a last question from me because I want to then turn, we have actually a lot of questions in the chat. Um, so just so if you fast forward 10 years and like everything works out as planned, what, what does that look like in terms of, um, you know, the life of like developers and data scientists and how does that get changed?
A We all have a smile on our face. We'll have a lot of fun to do the work. So that's the first, first one. I think what I, what I, I mean, none of these, I mean, I'm looking at a bit of a crystal ball and what I have seen has happened in the past. So hopefully it makes sense to the audience, but I think one of the big changes would be, we'll move from this specialized and specialization to generalization. So some of the things that we consider specialization today, like data engineering, a large portion of what we call data science becomes basic engineering. So I think that's, uh, that has to happen. Otherwise we will never scale to this, to meet the aspirations that we have. So, so move from specialization to generalization is one. I think we will rethink if data mesh happened, When we talk about data, we imagine something very different. We talk, we imagine this kind of lively, ever-changing thing that has an agency to govern its policies and to, uh, to keep the data alive and provide, you know, APIs. We don't think about it as a by-product we dump somewhere and we build technology on top of it to get access to it. So I think, I hope that our imagination around what constitutes data changes. Um, and you know, we, we really truly become data-driven in a way that we can get access to data safely and securely, um, no matter where that data is and who owns that data, as long as, of…
AI assessment note: “one of the big changes would be, we'll move from this specialized and specialization to generalization.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Yeah. There's no lack of tools and platforms and solutions to try and tackle this whole complexity. So what, what is the industry currently doing wrong? Uh, that, uh, you know, needs to be improved.
A Yeah. I mean, you're right. Uh, I think at the industry, when I actually, I use the, uh, one of your publications where you put the landscape of all this data and I feel Kind of dizzy. Uh, when I look at it, um, it's almost unreadable. And so I do agree that what we have done right up to now is a lot of innovation in the bottom layer of kind of our stack. If you think about our technology stacks where the Data driven organizations and consumers sit at the top and the machines sit at the bottom. We've been building a lot of technologies kind of try to re solve really hard, low level problems. The problems of data processing at scale or data storage at scale and distributed kind of computing and storage at the bottom layer. And that's great. What we've got wrong is a set of technologies that scale out nicely. Um, With the organization, with the growth of organizations. So what we haven't got right so far, I mean, what we've built, I suppose, has led to Um, I don't know if actually what comes first, whether the organization comes first or technology comes first, but what we've built is suitable for organizations that are functionally divided. You know, you run your business in one side, and then you deal with the data in the other side, so you put a wall between what's data-driven and what's not data-driven, so I think that functional separation, what we've built has led or has em…
AI assessment note: “What we've got wrong is a set of technologies that scale out nicely”
Answered raw tape
D 5 · C 3 · P 4 · Cm 3 3.85
Q to this point, are there, um, can you build a data mesh with the existing tools? Like you, you mentioned like some of the existing things like battery beam flow, like all, all the things, um, uh, To be successful at drawing out a data mesh, does it mean that new tools need to be created, like new standards, uh, or can you make do with what we currently have?
A Yeah, I think, uh, we have no choice, but start with, like, if you're starting today, like we started three years ago, we used a ton of stuff that already exists, but we also built a ton of stuff. I mean, I don't like to go to every client and say, look, this is great. You can use the technology that you already have, but you have to commit to this challenge. Oh Siri. Um, you have to commit to this two, three year kind of program of building out capabilities in your platform that you just simply can't buy. So hopefully that we still have to Uh, get the technology to fill the gaps and where those gaps are, uh, to me, you know, a lot of it is around interoperability. We are, we have that wonderful diagram of tools that you have, uh, you know, in your, in your landscape, but we, if you, if you zoom in, the very few of tools actually interoperate and play nicely with the, with the rest of the tools, uh, and the standards that it gets to be created are around, Kind of the expression of, um, uh, storage agnostic modeling of the data. So I talk about, uh, in this data quantum, uh, how time access has to be, uh, uh, ever present, um, parameter in the data sharing, because the only way we can have our cake of distribution and coupling data as you wish and have the, and it's the, oh, this, this is going really bad, the analogy of cake, but what I'm trying to say is that the only way we c…
AI assessment note: “we used a ton of stuff that already exists, but we also built a ton of stuff.”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Question from Jacob around immutable data. So could you elaborate on how immutable data would be used for joining? Uh, if say a consumer wants to join with a current state of entities and reading the question, are there patterns or resources that already exist that explain best practices dealing with immutable data?
A Yeah, I think this is a big conversation to have. The data can be only immutable if we build two time timestamps into every single, every single, um, kind of representation of the data, and those two timestamps are when something actually happened, and when we process that information, our understanding of that data. That three pieces, like actual data, the event or state at the time that it happened, and the time that we process it, that piece is immutable. And you can see any of those parameters can change. For example, as an understanding of the, let's say Matt and I talk to each other, we, we have this conversation and this, um, and you read, you view it tomorrow. So you're, let's say your system is processing this video tomorrow. Uh, and let's say there was a mistake in that processing and you have to reprocess that video. Maybe the transcript, you were processing the transcript and you have to reprocess the transcripts because it was a mistake. So then what the new piece of data would be this video. Uh, the new transcript that was happening, happened at the same time today, but it was processed tomorrow, right? The next time. So, so that's what immutability really means in data mesh. Um, and you can always arrive at the state at the point in time, or look at the differences between two points at a time. If you want to join, you can always join across a point in time acros…
AI assessment note: “you can always join across a point in time across all of those data products”
Partly raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q of all of this, uh, and I'm very curious about what that, how that manifests, like, data as a product, is that, like, a, Is that like a unit of like data people and product and technologies for that product? And then there's gonna be like the next one and how do they communicate? You alluded to like a lot of this, but I'd love to get into the details.
A Sure. And it is, it is a long conversation, so I don't think there's one right place to start it from, but if we, if we start the conversation around what are the affordances or capabilities that we really want to provide at the end of the day, no matter it's centralized or decentralized, let's look at those and, and what we really want, what are, what are those affordances? And then let's go deeper and say, how do we provide those affordances with this idea of decentralization, right? So one of those first Let's imagine the experience of a data scientist, like from a data scientist perspective or data consumer perspective. The first thing that they want, they do is a hypothesis. They have a hypothesis. Can I make, I don't know, recommendations around Kind of playlists and musics based on, on music profiling, and can I do music profiling based on the mention of the music in various blog posts and so on, if I'm in a digital business, you know, streaming business. So starting from that hypothesis, what I need to be able to do is I need to be able to discover the data that I'm looking for, right? No matter where that data actually physically is. So we need to have ways of, uh, ability to discover the data and hence it needs to have been registered with, with some way. Address it, discover it, connect to it, get to the output of that data or the actual data set itself, no matter wh…
AI assessment note: “what are the affordances or capabilities that we really want to provide”