The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jack Berkowitz argument clarity score 4.5/5 from 9 exchanges on raw tape · average scores: directness 5 · coherence 4.8 · precision 4.3 · compression 3.9 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
9exchanges match
9on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So then clients license the data. How do, how do they consume those, those benchmarks?

A Yeah, clients can consume the data in a lot of different ways. So we have a SaaS application that they can do it. So we have tens of thousands of clients who consume it through a SaaS application. They can also license the data directly. And so we have distribution partnerships, um, either through systems integrators or directly to clients. And we're on some of the data exchanges, you know, like the AWS one, or the Snowflake one, for example. Um, and then we also have APIs, and so we have a lot of clients that will hit APIs, and then access data from us via APIs. Now all of that's permissioned by the consumers, but it's a, a really big thing that, that, that goes on.

AI assessment note: “clients can consume the data in a lot of different ways”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And what about, last question on the vendor, uh, front, like, I don't know, data quality, data lineage, um, any of the cool startups or more established companies in the mix?

A Yeah, so this is the one that, that gets me. So, uh, uh, Mitt Walia and I are good buddies now. Uh, we are, we went with Informatica, and the reason why we're on Informatica is very clear. Um, Not, as Spencer was saying, not everything is in the cloud. And unfortunately for startups, they're making decisions about where to build, and so they're building in the cloud. Informatica has built for the cloud. They've also built on-prem. And so they can help us with all the on-prem systems. We run data centers that are massive, right? And our data centers, we need to be able to connect to old sources. I need to connect to DB two. I need to connect to Oracle. I need to connect to other systems. And so Informatica solved our needs there.

AI assessment note: “we went with Informatica, and the reason why we're on Informatica is very clear.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What about, um, machine learning? Are you, like, a SageMaker shop, or, like, what do you use? Or, like, homegrown, open source?

A Yeah, so, ah, we spent a long time looking at different things. So, ah, we do some work in SageMaker. We don't want to put constraints on folks, but, ah, we do use an awful lot of MLflow for our MLOps. Um, we do SageMaker work. We're working with this really interesting company called Robust Intelligence in California that's helping us with, um, bias monitoring, compliance, supply chain attacks on machine learning, which is, um, a scary, scary part of machine learning. Anybody can sit down and write pip install, and before you know it, you've got a corrupted library. Um, and so we're working, uh, with a lot of different vendors to build, uh, an MLOps pipeline that works for us.

AI assessment note: “we do some work in SageMaker. We don't want to put constraints on folks”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q got so many Follow-on question to, to, to this, but like, let's, let's keep going down the, the, the CDO and the organizational part of the, of the discussion. Um, so, uh, do you have teams, I guess, what, what does data, uh, from an organization, organizational perspective look like at ADP? Are you, do you have a centralized organization? Do you have a decentralized organization? How does that work?

A Yeah. So, we practice a hub and spoke. So, um, Ah, my core team, and I have a core team that builds data platforms, really think about them as the hub. But then we have, ah, development teams. In fact, one of the development teams is headquartered here, around the world. And those development teams and those technology teams, but also it could be the finance team, or the, our own HR team, or our sales team, they're building their own data pieces, but they're doing it instead of as a set of Separated systems actually declaring semantics and attaching their data into the data platform. Uh, at the same time as they're building independent models, but then sharing models. So, you know, a good example of that is benefit codes. Everybody gets benefits here. You have your medical benefits. Every company actually codes those differently. So one of the teams built a benefit code classifier. Then we advertise it. On, and then other teams can use it. And so we have a, a bit of a data ecosystem, but also an API ecosystem, or a machine learning API, a reasoner ecosystem across the company. And that's growing leaps and bounds. So when we started this, You know, I said, hey, let's do a hub and spoke. My boss, Don Weinstein, said, hey, let's do a hub and spoke. So we had a hub, and my team was the first spoke, right? But we didn't know if we were going to have a lot of other spokes. But here w…

AI assessment note: “So, we practice a hub and spoke.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And then, by all means, like, stop me for whatever is confidential or not, but, um, I don't know, what would you use for BI? Are you a Tableau shop? Are you a looker shop? Are you trying new things?

A What do we use for BI? So, as you know, I have a background in BI, so I'm a little bit biased. Um, but no, inside the company, we use Tableau. We also use Power BI quite a bit. Um, but we also build our own BI system. Right? And so, and we have two flavors of it. So the one that we give to our clients, uh, the, the SaaS application for people analytics, we've built that from ground up ourselves. Now it runs on some commercial databases, but we've built that ourselves. We've also built our own Pretty incredible reporting engine. So reporting, which is something that seems a little bit boring and everything, actually is the lifeblood of enterprise companies, right? So our clients use reporting more than you would, it's amazing. We couldn't get the performance characteristics that we wanted in a commercial.

AI assessment note: “inside the company, we use Tableau. We also use Power BI quite a bit.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, I think I read somewhere in preparation for this that you created, you had an ability to create ontologies automatically, is that, is that?

A Yeah, yeah, so we have, um, We took a look at it. You and I know each other from years ago where we worked in ontologies, and the problem in ontologies is they're super powerful. The problem is they're super hard to build. Um, so we took a look at it and we said, well, wait a second. Can we do, and can we solve the classic attachment problem? In other words, can we have a very shallow top-level ontology, but have the data automatically be processed through machine learning and then attach To the top nodes, and sort of become a self-organizing system. And so, over the past couple years, we built something we call the skills graph, or skills cloud, which is essentially that, as opposed to other attempts at this, and skills is the hot topic in HR technology and recruiting and everything. Instead of having it all done by hand, it's, pfft, 95% automated. And so as new skills come into the market, new job titles come into the market, they're realized, skills are attached, and then you get a massive graph that then you can navigate and give, you know, build applications on top of. We can even tie the compensation data to it, because the compensation ties to the graph. So I can tell you, for example, if a nurse gets a new certification, what's the incremental value of that to her paycheck over the next year? And so that's what we do. Um, uh, automated.

AI assessment note: “Instead of having it all done by hand, it's, pfft, 95% automated.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And you know, there's, there's been this, uh, ongoing discussion, um, around the data mesh, right, and this idea of, like, data products, and like, is, is there some parallel to this?

A Totally. Totally. So, the declared semantics are the piece that we've taken from the data mesh. And the responsibility for the data quality and the data freshness. So if a, if a spoke is going to contribute their data, they gotta stand behind it. They can't say, well, I put the data in, but, you know, this month we're not going to fill that column. Right? That's a responsibility. The piece we didn't bring from the data mesh was the notion of federated query. Um, because the problem with federated query is the latency of the query is the worst performing member of the federated group. Right? And so we still have a centralized data repository, particularly for analytics and machine learning, or those types of things that aren't sub-second. Um, but then the declared semantics are all federated, and the, the data governance is all federated.

AI assessment note: “the declared semantics are the piece that we've taken from the data mesh”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Since we're on the topic of startups and big companies and that's very much your professional story, like any, any, uh, lessons learned or anything that, uh, surprised you, uh, transitioning from being a startup Person to, uh, launch company exec?

A Yeah, so, so I, I met Matt when I was in a startup, and he wasn't a startup. He, he, he wasn't a startup, and he had moved and sold, right? Um, so, so there were two things that really surprised me. One was, I wanted to understand scale, but I didn't understand the scale in terms of how fast things could be sold and have to be supported at massive, massive things. So, I can launch a product today and have 5000 clients on it in a year. And it's, and so there's a, a scale factor of big companies that I was surprised at. Um, which is great. Um, the opposite side of it though is, is you can do very innovative things in big companies. Uh, very, very innovative things if you stay focused On wanting to innovate. And I've had, you know, three phases of my career, but the past 10 years, 11 years, being in Fortune, 200 companies, um, I've gotten to do some really, really cool things. You know, some of the most cutting edge stuff in the industry, and the companies have embraced it, as opposed to said, no, no, no. And that was the biggest surprise to me, that companies want to innovate, and they really do.

AI assessment note: “there were two things that really surprised me. One was, I wanted to understand scale”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q And is it like a deep learning kind of model where, where you, you can't really tell which factor leads to that prediction?

A You know, we have, we have a variety, just been around a bit, we have a variety of different models. Explainability is important for certain things. So, I remember once, uh, somebody who's the CEO of one of the hyperscalers saying to me, well, that's conventional machine learning. Like, yeah, but I can explain it. Right? Um, so we do. We, we have work in, ah, the turnover predictor, for example, uses 57 attributes. Um, some of those attributes are, are internal data from companies, but so, some of it's external data. So, for example, commute patterns. So the commute patterns actually shift over time. So it's important we retrain this, right? Before COVID, after COVID, you have retrains on those models. Now at the same time, we, we, we build transfer layers on top of large language models. And maybe we can't explain all those at times, right? Um, but we, we, we have a range of technologies and techniques that we use, because what we're really trying to do is go solve a problem. We're not, we're not, you know, bound to any one technique, or even any one database for that matter.

AI assessment note: “maybe we can't explain all those at times, right?”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.