The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Dave Burgess argument clarity score 4.4/5 from 11 exchanges on raw tape · average scores: directness 4.7 · coherence 4.7 · precision 4.5 · compression 3.9 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
11exchanges match
11on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q As a quick insight, it's sort of amazing actually how important Yahoo has been to this whole, like sort of big data, data ecosystem, right? Like so many fantastic people, uh, have come out of, uh, of Yahoo.

A That's right. And so Yahoo created Hadoop for those that don't know. And so Hadoop was born in Yahoo, and then it was sort of spun out as a separate company, but within, uh, that data organization, Things like Kafka eventually came out of LinkedIn that was based on work that we had been doing at Yahoo and many other things had come. We were doing machine learning and behavioral targeting back, uh, 15 years ago, more than 15 years ago in advertising. And so we were the first in a lot of things and also did a lot of data privacy at that stage, which is only now, really in the last couple of years, coming to the forefront of being important. So The great, one of the great things about Yahoo is that there's so many people spread around the Bay Area, probably New York as well, that have been in Yahoo. So there's a really good network of folks, and, uh, it was a real pleasure to work there. I was there for six years.

AI assessment note: “That's right. And so Yahoo created Hadoop for those that don't know.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Great. And then so ultimately all of this, um, is for, Analytics. What are some of the use cases?

A Actually, yeah, yeah, actually we have many, many use cases. Uh, so analytics is a small part. We, we have like just about everyone in Pinterest is using it for analytics every day, but we also have a massive, uh, number of experiments, right? We have got our own experimentation platform where we're doing AB experiments all the time to try and improve Pinterest and try and improve our ads, our ad, uh, relevance. And so we have about a thousand experiments running, uh, in parallel at any point of time. And so we're, we're constantly iterating. Uh, so that's another use case. And then for, we use it a lot for machine learning. So we have about 80 different use cases of machine learning. And that's, you know, mostly from this data that comes into, that comes into S three and AWS.

AI assessment note: “we have got our own experimentation platform where we're doing AB experiments all the time”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So, uh, since we're talking about machine learning, what are some of, some of the tools that are being used by the machine learning folks and frameworks?

A Yeah. So in machine learning, so like two or three years ago, it was, uh, a Pandora's box. Everyone was trying out everything. And so we had, uh, with all these different use cases, people were, were, uh, you know, it started to become more of a maintenance burden. And so what we wanted to do, what we've done is created a machine learning platform that kind of glues together the best of breed, uh, open source products out there. So we use TensorFlow from Google. We use PyTorch, scikit-learn. We use MLflow, which is a, from Databricks, is a registry of machine learning models. And, uh, we also, uh, have our own format for storing features within, within, uh, Pinterest, both for doing batch learning of models and then doing the serving models. So we, what we can do is within a few hours, you could do, uh, create a, a model, train a model, and then deploy it to our production service with thousands and thousands of servers, uh, by using this MLflow repository.

AI assessment note: “So we use TensorFlow from Google. We use PyTorch, scikit-learn. We use MLflow”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q will ask the question live. So that's the best way of doing it. Um, What one more from, um, me, uh, so I'd love to dive a little bit into the organizational aspect, um, of this. How, um, is the data team organized, machine learning team organized, uh, are they separate that they work together? Is that centralized? Is that throughout the organization? How does that, how is it organized?

A So data engineering at Pinterest comprises the, the serving part, all the online systems. So that includes the online databases like MySQL and key value stores like RocksDB and HBase, and also includes Druid platform. So we have our online systems. We also have all the batch systems and analytics and experimentation platforms. So we discussed that previously. And then the machine learning. Piece as well as a, as a platform. And so we provide the, the glue and the underlying infrastructure, uh, with TensorFlow and PyTorch and all the, all the other components available to our internal customers. And so with machine learning, machine learning is actually distributed everywhere. We have, uh, Many, many machine learning engineers and many, many use cases. And those mission machine learning engineers are embedded within each organization. So there's quite a few within shopping, quite a few within our advertising business and also different parts of our product and trust and safety. And then we have a separate product analytics and data science team as well that uses all the capabilities of, of data engineering.

AI assessment note: “machine learning engineers are embedded within each organization”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q news today. The open sourcing of, of, of query book. Um, so I'd love to, uh, take a step back from this specific tool. And, um, if you could help us, uh, you know, give us a, a, a broad picture of the data engineering analytics stack at, um, at Pinterest, what, what do, what do you use? You know, the various, the various tools and the various systems. Sure.

A Sure. So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. We're moving more and more to spark now, uh, because it's faster, uh, but Hadoop still scales really, really well. Uh, we have a query platform, which is Presto and Spark SQL. We used to use Hive, but we're migrating of Hive to Smart SQL, so we just have the, the two main engines. And then we have a workflow system on top of that that's built with Airflow, which is also open source. And the way that we get all this data is using Kafka. So Kafka, uh, has, uh, clients in all of our serving systems and gathers the data from these clients and then pipes it to the backend. Now we're, all of contrast is, is running on EWS. So we put all of this data into S three and it's a massive idea. We have more than 400 petabytes of data. Uh, And so we, we spend quite a bit of time, uh, putting it into different buckets, and we can show that the buckets are partitioned correctly, so that the, these engines will perform at scale. We also, uh, it's important to put it in the right format as well, and so we want to try as much as possible to put it into a parquet column or format, uh, for these query engines to be even faster.

AI assessment note: “for the analytics itself, uh, we focus on, uh, using Hadoop and spark.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q source, uh, frameworks from like Druid to Kafka to Flink. Um, does an organization like Pinterest work directly with the open source, or are the vendors involved as well? Because Kafka's Confluence, Drill, Flink is, I guess, was part of Alibaba now. Um, I think, like, do you, Do you do, because you have the engineering resources, do you take the open source and work with it directly? Open source?

A Yeah, we do. I mean, we're fortunate to have the number of engineers to do that, and one of the reasons why we, we work directly either with the open source community or the companies that are also working on that is that we want Pinterest to be up all the time, and so if there's a problem, we want our engineers to be able to go and look at the code and be able to fix it and deploy it into production. And have very, very little downtime. And so if we were dependent on a third party for that, then we, maybe the, the downtime would be a little bit longer. And we also want to develop our own features and our own capabilities that are needed for Pinterest. And so building on top of open source and then, then putting it back into the community is a, is a great way to, for the, for both for the engineers in terms of their careers, And what they enjoy doing. And then also great for the community. And it's also great for Pinterest because we stay up to date with what the open source latest versions are and get the benefits of that as well.

AI assessment note: “Yeah, we do. I mean, we're fortunate to have the number of engineers to do that”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q to close up with, um, like a rapid fire, sort of like bonus question type things, which are sort of unrelated to, to Pinterest, uh, Uh, specifically. So, uh, first question, what's new data trend or product in the overall ecosystem? So maybe use it, maybe you don't use the product. Yeah. Um, are you most excited or, or, or even just curious about that's come up on your radar?

A Well, I'd say there's, there's three, if I have to choose, choose one. And so I'm, I am really excited about machine learning capabilities that are being, that's being advanced very, very quickly. And, uh, it's great to see lots of companies, uh, basically providing offerings for, for improving what we do with EMMA. And to get to a point where we can have people that are data scientists that don't even code, that can just build models and be able to deploy those to production. So that's really what we want to, to get to within Pinterest as well. And so this is an area that is really hot. It's been hot for several years now and continues to be hot and it's, it's changing a lot. The next area I would say is, uh, data privacy. I think data privacy is really, really important to, to everyone. And we all want to have our, our data protected and know that it's being used in the right way and to have more controls around, uh, your own data. And so there's quite a few products that are coming out now that are doing data privacy compliance. They're doing data lineage. They're doing data leakage detection. Um, you know, to help you with GDPR. And so I'm kind of excited about that area because there's a movement, I think over the last, certainly over the last year within the US, uh, earlier in Europe and earlier in California, but now in the whole of US about what data privacy really mean…

AI assessment note: “I am really excited about machine learning capabilities that are being, that's being advanced”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Very good. All right. A couple of last ones. Um, question from Gedalia. Can you please speak more to use of Kafka and airflow? I guess you, you probably covered Kafka. Um, Airflow as a scheduler, like, do you use Airflow as such, or did you build something on top?

A Well, we've, we've done both. So we actually had an existing workflow system that was built in-house about seven years ago, and then just over the last year, we migrated to Airflow. We're still using the APIs for the existing system so that we could migrate, but we have over 2000 workflows running in production. And, uh, so my migrating that alternative Airflow was difficult. We've also built some UIs on top of Uh, airflow. So we have one, there's some internal names here, but there's, there's one that basically, uh, allows you to compose a workflow really easily just to like a drag and drop interface. And then we have another one for being able to take many, many workflows and optimize which, uh, which jobs to run and when. So, the, when you've got many thousands of workflows, you can start to decide, well, I'm not going to rerun a job in order to compute for this workflow, and that's why I'm going to, going to run it just once, and it can be used for both, uh, downstream. So, a couple of tools, which we're also considering open sourcing, uh.

AI assessment note: “Well, we've, we've done both.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Okay. I was actually going into that next slide. What are some of those examples of, uh, of machine learning? I know, I know that the visual search at some point was, was a really important, um, sort of breakthrough.

A Yeah, it is. And visual searches, it's actually, if you do like a, a comparison between the visual search and Pinterest, so just take a photo or something or, or get a picture and then upload it and see, see what the results are between Pinterest and Google and Bing and other search engines. You'll find that Pinterest is usually the one that comes out of top. So we're really, really proud of that. And, and we use a lot of the data that we've, we've got in order to do that. So our Pinterest users, which we call pinners, Are creating these boards of pictures and videos, and usually the boards have similar images within them, or at least a similar topic, and so we can use this underlying data that's been organized by all our users to come up with really, really high relevance results, and so we use it for things like classifying pins, like what, what, ah, which is the picture, what is it within the pin, so we can identify objects within, within an image or within a video, And we can also, uh, identify whether it's content that is safe to show our users. We want to have safe content. So we remove, uh, content that, that, uh, doesn't meet that bar. And we also have recommended pins. So one of the things people love about Pinterest is what they love being inspired and, you know, finding something that they enjoy and then finding something else they enjoy and continuing to continue. A…

AI assessment note: “we use it for things like classifying pins, like what, what, ah, which is the picture”

Partly raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q uh, let's, let's see, let me do this, uh, sort of real time. Um, where was this one? Um, can you talk about, to the extent there's nothing confidential, can you talk about your experience with Droid so far? Um, how did the migration from HBase to Droid work out? Also curious if you looked into additional data stores such as ClickHouse or Apache Pino. There's a question from Rohit.

A Yeah. And so, yes, we, we, we did a migration from HBase. We were using HBase for analytics and for many other things as well. And so we, uh, we started out with, Using Druid for one use case within Pinterest and proving out that that worked really well. Being able to slice and dice with kind of millisecond or second latency, uh, was really good. And so it works really well for this first use case. And then we decided, okay, we're going to bring this into data engineering and be part of the platform. And then we have a bunch of other use cases. Uh, we were also starting for experimentation platform. We had so much data within our, our HBase cluster that it was getting really slow and very, uh, difficult to manage and overhead for our, our engineers. And so we decided to migrate that to, to Druid. And so the maintenance cost has gone way down. The latency has gone way down. The actual cost of running the infrastructure way down. Uh, so really, really happy.

AI assessment note: “maintenance cost has gone way down. The latency has gone way down.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q Great question from Pierre. How do you observe and monitor the data flowing across these different systems and tools?

A That's a, that's a great question. Uh, so there are different kinds of monitoring. There's, there's kind of systems based monitoring where you want to, uh, see that each system is, is performing well. There's also monitoring of, uh, data quality going from one system to another to make sure that you're not losing any data along the way. Uh, And there's, there's other kinds of monitoring. So we use different tools for the different, different things. So we actually, uh, for kind of data monitoring, you'd want something like Splunk or another kind of tool like that. We built, we built our own, it's called Goku. And we're actually contemplating open sourcing Goku. It's a time series database, uh, which you can, you can do time series analysis. And so it's great for production monitoring system. Uh, people are interested in that, then let me know, and, uh, we might consider open sourcing that, and that was based off, uh, Facebook's and gorilla, uh, compression. And so we, we, we built a time series database around it.

AI assessment note: “for kind of data monitoring, you'd want something like Splunk or another kind of tool”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.