Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What about, um, machine learning? Are you, like, a SageMaker shop, or, like, what do you use? Or, like, homegrown, open source?
A Yeah, so, ah, we spent a long time looking at different things. So, ah, we do some work in SageMaker. We don't want to put constraints on folks, but, ah, we do use an awful lot of MLflow for our MLOps. Um, we do SageMaker work. We're working with this really interesting company called Robust Intelligence in California that's helping us with, um, bias monitoring, compliance, supply chain attacks on machine learning, which is, um, a scary, scary part of machine learning. Anybody can sit down and write pip install, and before you know it, you've got a corrupted library. Um, and so we're working, uh, with a lot of different vendors to build, uh, an MLOps pipeline that works for us.
AI assessment note: “we do some work in SageMaker. We don't want to put constraints on folks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Okay, so that's Snowflake. Um, let's talk about Fivetran and ETL, and maybe just in one minute, uh, what is Fivetran and what is ETL? We had George, uh, Fraser, the CEO, uh, at this event online during the pandemic, but, uh, maybe as a refresher.
A So, okay, so five trend is like the far left of this diagram you all just saw. Um, it is, you got a bunch of data in third party sources or in data warehouses. You want to centralize it into your central warehouse, be it Snowflake or Databricks or BigQuery or whatever. Um, the way you had to do that before, the first data team I worked on in Silicon Valley did this. You had to basically write a bunch of stuff to scrape things out of APIs of these services. So you'd have to basically hire an engineer to scrape stuff out of Salesforce's API. It was an enormous pain. Uh, the API is actually decent, but it's still like, you have to manage it when things change, you have to fix it. Fivetrend does it all for you. So Fivetrend is basically like, pull data out of various services, they connect to a couple hundred now, I don't know how many. Um, you push a button, you say sync the data from this service into your warehouse, and they just do all of it for you. So it's essentially like a copy it from thing that doesn't quite look like a database into a database, and then you can build all the stuff you just saw on top of it.
AI assessment note: “Fivetrend is basically like, pull data out of various services”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Uh, and it's, uh, a company that's been around for, like, about 10 years, and is actually, uh, as far as I know, one, one of those, like, companies are over a hundred million in, in, in revenue. So what's, what's the case against, not necessarily them, but, like, that space?
A So the, the, to me, the potential question there is, it's like a little bit of an awkward thing for a company to be sitting as this middleman, where, what they essentially do is they sit in between, take Salesforce and Snowflake, they sit in between those two, They have to maintain a connection to Salesforce's APIs. When Salesforce changes it, which Salesforce doesn't care what Fivetrend does, like, I mean, Fivetrend's maybe big enough now that they do a little bit, but like, third party services aren't gonna go call Fivetrend and be like, hey, we're changing our API. Fix it. Um, so Fivetrend basically has to maintain that. The way they also get data out of it is, is they scrape it. They, some, some companies provide ways for like, we are making changes. They push it to, to other services. But a lot of times it's just like, run a script against the API, check the differences, and like, put the thing back into the database and batch. That's kind of a clunky way to do this. Like, it would be sort of more sensible, uh, if you could design this in a perfect world, that Salesforce just writes it to a database. Now obviously they didn't do that way back when because nobody wanted it, but now it's become such a thing to say, hey, we want our database, our data out of your SaaS software into a database, not for the sake of migrating away from Salesforce, but for the sake of all the ana…
AI assessment note: “it's like a little bit of an awkward thing for a company to be sitting as this middleman”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And obviously a lot of the excitement around dbt from, you know, market and investment respective has been that dbt labs, the company and dbt core, I guess the project has, has, has owned this transformation layer. Do you want to explain what dbt actually is and what it does?
A dbt is the T in ELT. I was just talking about how the, this re-architecture. So dbt does not ingest data into your warehouse. It transforms it once it's in your warehouse. The funny thing about that is that if the data is already in the warehouse, then the only thing that you need to do to, uh, transform that data is write SQL, and you can do that in a couple different ways. You can, like, Create a view that abstracts some business logic, or you can create a table that stores the results of a query, or you can incrementally update the data in a certain table, but what dbt does is it allows data analyst, analytics engineers, data engineers to write these small bits of logic, modular business concepts, and slowly build up a Directed acyclic graph, a DAG of these concepts, and you go from left to right, and you start at the source data, and you slowly build up all of these concepts where you, and you eventually get to a place where you're dealing with business concepts that can be productively analyzed, and dbt is the framework that allows you to both, like, express all of that in code, but then also to run it against your database and, like, materialize all that stuff.
AI assessment note: “dbt is the T in ELT. I was just talking about how the, this re-architecture.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you want to talk about what decision intelligence, decision automation, what, what all of that means?
A Yeah, definitely. So I think it's, it's kind of easy to start with the, the fact that billions have been invested in, uh, you know, both data science and AI. And, uh, and this is a fact and like people have been talking about this space for awhile, you know, you built on, uh, kind of the digitization era, which was like, make sure we have event data for everything. Right. And then from there you went to kind of like BI. So, you know, what's happening in my world, right. Can I understand what happened in the last like week, et cetera. And then you went to data science, which was fundamentally like answering the question of like, what is possibly going to happen or like predictive modeling and really where decision intelligence is like the next layer on top of that. Uh, so it is the space around what, what should I do about it? Right. And so that's actually why we picked the name next move. So what's your next move? Uh, and you know, it's, it's really that space. And so, um, if you think about that, it's kind of the next evolution of a data stack. I like to think about a simple example could be like a subscription box, like I am a user of Stitch Fix or Birch Box, right? And these types of companies, they may have a data science group that's working on what is the likelihood that I'm going to like an item, right? And so they're trying to predict, right, if I'm going to like this s…
AI assessment note: “decision intelligence is like the next layer on top of that”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q if I'm a technology Personally, the company and I like data, data problems. How do I figure out what I use for different problems, right? So you have, you have key value stores, you have, uh, document databases, you have relational databases, you have graph databases. How does that, how does, how does that, or how do I choose the right tool? Uh, and how does it all work together?
A Yeah, so it's, it's actually pretty simple, right? You start with the shape of the data, and you look at the query workloads that you want to run in that data, right? And so if that data is very tabular, if it's a payroll system, and you want to record all the individuals, and they're all well structured, all of them have exactly the same schema, right? And you want to calculate average salary, and blah, blah, blah, stuff like that. Awesome. Relational database. Go, right? Or if you have a bunch of JSON documents sitting around and you don't really care how they're connected, right? Document database, go, right? Or if you have a data set that is highly complex, that is evolving, where the business requirements change, where the values and how things fit together, like a shopping cart, which is connected to order items, those order items are connected to product, which sits in a product hierarchy, and how things fit together, a graph database is your best bet, right? And so, so that ends up being kind of the The first go to move, look at the shape of the data, and then the queries you want to run on that, and that'll clue you in very rapidly, kind of where you should try to evaluate first.
AI assessment note: “You start with the shape of the data, and you look at the query workloads”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What does that, what does that mean for people that don't spend time in that world?
A Yeah, it means you don't take any, um, net exposure to the market, meaning, uh, every hundred dollars of long positions you have, In different companies. You, you also have to have a hundred dollars of short positions. Um, so we're never taking any, um, directional exposure to the market. And, um, you know, for example, in March, 2020, when the market fell 30% in about 20 days, um, we were down one and a half percent. Uh, so whatever's happening in the market would not give you an indication of how well we're doing. Whereas most funds that you heard of, or Cathie Wood, ARC, these types of funds are very exposed to basically risk factors in the market. So we built a market neutral strategy and it's a sophisticated strategy that's really suitable for institutional investors who already have tons of market exposure. They don't need us to buy stocks for them. They want a sophisticated uncorrelated strategy. And so that's what we've built.
AI assessment note: “it means you don't take any, um, net exposure to the market”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q You know, which was going to be like while we're on the topic, going to be my next question, like, what license is Clickhouse under, and how do you, um, yeah, how do you plan on, on, on working with the, the hyperscaler, the cloud providers?
A So it's currently governed, governed under an Apache two license, which, you know, is extremely permissive, allows for redistribution, modification. You can build a managed service around it. It obviously aids in, in growth of the project and popularity, but it comes with obvious risks that I just described. There are a variety of different licenses that we have considered, um, that we have not yet adopted, like AGPL version three, uh, or SSPL, um, which other open source companies have, have deployed. We feel that today the right decision for the community is to stay with an Apache two license, and we're excited about that. Like I said before, you know, we're not moving away from open source, quite the contrary. We're going to double down and double the size of the team of the core contributors and recruit new people into our company that understand the technology and that have been contributing to the projects in the past. Um, and at the same time in parallel, we are going to build A multi-tenant managed service in the cloud, which will inevitably be deployed on a variety of different cloud platforms, whether it's AWS, GCP, Azure, whether we go to China, like I've done in the past and partner with companies like Alibaba and Tencent and enable them to, you know, take all of the orchestration framework that we develop, which will in some of which will be open source, some of wh…
AI assessment note: “So it's currently governed, governed under an Apache two license”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So I'd love to, uh, go down memory lane a bit and go back to the origins of the company. So it's, it's started, the open source project started at Yandex, right? Which is the, the, the Google of Russia. I'd love for you to, uh, tell us the story.
A Sure. So it first came on my radar a few years ago, uh, the technology itself, um, as I started to observe its increase in popularity, uh, in the market. And earlier this year, I was introduced to Yandex, the CFO specifically, who then quickly brought in the founder and CEO of Yandex, the co-founder, Arkady Velos. And that was literally the start of the calendar year back in early January. And Arkady and I Started to romanticize about what it might look like to spin ClickHouse out of Yandex, as well as the core engineering team and the creator of the project, Alexa, and form a new company around this extremely popular open source database technology. And I immediately reached out to two investors who I had worked with in the past, specifically Mike Volpe at Index Ventures and Peter Fenton at Benchmark, two people that you, I believe, have interviewed in the same forum in the past.
AI assessment note: “Arkady and I Started to romanticize about what it might look like to spin ClickHouse out”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So why do those, um, issues occur at a very simplistic level? I mean, what's the range of things that tend to go bad leading to the quality issue?
A Um, it's an enormous range of things. I, I think the way I would describe it is actually to kind of like pick up on Nick's language a little bit. He talked about how, you know, 80, 90% of, uh, time spent by data scientists, data engineers is, you know, quote unquote data cleaning. When you look at what's happening in that data cleaning process, people are taking raw data, and raw means somebody else generated it, and you don't, you know, often don't know exactly where it came from. So in the data cleaning process, you're trying to form a mental model of, okay, where did this data come from? And therefore, um, you know, what should I do with it in order to make it appropriate for my use case? And so these are decisions like, oh, maybe I'm going to drop outliers, or I'm only going to pay attention to certain columns, or I'm going to say that certain events are good and certain events are bad, or lots and lots of little decisions. And when you look at kind of the work that goes into data science and data engineering, a huge amount of it is just making those Very small granular decisions about how you interpret this, you know, lump of bits that came to you and what you believe that that means about the upstream world. And now any time when one of those decisions changes, either in your pipeline or because somebody upstream is doing something different, that could break things downs…
AI assessment note: “any time when one of those decisions changes... that could break things downstream”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yep. There is a concept of extensibility for the assertions. What does that mean?
A Uh, it means that every, every one of these expectations is its own little module, and over time we actually expect for a lot of these to be built. Um, one of the, one of the things that we're asked about most frequently is, hey, can we extend the library of great expectations to do geographic data, expect point to be within region, or time series data, expect trend to be increasing by x percent. Um, we've done a bunch of work to get the core of the library in good shape, And one of the things that we're really excited to do, uh, that now that we're through our Series A round is, uh, kind of resource better, uh, collaboration with the community to build out that library. Great.
AI assessment note: “it means that every, every one of these expectations is its own little module”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q level startups like Calm and Komodo Health, um, but now public companies like Vimeo or like more traditional like Fortune, 1000, uh, companies like Heineken, um, like presumably those are like pretty different types of data, um, Stacks or, or maybe not. I don't know. Um, does the product work with like any kind of environment or, uh, is it particularly appropriate for like a certain type of, of customers?
A So, I mean, I would love to say it works everywhere. Uh, the truth is there are some things that are, are fully supported and some things that are more experimental. Um, so for example, Dask, uh, we, we know that there are a number of companies that have deployed and use great expectations on Dask. There are a lot of Prefect users, for example, who are using both tools together, and we don't currently, as part of our testing infrastructure, run all of the checks against Dask, so it's possible that at some point, you know, we'd accidentally break something. Now, that's the thing that we plan to fix eventually, but so I just want to point to this gray area of places where the tool is being used, but we as maintainers of the open source project haven't fully shouldered the burden yet of maintaining it. Um, now that said, we've always believed that it was very important to have, ah, kind of a global view into what is going on in your data. And, I mean, as we've heard from Nick and also from Daveris, like, data engineering is like a multi-application, like, multi-complicated stack world. And so having tests that can travel with you across that is important. Uh, the primary backends that Great Expectations runs on are Python pandas. Which allows you to bring in a lot of notebooks, a lot of machine learning type workflows. Uh, we also do Spark data frames, um, and then SQL, uh, throug…
AI assessment note: “there are some things that are fully supported and some things that are more experimental”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q think you said, um, if you look at the evolution of the data space that, um, a lot of the, um, sort of early wins in the data space over the last few years have already been around managing scale, uh, and that we're now switching to a phase where, um, the primary challenges are higher in the stack around productivity testing integration. Is that, is that, is that fair?
A I think that's fair. Um, you can think of it sort of like Maslow's hierarchy of needs. Uh, and when the amount of data started to explode, there was a very critical need that you couldn't even process it on a technical level. Like you could not scale out compute. And that's why, you know, in the big data revolution, you called it big data. You started with You know, Hadoop and then spark and now the cloud data warehouse. So now if you have petabytes of data, you are able to process that efficiently at scale, but now the problems are kind of higher up on Maslow's hierarchy, so to speak. So now it's okay. We can actually process the data. Okay. What is it? Does it have high quality? Are the people that actually build the computations that process that data productive? Can you track that data throughout the ecosystem? And so on and so forth. So, you know, you know, an amazing engineering achievement happened in the early, you know, throughout the 2010, which was solving these massive pure technical scale problems. But now we're talking about organizational scale, uh, dealing with complexity and dealing with developer productivity and any number of other dimensions.
AI assessment note: “I think that's fair. Um, you can think of it sort of like Maslow's hierarchy”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q you know, clash, upcoming clash between Snowflake and Databricks as, you know, two gigantic companies in the space. So is your vision of the future that the data lake eventually, the lake house eventually becomes the paradigm, and then everything else over time gets absorbed? Or do you view a future that's more Hybrid, where you have data warehouses to do certain things, and like, hazards to do other things?
A I'll answer it in two ways, and I really do mean both of the ways. Um, you know, I'll start by saying, you know, it's kind of like people make it about zero sum, but if you answer it like this, do you think Google Cloud will eliminate Amazon Cloud and Microsoft Cloud, or do you think, you know, Amazon Cloud will eliminate the other clouds? Nobody thinks that, right? They're going to be around. They're all going to be successful. The data space is huge. There's going to be lots of vendors in it. I think Snowflake will be successful. I think they right now have a great data warehouse. You know, it might be the best data warehouse in the market. Maybe BigQuery would give them a run for their money. Um, but, uh, it's a great data warehouse. Um, it's certainly going to coexist, and it already coexists with Databricks in probably 70% of the accounts we're in. Uh, I think that's going to continue to be the case, and people are going to use data warehouses for BI. But if you ask me long term, the answer is yes to your question. Long term, I think the lake house paradigm will win. Now, you know, it might be that the other vendors like Snowflake completely embrace it and revamp what they have to become that, uh, or other players come along in that space. But in the long run, this is going to be the architecture that wins. Why? Because the data has so much gravity. All of it is sitting in…
AI assessment note: “Long term, I think the lake house paradigm will win.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Remarkable. And then you keep adding more stuff to it, like more features, more like product expansion, so data lake, search, online archive. And Realm, which was an acquisition, I believe. Do you want to talk about the, the, the product expansion thinking?
A Yeah. It really comes from a lot of customer feedback and our intuition about the market. So let, you know, you, you mentioned a bunch of products. So search, search is basically what we found. A lot of customers were dual homing data to both MongoDB and say elastic. And they said, this is crazy. We don't want to do that. And they said, we really want to, you know, embed all, you know, our OLTP functionality along with application search in one platform. It's much easier to maintain. We don't have to manage two systems, learn two platforms, blah, blah, blah. And so we've now embedded that. It just makes life so much easier for customers. And Elastic themselves are going more towards a security space than application search. Um, the, the second area around, uh, Realm, Realm.
AI assessment note: “It really comes from a lot of customer feedback and our intuition about the market.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q really interesting when I was, uh, when I was researching this and ahead of this talk, like I actually, I hadn't realized to which extent the, the MongoDB open source project has continued to just like explode in, in popularity. Um, what, what are the key drivers, uh, behind that? Like, what, why, why do you think, um, it's, it continues to be so incredibly popular and, in fact, accelerating?
A Well, I think the first thing is like, one, you've got to have strong product market fit. And what essentially MongoDB did was address two fundamental problems with the relational database. One, the relational database where you're trying to basically disaggregate data into rows and columns is a very counterintuitive way to model data. It's frankly, it's not aligned to how developers work. And it just, and as the data model grows, it becomes more and more brittle and it's It's harder to change. I mean, talk to anyone who wants to do a schema change and they'll roll their eyes because they know how painful that is. Well, we essentially solved that problem by coming out with our document architecture, which is much more, uh, um, conducive to the way developers think and the way developers code. So it just makes it so much easier to use MongoDB to build applications, to do it quickly and to do it more powerfully. The second big challenge with relational databases that are not designed for scale. I mean, Oracle was founded in 1977. We're talking about a company, a company that's now 44 years old, and they still haven't solved a scalability problem. You know, they've had like Oracle Rack, SiteCard, you know, and other ways to try and replicate data. They're super expensive because the relational database was designed to be a single node system. So as data volumes grow, their archite…
AI assessment note: “So that's why we've gotten so much popularity”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Amazing. Um, so if part of the evolution of the license was to sort of clarify the relationship with the, the cloud vendors, um, how do you work with cloud vendors these days? Like what, what can they do? What, like how do you incorporate or partner with them?
A So obviously with Atlas, it has to be deployed and delivered to the cloud, so we have to work with cloud providers. In fact, the cloud provides helps subsidize the R&D effort to deploy Atlas on the cloud platforms, because remember, They're the beneficiaries of these workloads moving to their platform. Not only do they get, you know, revenue from the underlying consumption of storage and compute and network services, but they also, they see customers for every dollar to spend on Atlas, they probably spend another five to 10 dollars on other ancillary services. So, so they're motivated to drive, you know, those workloads to the cloud. Now, Amazon and Azure in particular have competitive offerings, but we feel we're very, Well positioned when it comes to going head to head against those, but there's a, there's some spirit of competition. And that's, that's like we've been dealing with the big guys, Oracle from day one, Microsoft, and obviously now Amazon. Um, and that we recognize, you know, the market's large enough and we feel like we're well positioned. Um, and, and they've actually helped subsidize and drive a lot of marketing, right? They give us marketing dollars to drive customer demand. We partner with them on events. We partner with them on deals. They align the sales compensation to incentivize working to have their salespeople work with us. So there's a lot of, you kno…
AI assessment note: “We partner with them on events. We partner with them on deals.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q um, so you mentioned that you use open source as a freemium, um, kind of, um, strategy. Is that, is that for a company of the, you know, global scale of Mongo, is that, is that still the core motion? Is that like sort of bottoms up or is there, Also top down of like, uh, you know, AEs are going directly to CIOs. Like how do those two motions?
A Yeah. So, um, so our business starts to stop with developers. If developers don't use our product, we don't have a business. So we have to make sure we drive a lot of developer enthusiasm by using MongoDB. That being said, developers don't necessarily have the right to say yes. They have the right to say no. They can kill a deal by saying I'll never use MongoDB and you're done. But just because they like MongoDB doesn't mean that you're going to get a check in the mail from the customer, right? So you have to then Translate developer enthusiasm and interest into a real business, you know, value proposition for some decision-maker. It's not always a CIO. It could be the, you know, VP of infrastructure. It could be the line of business, depending. And in our business, it's a land expand. So we're not like a Workday or Salesforce. We're making an enterprise-wide decision day one for the whole enterprise to use your platform. You know, we coexist with Oracle, we coexist with Microsoft, we coexist with other players, right? And we want to get that first workload. And we find the relationship is kind of tends to be in three stages. First, you kind of land, land a workload or a few workloads. Then you find a JCC to expand into as they see successes. Okay, I should use MongoDB for the next application or re-platform an existing application. And then maybe over time, maybe even years, t…
AI assessment note: “Translate developer enthusiasm and interest into a real business, you know, value proposition”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q and then it progressed to representing 23% of revenues in, uh, fiscal year, 2019. 39% in twenty-twenty and now 46% in twenty-twenty-one. So, uh, it's gone from zero to being almost half of the company's revenue. Um, I just, I just think it's absolutely fascinating and super impressive. So, uh, maybe as a level said, what does Atlas do so folks can, uh, get their, wrap their minds up? Sure.
A What Atlas essentially allows people to do is consume MongoDB as a service. It's a managed offering delivered, um, from the cloud, and we run on all the three major cloud providers. And you pay as you go. So you can start the, you know, the cheapest paid tier is like nine bucks a month. So it's literally like, you know, like having two cups of coffee a month to, we have customers now spending, you know, seven more than seven figures with us on Atlas. So depending on how big your deployment is, how much resources you need. And it really, um, ranges the gamut in terms of, uh, a variety of customers. I would say there's two big reasons why customers love Atlas and why I'm one big reason why It really helps MongoDB. So the two big reasons why it helps, um, uh, customers is one is it just makes it that much easy to consume, um, um, MongoDB because MongoDB has had a very strong product market fit. And one of the other things customers care a lot about is getting rid of undifferentiated heavy lifting, right? And managing a distributed database can be quite challenging. And so when they, when they basically say, you know, why don't you, the experts manage it for us, it just alleviates them. And a customer will always tell us, you know, my competitive advantage is not by hiring more DBAs and more support staff. My competitive advantage is by building great product, adding new features, …
AI assessment note: “What Atlas essentially allows people to do is consume MongoDB as a service.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Okay. Thank you. And data lakes is part of this as well, right?
A Correct. Correct. So as, so we're leaving people to use the power of our query language to query data, not just from MongoDB, but from other sources. And it's not necessarily trying to, we're not trying to be a data warehouse. We're not trying to compete with With, um, Snowflake. It's more to help developers, because we believe that applications of the future, even today, are more and more instrumenting the business. So applications are not embedding analytics, especially real-time analytics into apps. Think about, like, a leaderboard for a gaming application. Think about financial app. Think about supply chain apps where you have to have real-time, you know, information. So being able to embed analytics is going to be even important, and our focus is really helping developers use data to build amazing applications.
AI assessment note: “Correct. Correct. So as, so we're leaving people to use the power”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q As a quick insight, it's sort of amazing actually how important Yahoo has been to this whole, like sort of big data, data ecosystem, right? Like so many fantastic people, uh, have come out of, uh, of Yahoo.
A That's right. And so Yahoo created Hadoop for those that don't know. And so Hadoop was born in Yahoo, and then it was sort of spun out as a separate company, but within, uh, that data organization, Things like Kafka eventually came out of LinkedIn that was based on work that we had been doing at Yahoo and many other things had come. We were doing machine learning and behavioral targeting back, uh, 15 years ago, more than 15 years ago in advertising. And so we were the first in a lot of things and also did a lot of data privacy at that stage, which is only now, really in the last couple of years, coming to the forefront of being important. So The great, one of the great things about Yahoo is that there's so many people spread around the Bay Area, probably New York as well, that have been in Yahoo. So there's a really good network of folks, and, uh, it was a real pleasure to work there. I was there for six years.
AI assessment note: “That's right. And so Yahoo created Hadoop for those that don't know.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. And then so ultimately all of this, um, is for, Analytics. What are some of the use cases?
A Actually, yeah, yeah, actually we have many, many use cases. Uh, so analytics is a small part. We, we have like just about everyone in Pinterest is using it for analytics every day, but we also have a massive, uh, number of experiments, right? We have got our own experimentation platform where we're doing AB experiments all the time to try and improve Pinterest and try and improve our ads, our ad, uh, relevance. And so we have about a thousand experiments running, uh, in parallel at any point of time. And so we're, we're constantly iterating. Uh, so that's another use case. And then for, we use it a lot for machine learning. So we have about 80 different use cases of machine learning. And that's, you know, mostly from this data that comes into, that comes into S three and AWS.
AI assessment note: “we have got our own experimentation platform where we're doing AB experiments all the time”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So, uh, since we're talking about machine learning, what are some of, some of the tools that are being used by the machine learning folks and frameworks?
A Yeah. So in machine learning, so like two or three years ago, it was, uh, a Pandora's box. Everyone was trying out everything. And so we had, uh, with all these different use cases, people were, were, uh, you know, it started to become more of a maintenance burden. And so what we wanted to do, what we've done is created a machine learning platform that kind of glues together the best of breed, uh, open source products out there. So we use TensorFlow from Google. We use PyTorch, scikit-learn. We use MLflow, which is a, from Databricks, is a registry of machine learning models. And, uh, we also, uh, have our own format for storing features within, within, uh, Pinterest, both for doing batch learning of models and then doing the serving models. So we, what we can do is within a few hours, you could do, uh, create a, a model, train a model, and then deploy it to our production service with thousands and thousands of servers, uh, by using this MLflow repository.
AI assessment note: “So we use TensorFlow from Google. We use PyTorch, scikit-learn. We use MLflow”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q will ask the question live. So that's the best way of doing it. Um, What one more from, um, me, uh, so I'd love to dive a little bit into the organizational aspect, um, of this. How, um, is the data team organized, machine learning team organized, uh, are they separate that they work together? Is that centralized? Is that throughout the organization? How does that, how is it organized?
A So data engineering at Pinterest comprises the, the serving part, all the online systems. So that includes the online databases like MySQL and key value stores like RocksDB and HBase, and also includes Druid platform. So we have our online systems. We also have all the batch systems and analytics and experimentation platforms. So we discussed that previously. And then the machine learning. Piece as well as a, as a platform. And so we provide the, the glue and the underlying infrastructure, uh, with TensorFlow and PyTorch and all the, all the other components available to our internal customers. And so with machine learning, machine learning is actually distributed everywhere. We have, uh, Many, many machine learning engineers and many, many use cases. And those mission machine learning engineers are embedded within each organization. So there's quite a few within shopping, quite a few within our advertising business and also different parts of our product and trust and safety. And then we have a separate product analytics and data science team as well that uses all the capabilities of, of data engineering.
AI assessment note: “machine learning engineers are embedded within each organization”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q news today. The open sourcing of, of, of query book. Um, so I'd love to, uh, take a step back from this specific tool. And, um, if you could help us, uh, you know, give us a, a, a broad picture of the data engineering analytics stack at, um, at Pinterest, what, what do, what do you use? You know, the various, the various tools and the various systems. Sure.
A Sure. So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. We're moving more and more to spark now, uh, because it's faster, uh, but Hadoop still scales really, really well. Uh, we have a query platform, which is Presto and Spark SQL. We used to use Hive, but we're migrating of Hive to Smart SQL, so we just have the, the two main engines. And then we have a workflow system on top of that that's built with Airflow, which is also open source. And the way that we get all this data is using Kafka. So Kafka, uh, has, uh, clients in all of our serving systems and gathers the data from these clients and then pipes it to the backend. Now we're, all of contrast is, is running on EWS. So we put all of this data into S three and it's a massive idea. We have more than 400 petabytes of data. Uh, And so we, we spend quite a bit of time, uh, putting it into different buckets, and we can show that the buckets are partitioned correctly, so that the, these engines will perform at scale. We also, uh, it's important to put it in the right format as well, and so we want to try as much as possible to put it into a parquet column or format, uh, for these query engines to be even faster.
AI assessment note: “for the analytics itself, uh, we focus on, uh, using Hadoop and spark.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And you mentioned examples, um, with videos, which, uh, qualifies as well as document and sounds that qualifies as unstructured data. Uh, but, but, uh, deployment can be used on, on tabular data as well, right? What, what are some examples and some use cases for tabular data?
A Yeah, I mean, that's my favorite topic. If you guys look at Abacus, you will see that our first focus was all on tabular data. And the reason we focused on tabular data and why tabular data is actually pretty effective with deep learning is, uh, uh, one, uh, because every company has multiple different, uh, uh, examples or machine learning, uh, models that they want to develop using tabular data. At its most basic level, we have something called predictive modeling. Predictive modeling is very similar to Finding a variable, uh, which is what we, what we call a dependent variable based on other independent variables. Let me give you a concrete use case. If you go and look for a home on Zillow or Redfin, you have something called the house property price, right? Estimate, the Z estimate. I'm sure everyone's seen this. The Z estimate comes because they have a machine learning model, which is a predictive model, which is using all kinds of properties, like say bedrooms, bathrooms, square footage, neighborhood, location, and whatnot. And trying to determine the price of the house based on recent homes, which have sold, uh, you know, with those attributes. So it's learning constantly based on recent sales, and then trying to predict based on the attributes of your home, what the price would be, right? That's a predictive model. And these predictive models, there are very many methods…
AI assessment note: “If you go and look for a home on Zillow or Redfin, you have something”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. Who's an ideal customer for the platform? Uh, is that a, like a fortune 1000 company that, uh, doesn't have any general resources that it started? Like who's the ideal customer?
A So we have a bunch of, I would say, startups and medium-sized companies. So we have, I would say, two classes of ideal customers. One is the startup of the medium-sized company who basically wants to go, go, go, right? They have like seven different ML use cases. They may have anywhere between two to five data scientists, sometimes zero. So let's say zero to five data scientists, and they want to get things in production. What we've seen all the time is, I'll give you a classic example of this. This is a medium-sized company, not really a startup. It's a company called USIC LLC. I'm picking this one because they're the eight one one company. When you call them, you call them every time you want to dig somewhere because they have to come and mark exactly where a particular, you know, utility is so that you don't dig, dig it and like, you know, harm the utility. They have, they're a pretty large company. I think they're about 500 plus employees. They have a ton of data. They want to be able to like build models which are like, Probability of a particular site getting damaged. Time to, like, go and, you know, make that inspection. So they have multiple different machine learning models, and they have a lot of scale, and, you know, they would rather use Abacus than not. So that would be our, like, you know, I would say a sweet spot customer in a small to mid-sized company. Then we …
AI assessment note: “So we have, I would say, two classes of ideal customers.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q learning model. Your data is your leverage and your best most against competition. Don't wait to hire a data science team. Use a low-code ML platform and get started ASAP. Um, so you've, you've seen people do this, um, successfully or what, what, how does one, uh, get started? I guess what, what's, what are some examples of like low-code ML platforms and how do people go about doing that?
A Well, it's a, it's a self-promotional tweet. So an example of a low-code ML platform would be Abacus. Um, I mean, I'm sure there are others as well. Uh, but, um, you know, I'll give you an example. One of our customers, a startup called Daily Look, uh, you know, they're kind of like Stitch Fix, uh, except their, their business model is slightly different in the sense that, uh, I think you, you get to choose certain things that you want in your box. But presumably every month you get a box of clothing that you might like. You keep some, you, you, you, you, you know, you return some. So, you know, even when they were pretty small, they had a few thousand customers, and they didn't have a machine learning team, they started using us. And they've been using us for a year. They're one of our oldest customers. And, you know, now they're up to model three. And, you know, it's worked really well for them because they kind of knew what the problem was. They wanted to be able to optimize the amount of, Uh, uh, you know, the number of things that you kept every month, clearly, right? So you have to pick the right, uh, you know, style items to send them on a monthly basis. Clearly that's a machine learning problem. You don't have to wait until you have millions of data points. Even hundreds of thousands is too much. If you have 10,000 data points, that's my rule of thumb. If you have 10,00…
AI assessment note: “An example of a low-code ML platform would be Abacus.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great question that is going to enable you to like dish out on the competitors. What are some of your competitive advantages or even drawbacks compared to H two other AI and data robot?
A Yeah. So if somebody actually knows our space, but HTO and data robot are really, we, I mean, maybe they consider themselves our competitors. We like to think of them as non-competitors. The reason for that is their real focus, as you probably may know, is on the auto ML or the automated machine learning part of it. Our focus really is to do that end to end. We've spent a ton of time in the data wrangling data management, you know, and doing that across multiple sources. So literally you can connect, you know, your data from Snowflake, from Redshift, and from your data lake and Salesforce together and bring that all together. So I think our biggest competitive advantage is that data wrangling and the real time feature store, which is part of our service.
AI assessment note: “our biggest competitive advantage is that data wrangling and the real time feature store”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Jumping into this, do you want to talk about Kafka and, like, explain for, like, the non-technical part of this group? What's Kafka? What's the big deal? And what, I guess, what it does?
A Right. So Kafka is an open source technology, which is a message broker. So if you have, if you have Various systems that are producing data. You have various other systems that are consuming data. You can put Kafka in the middle in, and you can have the producers of data publish various feeds, and you can have consumers of data in a fairly, In a fairly decentralized fashion, choose which feeds of data or which topics as they call them in Kafka to subscribe to. And what this means is that you can sort of create this, um, real time clearing house infrastructure in your data In, in, in your company, um, to allow various teams that may not even be coordinating or, or, or with otherwise in the absence of a system like Kafka have to, you know, essentially coordinate and talk to each other and figure out a way to, to, to move that data and closely integrate their systems. They don't actually have to, um, you know, do that coordination and they can just sort of subscribe to those data feeds that are available and, and potentially come up with use cases. Um, that make use of that data in real time. You could imagine how this would happen in batch, right? So, so everybody at the end of the day would, would put all their data in, in a data lake, and then tomorrow anybody can pick up that data and do something useful with it. Kafka basically moves this, um, into real time, um, and allows …
AI assessment note: “Kafka is an open source technology, which is a message broker.”