The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Matei Zaharia argument clarity score 4.5/5 from 10 exchanges on raw tape · average scores: directness 4.8 · coherence 4.9 · precision 4.4 · compression 4.1 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
10exchanges match
10on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q What was the limitation of, like, why wasn't what was in place MapReduce, I guess you were describing, enough?

A So MapReduce was, was actually great for running sort of large batch jobs that kind of scan through the whole data and, and summarize all of it and give you an answer. But it was designed mainly, you know, for jobs that take tens of minutes to hours. MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web. But what Facebook wanted to do was different. They had a lot of questions that they wanted to ask almost interactively, and there's a person sitting there who launches a, you know, a question at the cluster and needs to get an answer back, and it wasn't very well suited for that.

AI assessment note: “it was designed mainly, you know, for jobs that take tens of minutes to hours.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q What was the limitation of, like, why wasn't what was in place MapReduce, I guess you were describing, enough?

A So MapReduce was, was actually great for running sort of large batch jobs that kind of scan through the whole data and, and summarize all of it and give you an answer. But it was designed mainly, you know, for jobs that take tens of minutes to hours. MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web. But what Facebook wanted to do was different. They had a lot of questions that they wanted to ask almost interactively, and there's a person sitting there who launches a, you know, a question at the cluster and needs to get an answer back, and it wasn't very well suited for that.

AI assessment note: “it was designed mainly, you know, for jobs that take tens of minutes to hours.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So is there just like a lot of documentation on the project then? Or, I mean, what really makes them easier to be able to contribute then? Like what concretely happens?

A Yeah. So there's several things. So first, you know, first, even before anyone contributes, they have to be able to use it. So there's been a lot of focus on making Spark very easy to download and use and having as much documentation and examples out of the box Uh, as possible. And we're still doing a lot, uh, to, to expand this actually. The second thing you, you need is, uh, you know, once people, uh, are actually trying to, to send in patches or to, to, to try to understand something about the project, uh, you do have to talk with them to review the patches and so on and, uh, you know, and help them actually get them in. And the third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this This kind of infrastructure, much the same as you do in any other engineering organization, uh, you can then, uh, end up moving a lot faster. So these are the things we, we spend time on.

AI assessment note: “focus on making Spark very easy to download and use and having as much documentation”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You know, let's take a step back for a moment and talk again about your experience at Facebook, though. Why was the problem challenging there? I mean, big data has been around forever. So what about that problem was interesting and different that made you want something better than what you already had? Like, what was it about that data, I guess?

A Uh, so Facebook, uh, like, like other companies starting to, you know, to, to use, uh, business data was able to, to collect a lot of very, uh, valuable data about how users are interacting with it. And Facebook was also growing very quickly. So they were adding, you know, many tens of millions of users every few months, and, um, you know, they definitely couldn't just talk to every user or even send someone to every country where Facebook was used and figure out how people are using the site. So they needed to use this data to, uh, improve this user experience. So there were two things that made it especially challenging. One was the scale of the data, which was, you know, much higher than you could do with traditional tools, and the second one was how many Uh, different people within Facebook wanted to interact with it. It wasn't just one person or one team doing it. It was many people often with not, uh, you know, not, not that many technical skills, so it needed to be very easy to work with this data.

AI assessment note: “there were two things that made it especially challenging. One was the scale”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So let's talk about another transition here, which is the transition from an open source project to becoming a commercial one. And, and part of what you guys are doing, obviously, is, is involved in that. Can you talk a little bit more about the transition and what that takes to move from open source to commercial application that people actually use and have expectations of?

A Yeah, definitely. So, uh, you know, so as we saw, uh, Spark do very well in, in the, in the open source domain, we, we wanted to start a company around it to really harden it and to bring it to a much wider class of, uh, commercial users. And, um, you know, we really wanted to find a model that lets it continue to be fully open source and continue to be a successful project that way for, for everyone participating in it. And, um, traditionally it's, it's always been a tension in, in companies that, uh, you know, that try to commercialize open source projects because, uh, you know, they build all this great stuff and then they kind of give it away for free. And, uh, you know, they, you know, it's, it's this tension between, oh, are there some things we just shouldn't put into it? Or, you know, how, how, how else can we actually have, uh, you know, power as a successful business around it? So the way we're doing this at, at Databricks is, uh, is actually quite different than I think it's a, it's an, A very nice, a very powerful model for doing this, which is that we're offering Spark in a cloud service.

AI assessment note: “we're offering Spark in a cloud service”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You know, let's take a step back for a moment and talk again about your experience at Facebook, though. Why was the problem challenging there? I mean, big data has been around forever. So what about that problem was interesting and different that made you want something better than what you already had? Like, what was it about that data, I guess?

A Uh, so Facebook, uh, like, like other companies starting to, you know, to, to use, uh, business data was able to, to collect a lot of very, uh, valuable data about how users are interacting with it. And Facebook was also growing very quickly. So they were adding, you know, many tens of millions of users every few months, and, um, you know, they definitely couldn't just talk to every user or even send someone to every country where Facebook was used and figure out how people are using the site. So they needed to use this data to, uh, improve this user experience. So there were two things that made it especially challenging. One was the scale of the data, which was, you know, much higher than you could do with traditional tools, and the second one was how many Uh, different people within Facebook wanted to interact with it. It wasn't just one person or one team doing it. It was many people often with not, uh, you know, not, not that many technical skills, so it needed to be very easy to work with this data.

AI assessment note: “One was the scale of the data... second one was how many different people”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q the ecosystem you're describing, not, it's just not random chance that it ended up there. I guess what I'm really interested in is that you are the creator of one of the most popular and fastest growing open source projects ever. What are some of the ingredients of a successful open source project? Like what does it take to kind of get here for other people working in open source?

A Yeah. So I should say, you know, from the beginning that, you know, we didn't, we certainly didn't, uh, imagine that Spark would be this widely used, uh, when we started. And it's only, you know, it's kind of been a feedback loop as we saw people being excited in it. We, uh, also, uh, you know, decided to spend a lot of time to make it better and, uh, and to foster the community around it. But I think there are several, uh, things that helped. So first of all, uh, Spark was actually tackling a problem, a real problem that people had and that more and more people were beginning to To have in the future, which was the problem of working with really large scale data sets.

AI assessment note: “Spark was actually tackling a problem, a real problem that people had”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So is there just like a lot of documentation on the project then? Or, I mean, what really makes them easier to be able to contribute then? Like what concretely happens?

A Yeah. So there's several things. So first, you know, first, even before anyone contributes, they have to be able to use it. So there's been a lot of focus on making Spark very easy to download and use and having as much documentation and examples out of the box Uh, as possible. And we're still doing a lot, uh, to, to expand this actually. The second thing you, you need is, uh, you know, once people, uh, are actually trying to, to send in patches or to, to, to try to understand something about the project, uh, you do have to talk with them to review the patches and so on and, uh, you know, and help them actually get them in. And the third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this This kind of infrastructure, much the same as you do in any other engineering organization, uh, you can then, uh, end up moving a lot faster. So these are the things we, we spend time on.

AI assessment note: “having as much documentation and examples out of the box Uh, as possible.”

Answered raw tape D 4 · C 5 · P 5 · Cm 4 4.55

Q So what were some of the reasons for inventing Spark in the first place then? I mean, besides the problems and the limitations of that, were you just trying to solve the problems of MapReduce, or were you actually trying to do something different?

A Um, yeah, that's a good question. So we, um, started building Spark after several years of working on MapReduce and working with companies that were very early on using MapReduce. I was a PhD student at UC Berkeley, and we actually started working with Hadoop users back in 2007. And I did, for example, an internship at Facebook when Facebook was only about 300 people and they were just starting to set up Hadoop. And in all these companies, I saw that there was a lot of potential to putting together large data sets and processing them on a cluster, but they were all hitting the same kind of limitations and they all wanted to do more with it. So basically our lab created Spark to address these limitations.

AI assessment note: “our lab created Spark to address these limitations.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q the ecosystem you're describing, not, it's just not random chance that it ended up there. I guess what I'm really interested in is that you are the creator of one of the most popular and fastest growing open source projects ever. What are some of the ingredients of a successful open source project? Like what does it take to kind of get here for other people working in open source?

A Yeah. So I should say, you know, from the beginning that, you know, we didn't, we certainly didn't, uh, imagine that Spark would be this widely used, uh, when we started. And it's only, you know, it's kind of been a feedback loop as we saw people being excited in it. We, uh, also, uh, you know, decided to spend a lot of time to make it better and, uh, and to foster the community around it. But I think there are several, uh, things that helped. So first of all, uh, Spark was actually tackling a problem, a real problem that people had and that more and more people were beginning to To have in the future, which was the problem of working with really large scale data sets.

AI assessment note: “Spark was actually tackling a problem, a real problem that people had”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.