The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Ali Ghodsi argument clarity score 4.3/5 from 12 exchanges on raw tape · average scores: directness 4.6 · coherence 4.7 · precision 3.9 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
12on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q you know, clash, upcoming clash between Snowflake and Databricks as, you know, two gigantic companies in the space. So is your vision of the future that the data lake eventually, the lake house eventually becomes the paradigm, and then everything else over time gets absorbed? Or do you view a future that's more Hybrid, where you have data warehouses to do certain things, and like, hazards to do other things?

A I'll answer it in two ways, and I really do mean both of the ways. Um, you know, I'll start by saying, you know, it's kind of like people make it about zero sum, but if you answer it like this, do you think Google Cloud will eliminate Amazon Cloud and Microsoft Cloud, or do you think, you know, Amazon Cloud will eliminate the other clouds? Nobody thinks that, right? They're going to be around. They're all going to be successful. The data space is huge. There's going to be lots of vendors in it. I think Snowflake will be successful. I think they right now have a great data warehouse. You know, it might be the best data warehouse in the market. Maybe BigQuery would give them a run for their money. Um, but, uh, it's a great data warehouse. Um, it's certainly going to coexist, and it already coexists with Databricks in probably 70% of the accounts we're in. Uh, I think that's going to continue to be the case, and people are going to use data warehouses for BI. But if you ask me long term, the answer is yes to your question. Long term, I think the lake house paradigm will win. Now, you know, it might be that the other vendors like Snowflake completely embrace it and revamp what they have to become that, uh, or other players come along in that space. But in the long run, this is going to be the architecture that wins. Why? Because the data has so much gravity. All of it is sitting in…

AI assessment note: “Long term, I think the lake house paradigm will win.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q For the founders and entrepreneurs in the, in the audience, uh, one of the, Peculiarities of the founding story of, of Databricks is that, um, you guys have seven or eight co-founders, which is very unusual in, in retrospect to what were the pros and the cons of having a large group like this?

A Yeah, I think there are pros and cons for sure. I think if you know how to actually get a tight knit group of seven people to really trust each other and work well together, amazing things can happen. I think a lot of the success of Databricks was getting all these seven people to really trust each other and do great innovation. Very few companies have the pleasure of having, you know, that kind of critical mass of thought leaders together. The downside can be oftentimes Founders. This didn't happen at Databricks, but you see it all the time in other companies. The early founders, even if there's two of them, they fight and then they split off early on within a year or two. Um, that's the problem. So if, you know, it might be too many cooks in the kitchen could be, could be a problem. We found a way where we really kind of know each other's strengths and weaknesses, and it's made, uh, this journey an absolute pleasure for me. You know, they always say the CEO job is the loneliest job on the planet. I never felt that way. It's, uh, you know, I had, Lots of sort of early co-founders with me that were always there. So, uh, so for us, it's been an absolute strength. We wouldn't be where we are if we didn't have those folks.

AI assessment note: “I think there are pros and cons for sure.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And to, uh, add the one more layer of complexity, this, the whole open source to commercial dimension, right? MLflow, Delta Lake, Koalas, which we haven't mentioned yet. Um, Uh, where does that fall in the innovation camp or is that a sub layer of the commercial camp?

A No, it's, these are all innovation camp. So they're all in the innovation camp. Of course, some of these projects, when they get older, like spark, they move into the maintenance side. Um, and we typically also move the people around. So it's the same people that do the innovations over and over. We try to grow more of those innovators, but we try to move the sort of people that are really have a knack for cracking the zero to one into the next problem, and then hand over the existing projects to other people who want a chance in, you know, in their career to run, let's say spark, which is a huge successful project, right? It's a big, Career step up for someone to get, step up to get that responsibility, but we move the person that created it to something else to create the next thing, and we also find who are the ones that are good at the zero to one things, and we actually experiment. We give people in, in, in R&D, you know, a chance to go experiment with the zero to one things, and they don't always succeed, and it takes a couple tries until they become really good at it, so you have to think deliberately about this kind of high failure strategy.

AI assessment note: “No, it's, these are all innovation camp. So they're all in the innovation camp.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q at a personal level. As you've grown from CEO of a, by definition, small startup to now a mega startup and, you know, soon, soon enough, uh, a large public company. Uh, how do you, how do you scale yourself? How do you, um, have you learned along the way and how have you switched from the job of being like the visionary storyteller to, uh, running a global organization?

A Yeah, I mean, it really boils down to finding right leaders that you can trust and building trust with them. So that's as simple as that. Can you find the right leaders that you can trust? I could spend all my days on shows like this and the company will continue to run itself. Why? I have great sales team that's well-functioning. I don't have to be directly involved in it. I have great marketing. I have great engineering. So why do I have those great departments? I have great leaders in those departments. And I trust them and we built this trust over many years. So, so it really boils down to, that's really the, I know it sounds simple or silly, but that's the problem you have to find out. And I think a lot of early stage, and I certainly had this problem as well. In the early stages, you have this situation where you're like, these people that are running these departments don't know what they're doing. I have to do it. It's about me, me, me, me, me. And then you go in and you, you know, you have your fingers in the pie all over the place. Um, that doesn't scale. Because, you know, as your organization gets to 152 102 hundred people, Dunbar's number, now you can't anymore remember even what's going on. So you feel kind of completely inundated and behind all the time and frustrated, and then when you hit like a thousand people, it's a whole nother deal with, you know, Japan of…

AI assessment note: “it really boils down to finding right leaders that you can trust and building trust with them”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q customer of Databricks and says there was a huge reason why Databricks was a value add for our machine learning engineers to become more self-service scheduled jobs. Uh, and for hosted data engineering pipelines for data warehouse incremental loads. Um, so what's your, what's your perspective on, um, on managed cloud containers, um, is that part of your longterm strategy, uh, to support this, this alt cloud in the future?

A Yeah. Um, so I think the whole, um, if the question is around Kubernetes and containers and Docker and things like that, uh, I think it's great. It's again, like USB. It's another standardization layer that makes it easier to move things across these different things. Um, so we support it. We think it's great. We're going to offer it on wherever you want to go. However, our experience is that you need much more today. I mean, we're huge fans of Kubernetes, Docker. We standardize everything under the hood on Kubernetes. It enables us to move between the clouds, but in the data space, you need more. You need a catalog. You need data discovery tools. You need ways in which you can search for your different data assets. You need ways to do security on your data assets. You need ways to dashboard it. You need ways to query it. So, um, you need these as well. So, Uh, so in some sense we're trying to build Databricks such that it becomes that open standardized layer that you can move between the different clouds, but absolutely it will also have plugins that you can bring your own containers and run your own Kubernetes kind of, um, apps or operators on it as well. Uh, so absolutely. I mean, this is what I mean with the ecosystem of open infrastructure that's actually being built up. Uh, you know, it's, that's what's going on in the last decade or two, and we're excited about them. We'…

AI assessment note: “absolutely it will also have plugins that you can bring your own containers”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q late stage, um, startup people can still call you a startup. Uh, what does open source fit in, in the go to market motion? Um, and, you know, coming back to the earlier parts where like bottoms up versus top down, um, like who does, what do you have, uh, like a BDR group still versus the AEs? How do they all work together without stepping over each other's feet?

A Yeah. Databricks is a hybrid model, so there's a top-down and a bottom-up at the same time combined. We started, as I said, with bottom-up, but we've kept it. So yes, we have BDRs, SDRs. There is, they create opportunities that then they hand over, right? It starts, it's a funnel that starts with marketing, and the funnel bleeds in from marketing into the SDRs. The SDRs then have, they get some of the leads from marketing, some of it just directly outbound from the SDRs, then it goes to the sort of sales team, right? There is also a very interesting bottom-up, uh, sort of completely free, uh, sort of freemium, free tier funnel. So Databricks Community Edition is a completely free, use it all you want, never pay us funnel, where you can use all of Databricks, but you only get like a slice of a small machine. So you get kind of a taste of the real big thing, but you could use it forever, and that then generates leads That also feeds into the SDRs. So that also is a pipeline. That's really important. Half of our leads comes from that. So that's why open source is an important, uh, engine for us. It's half of the leads that come to sales comes from that. And if we were just doing spark, like we were on this show in, that would have been probably 10, five percent because, you know, over time, these technologies become mature and, you know, the excitement around them, uh, wanes. Uh, …

AI assessment note: “half of our leads comes from that. So that's why open source is an important”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How did you, how did you find them? Do you have a bias towards promoting people internally or do you think it works better if you bring in sort of snipers from outside who have done that stage of the company? What's your, how do you approach it?

A It's so hard to find great leaders that work with your culture and that you can build amazing trust with, that I think you should, you shouldn't exclude any options. If you can promote people from within, great. But if you just try to promote from within, you probably are not getting the experience that exists in the market. The experience can be super, super valuable. People have seen the movie. You need to also bet on them as well. Uh, what are the things we look for? We look for people who have seen build. So the joke I say is, do you have a driver's license? And people will say, I have it. Are you good at driving your car? Yes, I'm very good at it. Why are you asking? Can you build a car with your bare hands? And people say, okay, I get it. So can they build it? Not just drive it and maintain it. Uh, they have to have built the phase we're in now. I'm not saying they have to have built a twenty eight billion dollar company from zero. That's not what we're looking for, but they have to, in our case, taking a company to, you know, a few billion dollars of revenue, uh, or seeing that kind of phase in engineering or marketing or wherever it is. Um, so that's what we're looking for. Have they, we're really looking, did they build it? And then we look at, did they have, First principles thinking when they build it. Were they just joining in for the ride as those companies were go…

AI assessment note: “I think you shouldn't exclude any options. If you can promote people from within, great.”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q And what was the big technical breakthrough, uh, to enable, to provide that layer of structure on top of the, like, house with this Delta Lake? I think there was iceberg that came out around the, the same time. Uh, how does that work?

A Yeah, actually, the four technological breakthroughs that kind of happened at the same time, 2016, 17, at the same time, the one we contributed was Delta Lake, there was Hootie, there was uh, Hive acid, and there was icebergs. Four technologies at the same time kind of started, and with a lot of breakthroughs in science, that's kind of what happens, like, you know, several groups at the same time will, like, let the DNA crack the code of it, right, in the US and in the UK. So the idea, the problem was this. You had all this data in the data lakes that people had collected. It was super valuable, but it was very hard to do structured queries on it, basically SQL, basically BI. So for that, you need a separate data warehouse. Why was it so hard? Because the data lakes were built for big data, large data sets. They weren't built for really fast queries, so they were just simply too slow, and they didn't have any way to structure the data and give it sort of Tabular form. That was the problem. Um, so how do you take something like a big blob storage for data and turn it into a data warehouse? Uh, so that, that's the, that's the secret sauce was these projects. We basically figured out ways to work around the inefficiencies of these data lakes and enable you to get the same value you would get out of data warehouse directly there on your data lake. So that, that was those projects. …

AI assessment note: “figured out ways to work around the inefficiencies of these data lakes”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Are there any trade-offs to this approach?

A Not really. Um, you know, you can have your cake and eat it too. I know it sounds crazy, but you can, you can, it's reducing a lot of the techniques that were invented, invented in the eighties and nineties by data warehousing vendors, adapting them to making them work on the data lake. Um, you could ask why did this not happen 10 years ago or 15 years ago? Um, it's the ecosystem of open standards didn't exist. It's slowly emerged over time. So it started with the data lakes. Then there was a big actual technological precursor breakthrough to this that we're talking about here, which was standardized formats for the data. You know, they're called Parquet and ORC, but these are data formats that the industry kind of standardized all their data sets on. Uh, those kind of standardization steps were needed, uh, to get this breakthrough of the lake house. It's kind of like the USB. Once you had it, you could connect any two devices with each other. Um, that was what was needed for the industry. So slowly what's happening is that the open source realm An ecosystem is emerging where you can do all of your analytics in this lake house paradigm, and you don't, eventually it will be the case that you will not need all these other, you know, proprietary old systems that people have had since the eighties, the data warehouses and other systems like that.

AI assessment note: “Not really. Um, you know, you can have your cake and eat it too.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q If you were going to start, uh, another enterprise software company today, would you go open source first?

A Yeah, I think it's a superior, I think it's, if you think of it from an evolutionary standpoint, it's evolutionarily superior to the previous business models. Uh, why do I say that? Uh, because any proprietary software company out there is ripe for disruption by an open source, uh, uh, competitor. So anything that's proprietary can immediately be disrupted. I mean, just like Windows got disrupted by Linux, the same, I mean, that's as advanced as it gets, right? That's really complicated technology. Operating systems, right? Low level operating systems for different types of hardware. You wouldn't think, uh, some guy out of a university would invent that, and that would become the standard in the industry. Any proprietary software, uh, is ripe for disruption like that. The question is, can you make money on it? And that has been really hard up until, you know, Red Hat and all these companies that were doing supporting services, until Amazon Web Services cracked the code on the business model. The business model is, We run the software for you, you rent it from us. That's a superior business model, uh, because, um, you actually then can have a lot of IP that's very hard to replicate. Uh, so I think that's the, so I would next company I start would be that. And if you're going to ask me, I don't know what your next question is. If it's going to be, what would you start it in? Whic…

AI assessment note: “so I would next company I start would be that.”

Partly raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q Can you double click on SQL analytics, which is the most recent major release and major product addition, and including how you work with the existing ecosystem of BI solutions?

A Yeah, I mean, that's really our kind of business analytics, business analyst, warehousing offering directly on the lake. So it has all the classic pieces of a data warehouse engine. So we re-implemented. So in the past, when someone wanted to do SQL or warehousing on Databricks, we would offer them Spark. Spark has SQL, um, but Spark was written in Java, uh, you know, couldn't have the performance of the best in class, uh, data warehouses. So I think two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine. So basically a modern data warehousing engine written in C++ for modern hardware. It's called SIMD instructions, like modern, you know, modern hardware can do lots of instructions in parallel on the same data, right? So it's perfect. So it's really in the lake house building warehousing capabilities straight into it. Um, So that's what we announced last year. We're excited about it. You know, it's, uh, we're seeing huge performance improvements. We're actually gonna, uh, you know, reveal a lot of the performance improvements next week or in two weeks at our data and AI summit. So, uh, so that's really exciting.

AI assessment note: “that's really our kind of business analytics, business analyst, warehousing offering directly on the lake.”

Partly raw tape D 3 · C 4 · P 2 · Cm 3 3.05

Q know, chapter of the early years, um, how did you go from this academic, uh, very popular open source project, uh, which was Spark to A company, you know, from zero to call it ten million in LRR, was there like any, any defining moments, any perhaps any hacks, any growth levers that you guys use to, to, to go from zero to one, zero to 10 in that case?

A Yeah, I think the zero to 10 journey is very special. It's very different from the rest of the journey. I mean, so we went through three phases and I can explain each of the three, but the first phase is really the product market fit phase, figuring out, so you have a product, does it get, do, can you find fit between that product and some audience that really loves that product? Can you make that happen? And there were challenges around that. I'm happy to explain what they were. And then once you found them, you kind of figure out what's the channel That can connect that product with that market. So you have product market fit, but what's the channel to sell to them? There are different ways to actually set up the channel and we actually got it wrong initially. So it took us a couple of years to actually figure out what that right kind of, uh, sort of tweaks to it were. So those were definitely very special years where it's a lot of experimentation to figure out what the right model for Databricks is.

AI assessment note: “we went through three phases and I can explain each of the three”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.