Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q from a product standpoint, so, so that you guys, ah, actually you were the first one in the market to, ah, partner with Databricks on, on, ah, on Spark and all the things. What, what does Hadoop look like in, you know, three, four years? Um, because we've evolved from batch to real time. What are, what are the product needs and what technology address, will address those market needs?
A The system that we commercialized in nine, ah, in 2008, Looked very little like the original software that Doug Cutting and Mike Caffarella created and that Yahoo developed in 2005, 2006. It had already evolved a long way. These days when people talk about Hadoop, what they mean is HDFS and MapReduce, yeah, yarn for resource management. You need some ingest tools, so you need Scoop and Flume, and you might even be looking at Kafka right now. Um, everybody is super hot on Spark. The technology is way easier to program. You know, it's got some rough spots. It's not well integrated with security framework yet, but it will come along. We like Impala, but you go around the industry and you'll hear, you know, 30 different MySQL is better than your SQL stories. So what we've seen is a proliferation of processing and analytic engines with a whole bunch of supporting plumbing, right? Data ingest, oh, by the way, security and data governance and data lineage and so on. And a steady improvement in the capabilities of HDFS, right? This is a really different platform than we brought to market in 2000 eight. The only thing that it really has in common, in my view, is that it's still called Hadoop, right? HDFS MapReduce is so widely deployed that it will always be used. MapReduce is so widely used that it's always going to be there. But I think we're going to see most new workloads embrace so…
AI assessment note: “percentage of cycles spent on MapReduce in Hadoop clusters generally is going to asymptotically approach zero”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Fascinating. So with all that money, ah, where, where, where do you go? What sort of the, what, what's Collidera like in, in, in five years? Is the end goal to be the, the new Oracle?
A Ah, well, ah, I spent some time at Oracle, ah, and I liked it a lot, and I've got a bunch of friends there, but you walk into a room of IT buyers and say you want to be the new Oracle, and it doesn't go good. Um, look, I think that there's an opportunity for the big data market to be much bigger than the relational database market was, right? That, uh, if, if min max median on traditional numerical data stored in tables is worth a hundred billion dollars a year, right? It just stands to reason that a thousand times more data that we can analyze using machine learning and predictive modeling Uh, so that we deduce real intent. I mean, that's got to be worth more, right? So I believe that there's an opportunity for new companies to emerge and to lead in that space, and we absolutely want to be a long-term, independent enterprise software company that is fundamentally about data, and that over time has a rich ecosystem of those partners. So yeah, we want to grow up to be successful at that scale, but I totally don't need a boat.
AI assessment note: “we absolutely want to be a long-term, independent enterprise software company”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Switching gears a bit, uh, from a, an entrepreneurial, entrepreneurial standpoint, so you had been doing, you know, this for a while before starting Cloudera, um, anything that sort of surprised you, whether that was, you know, from product standpoint, a sales standpoint, or, um, you know, how fast it went, how slow it went, any other things?
A Well, the back-to-back experience of running Sleepy Cat software, bootstraps, you know, uh, little 25 person engineering team for six years to Cloudera, uh, which is a venture-backed rocket ship with more than a billion dollars in capital raised, um, there's no business school classes about that. Um, I would say we believed this market was gonna be very, very big, uh, and that the opportunity was gonna be You know, exciting for us. It's been shocking at how quickly that has come true. It was astonishing how quickly we had to deal with companies like EMC, uh, VMware, Spinout, Pivotal, and IBM, and SAP in the space, right, uh, elbowing us aside. I mean, usually you get to be seven, eight years old before the big guys really notice who you are. At two and a half years, IBM was looking at what we were doing and, and offering their own open source alternatives, so. The interest in the market and the speed with which the market has grown has been really shocking.
AI assessment note: “It's been shocking at how quickly that has come true.”
Answered raw tape
D 4 · C 5 · P 5 · Cm 4 4.55
Q Well, you could now, right? Um, why, uh, so you, you mentioned some insights into the roadmap, but I'm, I'm, I'm, Um, the VC in me can, you know, resist asking what, why this versus an IPO or, uh, like Horton Works just did last week.
A We did not undertake that partnership because we were looking for the money. Um, as you guys may know, uh, Intel had been in the market with a product of their own that competed with all the other vendors in the space. Um, we began our Series F process Looking to raise what's called a mezzanine round. So we wanted to attract some high-quality public market investors to buy our private equity in advance of a planned IPO, right? And that went great. So T. Rowe Price led. Henry Ellenbogen, great guy. We like him. A number of other public market investors, well-respected, joined. We reserved a small part of that round for strategic. So we wanted to attract some companies that we thought would be a good signal. So Michael Dell invested, and that was great. Google, The inventor of Hadoop, Google Ventures, the inventor of MapReduce and GFS took a stake. We liked that a lot. We approached Intel with the idea that they would take a modest stake in the business as well, and obviously we liked the, the idea of an alignment with a hardware company as a big scale-out software company. The discussion quickly got more strategic. Um, you know, Intel, I think, saw a way to increase its influence over the open source community, because we're very well represented, and we've had a lot of success in most of the world. Um, we saw the advantage of being able to drive the roadmap in the way that I de…
AI assessment note: “to attract some high-quality public market investors to buy our private equity in advance of a planned IPO”
Redirected raw tape
D 2 · C 4 · P 4 · Cm 3 3.25
Q know, from, from a business standpoint, but also from a product standpoint, And, you know, maybe let's start with a product standpoint. So, you know, this Hadoop, but then it seems to be evolving all the time. This Spark, this data flow, and, you know, maybe MapReduce is not what people need, and all of this is evolving. Where do you think this is, um, evolving from a product standpoint?
A Um, let me talk a little bit about sort of the, the, the market status right now, and then I'll, I'll talk some about product. So there are a few ways to, to look at where the market is right now. Gartner, the analyst firm, has something that they call the hype cycle, and the place where they plot Hadoop right now is just over the peak of inflated expectation, about to enter the trough of, uh, uh, disillusionment. A trough of disillusionment on its way to the plateau of productivity. So that all sounds like it's pretty bad, but actually, you know, what it says is the market has been aging for a few years, and we really see that, right? I mean, we see large enterprises in finance, medicine, retail, government, ah, energy, healthcare, adopting big data platforms, not any longer for one or two projects, but, but pretty broadly, right? So we see the market maturing in a way that's, we think, is very promising. Ah, and our existing customers Are growing their installations very quickly. So we've got a long way to go before this platform is mature in the way that, say, Teradata or Oracle is, but hey, you know, I like the direction that the wind is blowing, uh, and the ground is pretty smooth.
AI assessment note: “let me talk a little bit about sort of the, the, the market status right now”
Redirected raw tape
D 2 · C 4 · P 4 · Cm 3 3.25
Q And so, uh, when you compete with the other Hadoop distribution vendors, what's the, um, what is the positioning? Is that a more mature ecosystem, a better editorial choice of what technologies matter?
A So I don't want to talk about how competitor, or how our customers make their decision, because really the customers are the people who should think about that. And I, I don't want to talk about positioning versus competitors, because frankly, you know, that's just a vendor slap fight. But let me tell you how we think about ourselves and who we try to be. So the company started in 2008, uh, and really I credit Amr Awadallah from Yahoo and Jeff Hammerbacher with this insight. The Hadoop that was then available, which was HDFS and MapReduce, was powerful, transformative. It could attack a bunch of problems at scale that were never addressable before. You just couldn't find a piece of software that would attack those problems. Provided that you could tolerate batch mode, latency delays, and that you had a naturally data parallel algorithm. Oh, by the way, and you know, extra points for being an exceptionally good Java programmer capable of mastering a new development framework. Um, it turns out not everybody on the planet falls into that camp, right? What Jeff and Ama realized was that The key innovation that Google had was the scale-out store. And the distributed processing engine that they dropped on top was merely an example of how you would attack data at scale. Those guys knew from the very earliest days that there would be other engines, a la HBase, and Spark, and Impala, an…
AI assessment note: “I don't want to talk about positioning versus competitors”