why aren't all 40 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Bob Muglia: Hadoop will not see much incremental investment
“So I think Hadoop is, is, is a past technology. I think it's, although it's still gonna, people will still use it still has a place, I think it's not an area where there's gonna be a lot of incremental additional investment.”
Prediction Not checkable as stated
Groschupf: Data center operating systems will be the next Hadoop killer
“That's a data center OS, and I really think that's the next Hadoop killer.”
Prediction Not checkable as stated
Tech startup valuations in mid-2014 are unwarranted and unsustainable
“There's gonna be a lot of you know, broken hearts and tears are gonna fall, because I don't think that, that those valuations are warranted, are sustainable and that's too bad.”
Prediction Not checkable as stated
Mike Olson predicts MapReduce compute cycles in Hadoop clusters will approach zero
“I think the percentage of cycles spent on MapReduce in Hadoop clusters generally is going to asymptotically approach zero. That's not because there will be less MapReduce happening, but because there will be so much of the other stuff happening.”
Assertion Not checkable as stated
Unmodified on-premise Hadoop distributions fail in the cloud beyond 10 nodes
“If you just take a normal Hadoop distro and try to run it in the cloud, the chances are at 10 nodes it'll work fine as you start growing and, you know, as you start growing and growing and growing further. Things will start breaking because, you know, compute …”
Prediction Not checkable as stated
Matt Ocko predicts billion-dollar startups will solve core Hadoop infrastructure limits
“In each one of these, kind of, criteria, or vectors, or themes, there's a handful of billion dollar startups Ah yet to be, ah, yet to be built.”
Opinion
Ghodsi: Hadoop was terrible for machine learning tasks
“The people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.”
Prediction Not checkable as stated
Scholnick: AI commercialization will replicate the massive enterprise boom of Big Data
“And to me it feels like, Big data. Maybe six or seven years ago where companies were real waking up and realizing we have all these data assets. We need to do something with them. And that led to the rise of Hadoop and the Hadoop vendors and then, you know, a …”
Opinion
Groschupf: SQL on top of Hadoop was an unfortunate development
“Well, it's unfortunate, what I think is one of the most unfortunate thing that happened in the Hadoop space is kind of the introduction of SQL on top of Hadoop.”
Insight
Static Hadoop clusters in the cloud defeat the purpose of elasticity
“They just run long-running Hadoop clusters, which completely defeat the purpose of You know, how, how you can leverage the cloud to be dynamically adaptable to your workloads and things like that.”
Insight
Gislason: Hadoop and NoSQL are poor for quantitative data aggregation
“Actually we found that Hadoop and most, kind of, no, no SQL solutions are not very good for, kind of, quantitative data when you need to aggregate and, kind of, go across these things.”
Opinion
Ping Li: The tech market does not need 10 more Hadoop infrastructure startups
“The world doesn't need You know, another 10 companies trying to solve the problems of Hadoop.”
Opinion
Merriman: HBase is more directly competitive with MongoDB than Hadoop
“I think Hadoop, or HBase, which is a Hadoop subproject, that's more of a, that's something that's more of an alternative or competitive with Mongo, where you would look at A versus B”
Opinion
Mike Driscoll: Running algorithms via Apache Mahout on Hadoop is too slow
“I think the problem with Mahoot is that anything, it's, many of these things are, if you run in Hadoop, you're slow. You need to be able to run in an environment that's fast”
Prediction Not checkable as stated
Turck predicted the industry would realize Hadoop is extremely complex
“So everybody's gonna realize sooner or later that Hadoop is really complicated to install.”
Assertion Not checkable as stated
Hilary Mason: Hadoop is entirely impractical for real-time products
“All it is is a structure for running queries in parallel against data that you store in a redundant file system, and so it is entirely impractical for doing a real-time product.”
Assertion Not checkable as stated
Housley: Moving from Hadoop to cloud data stacks has been very tough
“One of the things, one of the transitions that Joe and I went through, which I think a lot of people in this room went through, was the transition from the Hadoop world, from the previous big data world, into this new, like, cloud-based data engineering snack,…”
Assertion Not checkable as stated
Prat Moghe: Hiring qualified Hadoop DevOps engineers is exceptionally difficult
“Can you actually hire a good Hadoop DevOps engineer? Is it easy? I mean, you saw somebody stand up here saying they're recruiting. There's a reason, and it's because it's really hard to find these people, right?”
Assertion Not checkable as stated
6sense co-founders built the third largest Hadoop instance worldwide
“They're a Y Combinator company, built the third largest instance of Hadoop in the world, a real time predictive ad serving tool.”
Assertion Not checkable as stated
Srivas: Almost every enterprise now has a Hadoop budget item
“Every company now has a Hadoop budget item. Almost every company.”
Insight
Deighton: Hadoop's schema-on-read innovation hasn't reached front-end users
“I think one of the great innovations of Hadoop is this idea of schema unread, but that's not realized through to the front end, to the user itself”
Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Assertion Supported
IBM Watson used Apache UIMA and did not replace Hadoop
“I wouldn't assert that it replaces Hadoop. In fact, it's based on UEMA. It's an Apache project.”
Assertion Not checkable as stated
Hadoop's shared-nothing architecture does not easily translate to OLAP or OLTP workloads
“Hadoop, this big scale-out, shared-nothing architecture, is good at much, but that architecture doesn't easily translate into OLAP or OLTP workloads.”
Insight
Hadoop was designed for new data problems, not relational database issues
“What we didn't understand at the time, and it's been a pretty common feeling, is Hadoop wasn't built to solve the problem we'd been solving with relational databases. It was designed to solve a new problem, and it turned out that new problem was going to be ve…”
Prediction Not checkable as stated
Large enterprise companies will eventually move production workloads onto Hadoop
“I think it will happen.”
Assertion Supported
Justin Borgman: Sears is making massive Hadoop investments to consolidate data
“Sears actually, there's been some interesting things written about Sears going in that direction. Which you think of, you know, major retail, you wouldn't think they would be, you know, compared to Facebook, but they are, and they're making huge investments in…”
Prediction Not checkable as stated
Borgman: Database market will see convergence of relational tech and Hadoop
“This is where the market's going. There's going to be this convergence of, you know, sort of relational database technology and Hadoop, and this is the future, and”
Prediction Not checkable as stated
Ping Li: Hadoop will be a definitive platform for big data workloads
“Hadoop I think will be a definitive platform for a lot of big data workloads.”
Disclosure
Goldman's compliance analytics rely on Hadoop and MapReduce batch processing
“So other than search, everything I described is batch processing. We use standard Hadoop. We use MapReduce.”
Assertion Not checkable as stated
Uber transitioned from ETL into Vertica to EL into Hadoop
“We went from an ETL model, where we scraped from, like, the original source, transformed the data and loaded to Vertica, to, like, just an EL model, where we just, like, just copy the data as soon as possible into, like, Hadoop, and all the transformation can …”
Disclosure
Srivas interviewed 50 Hadoop-using companies before founding MapR
“You know, well, before we started Mapper, I spoke to, like, about 40 or 50 people who were using Hadoop. 50 companies.”
Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Assertion Supported
Stoica: Early Hadoop was limited to batch processing
“So at that point, in big data space we there was Hadoop just started, but of course that was, by, back then it was mostly, you know, batch, computation, so you could do historical analysis, but not much more than that.”
Insight
Dix: Ship code to where data lives, not data to code
“This is the key thing that we learned from Hadoop and Google's MapReduce framework, which is You want to ship the code to where the data lives, not the other way around.”
Assertion Not checkable as stated
AppNexus processes 30 billion daily impressions on a 16-node Hadoop cluster
“We have a 16 node Hadoop cluster currently, and I have on my proposed budget for 2015, a 200 node Hadoop cluster so that we can really get our hands on all that raw data of the thirty billion impressions we're transacting daily.”
Assertion Not checkable as stated
Steier: Database engines are being partially replaced by Hadoop
“And in particular, this sort of, the database engine itself is being replaced to a certain degree with things like Hadoop.”
Assertion Supported
Tasso Argyros: Aster Data began developing its architecture in 2005
“We were thinking about this problem back in 2005, right? So that was pre-Hadoop”
Assertion Supported
Facebook built Hive to provide a SQL interface on Hadoop
“Facebook built Hive, right, because they needed a tool to sit on top of Hadoop, you know, to allow their business analysts to kind of sequel interface to this big data platform.”