Hadoop

40 statements across 29 episodes · 11 bullish · 15 bearish · 30 people on the record · first statement Dec 5, 2013 by Hilary Mason · across every show →

Everything said about Hadoop, oldest first

Dec 5, 2013 negative
Assertion Not checkable as stated
Hilary Mason: Hadoop is entirely impractical for real-time products
“All it is is a structure for running queries in parallel against data that you store in a redundant file system, and so it is entirely impractical for doing a real-time product.”
Hilary Mason Dec 5, 2013 ▶ 20:21 Hilary Mason, Bitly // Data Driven NYC #3 // Feb 2012
Dec 5, 2013 bearish
Prediction Not checkable as stated
Turck predicted the industry would realize Hadoop is extremely complex
“So everybody's gonna realize sooner or later that Hadoop is really complicated to install.”
Matt Turck Dec 5, 2013 ▶ 11:19 Matt Turck, Bloomberg Ventures // Data Driven NYC# 3 // Feb 2012
Dec 5, 2013 bearish
Opinion
Mike Driscoll: Running algorithms via Apache Mahout on Hadoop is too slow
“I think the problem with Mahoot is that anything, it's, many of these things are, if you run in Hadoop, you're slow. You need to be able to run in an environment that's fast”
Mike Driscoll Dec 5, 2013 ▶ 53:02 Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
Dec 5, 2013 neutral
Assertion Supported
Facebook built Hive to provide a SQL interface on Hadoop
“Facebook built Hive, right, because they needed a tool to sit on top of Hadoop, you know, to allow their business analysts to kind of sequel interface to this big data platform.”
Todd Papaioannou Dec 5, 2013 ▶ 20:08 Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012
Dec 5, 2013 neutral
Opinion
Merriman: HBase is more directly competitive with MongoDB than Hadoop
“I think Hadoop, or HBase, which is a Hadoop subproject, that's more of a, that's something that's more of an alternative or competitive with Mongo, where you would look at A versus B”
Dwight Merriman Dec 5, 2013 ▶ 41:07 Fireside chat with Dwight Merriman // Data Driven #8 // Sep 2012 (interviewed by Matt Turck)
Dec 5, 2013 bearish
Opinion
Ping Li: The tech market does not need 10 more Hadoop infrastructure startups
“The world doesn't need You know, another 10 companies trying to solve the problems of Hadoop.”
Ping Li Dec 5, 2013 ▶ 50:09 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
Dec 5, 2013 bullish
Prediction Not checkable as stated
Matt Ocko predicts billion-dollar startups will solve core Hadoop infrastructure limits
“In each one of these, kind of, criteria, or vectors, or themes, there's a handful of billion dollar startups Ah yet to be, ah, yet to be built.”
Matt Ocko Dec 5, 2013 ▶ 49:20 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
Dec 5, 2013 bullish
Prediction Not checkable as stated
Ping Li: Hadoop will be a definitive platform for big data workloads
“Hadoop I think will be a definitive platform for a lot of big data workloads.”
Ping Li Dec 5, 2013 ▶ 38:13 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
Dec 5, 2013 negative
Insight
Gislason: Hadoop and NoSQL are poor for quantitative data aggregation
“Actually we found that Hadoop and most, kind of, no, no SQL solutions are not very good for, kind of, quantitative data when you need to aggregate and, kind of, go across these things.”
Hjalmar Gislason (Halmar) Dec 5, 2013 ▶ 11:31 Panel discussion // Data Driven NYC #9 // Nov 2012
Dec 5, 2013 bullish
Assertion Supported
Justin Borgman: Sears is making massive Hadoop investments to consolidate data
“Sears actually, there's been some interesting things written about Sears going in that direction. Which you think of, you know, major retail, you wouldn't think they would be, you know, compared to Facebook, but they are, and they're making huge investments in…”
Justin Borgman Dec 5, 2013 ▶ 40:53 Panel discussion // Data Driven NYC #12 // Jan 2013
Dec 5, 2013 bullish
Prediction Not checkable as stated
Borgman: Database market will see convergence of relational tech and Hadoop
“This is where the market's going. There's going to be this convergence of, you know, sort of relational database technology and Hadoop, and this is the future, and”
Justin Borgman Dec 5, 2013 ▶ 1:46 Panel discussion // Data Driven NYC #12 // Jan 2013
Mar 3, 2014
Assertion Supported
Tasso Argyros: Aster Data began developing its architecture in 2005
“We were thinking about this problem back in 2005, right? So that was pre-Hadoop”
Tasso Argyros Mar 3, 2014 ▶ 2:22 Tasso Argyros, Aster Data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital)
Mar 3, 2014 neutral
Assertion Not checkable as stated
Steier: Database engines are being partially replaced by Hadoop
“And in particular, this sort of, the database engine itself is being replaced to a certain degree with things like Hadoop.”
Sandy Steier Mar 3, 2014 ▶ 3:43 Sandy Steier, 1010data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital)
May 27, 2014 negative
Assertion Not checkable as stated
Unmodified on-premise Hadoop distributions fail in the cloud beyond 10 nodes
“If you just take a normal Hadoop distro and try to run it in the cloud, the chances are at 10 nodes it'll work fine as you start growing and, you know, as you start growing and growing and growing further. Things will start breaking because, you know, compute …”
Ashish Thusoo May 27, 2014 ▶ 22:34 Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
May 27, 2014 negative
Insight
Static Hadoop clusters in the cloud defeat the purpose of elasticity
“They just run long-running Hadoop clusters, which completely defeat the purpose of You know, how, how you can leverage the cloud to be dynamically adaptable to your workloads and things like that.”
Ashish Thusoo May 27, 2014 ▶ 15:15 Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
Jun 26, 2014 bearish
Prediction Not checkable as stated
Tech startup valuations in mid-2014 are unwarranted and unsustainable
“There's gonna be a lot of you know, broken hearts and tears are gonna fall, because I don't think that, that those valuations are warranted, are sustainable and that's too bad.”
Chris Lynch Jun 26, 2014 ▶ 19:03 Chris Lynch, Atlas Venture // Data Driven #28 // June 2014 (Hosted by FirstMark Capital)
Oct 16, 2014 bullish
Prediction Not checkable as stated
Large enterprise companies will eventually move production workloads onto Hadoop
“I think it will happen.”
Mike Abbott Oct 16, 2014 ▶ 4:01 Mike Abbott, KPCB // Data Driven #30 // Oct 2014 (Hosted by FirstMark Capital)
Nov 20, 2014
Assertion Not checkable as stated
AppNexus processes 30 billion daily impressions on a 16-node Hadoop cluster
“We have a 16 node Hadoop cluster currently, and I have on my proposed budget for 2015, a 200 node Hadoop cluster so that we can really get our hands on all that raw data of the thirty billion impressions we're transacting daily.”
Catherine Williams Nov 20, 2014 ▶ 14:48 Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014
Dec 18, 2014 neutral
Assertion Not checkable as stated
Hadoop's shared-nothing architecture does not easily translate to OLAP or OLTP workloads
“Hadoop, this big scale-out, shared-nothing architecture, is good at much, but that architecture doesn't easily translate into OLAP or OLTP workloads.”
Mike Olson Dec 18, 2014 ▶ 8:45 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
Dec 18, 2014 bearish
Prediction Not checkable as stated
Mike Olson predicts MapReduce compute cycles in Hadoop clusters will approach zero
“I think the percentage of cycles spent on MapReduce in Hadoop clusters generally is going to asymptotically approach zero. That's not because there will be less MapReduce happening, but because there will be so much of the other stuff happening.”
Mike Olson Dec 18, 2014 ▶ 10:45 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
Dec 18, 2014 positive
Insight
Hadoop was designed for new data problems, not relational database issues
“What we didn't understand at the time, and it's been a pretty common feeling, is Hadoop wasn't built to solve the problem we'd been solving with relational databases. It was designed to solve a new problem, and it turned out that new problem was going to be ve…”
Mike Olson Dec 18, 2014 ▶ 5:10 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
Jan 15, 2015 neutral
Assertion Supported
IBM Watson used Apache UIMA and did not replace Hadoop
“I wouldn't assert that it replaces Hadoop. In fact, it's based on UEMA. It's an Apache project.”
Michael Karasick Jan 15, 2015 ▶ 23:14 Michael Karasick, IBM Watson // Data Driven #33 // Jan 2015 (Hosted by FirstMark Capital)
Apr 2, 2015
Insight
Dix: Ship code to where data lives, not data to code
“This is the key thing that we learned from Hadoop and Google's MapReduce framework, which is You want to ship the code to where the data lives, not the other way around.”
Paul Dix Apr 2, 2015 ▶ 5:16 Paul Dix, InfluxDB // Open-Source Time Series Database // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 neutral
Assertion Supported
Stoica: Early Hadoop was limited to batch processing
“So at that point, in big data space we there was Hadoop just started, but of course that was, by, back then it was mostly, you know, batch, computation, so you could do historical analysis, but not much more than that.”
Ion Stoica Apr 2, 2015 ▶ 3:03 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Ion Stoica Apr 2, 2015 ▶ 11:23 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 negative
Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Ion Stoica Apr 2, 2015 ▶ 4:08 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Sep 14, 2015 neutral
Insight
Deighton: Hadoop's schema-on-read innovation hasn't reached front-end users
“I think one of the great innovations of Hadoop is this idea of schema unread, but that's not realized through to the front end, to the user itself”
Anthony Deighton Sep 14, 2015 ▶ 8:47 Anthony Deighton, Qlik: Top 10 Requirements For Visual Analytics (Hosted by FirstMark Capital)
Dec 17, 2015 bullish
Prediction Not checkable as stated
Groschupf: Data center operating systems will be the next Hadoop killer
“That's a data center OS, and I really think that's the next Hadoop killer.”
Stefan Groschupf Dec 17, 2015 ▶ 13:21 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
Dec 17, 2015 negative
Opinion
Groschupf: SQL on top of Hadoop was an unfortunate development
“Well, it's unfortunate, what I think is one of the most unfortunate thing that happened in the Hadoop space is kind of the introduction of SQL on top of Hadoop.”
Stefan Groschupf Dec 17, 2015 ▶ 5:42 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
Dec 17, 2015
Disclosure
Srivas interviewed 50 Hadoop-using companies before founding MapR
“You know, well, before we started Mapper, I spoke to, like, about 40 or 50 people who were using Hadoop. 50 companies.”
M.C. Srivas Dec 17, 2015 ▶ 4:43 A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
Dec 17, 2015 positive
Assertion Not checkable as stated
Srivas: Almost every enterprise now has a Hadoop budget item
“Every company now has a Hadoop budget item. Almost every company.”
M.C. Srivas Dec 17, 2015 ▶ 25:50 A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
Jan 25, 2016 bullish
Prediction Not checkable as stated
Scholnick: AI commercialization will replicate the massive enterprise boom of Big Data
“And to me it feels like, Big data. Maybe six or seven years ago where companies were real waking up and realizing we have all these data assets. We need to do something with them. And that led to the rise of Hadoop and the Hadoop vendors and then, you know, a …”
Dan Scholnick Jan 25, 2016 ▶ 9:19 Investing in Data and A.I. // Dan Scholnick, Trinity Ventures (Hosted by FirstMark Capital)
May 23, 2016
Assertion Not checkable as stated
6sense co-founders built the third largest Hadoop instance worldwide
“They're a Y Combinator company, built the third largest instance of Hadoop in the world, a real time predictive ad serving tool.”
Amanda Kahlow May 23, 2016 ▶ 20:18 Predictive Analytics for B2B Marketing // Amanda Kahlow, 6Sense [FirstMark's Data Driven]
Sep 30, 2016
Assertion Not checkable as stated
Uber transitioned from ETL into Vertica to EL into Hadoop
“We went from an ETL model, where we scraped from, like, the original source, transformed the data and loaded to Vertica, to, like, just an EL model, where we just, like, just copy the data as soon as possible into, like, Hadoop, and all the transformation can …”
Praveen Murugesan Sep 30, 2016 ▶ 6:19 The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
May 24, 2017 negative
Assertion Not checkable as stated
Prat Moghe: Hiring qualified Hadoop DevOps engineers is exceptionally difficult
“Can you actually hire a good Hadoop DevOps engineer? Is it easy? I mean, you saw somebody stand up here saying they're recruiting. There's a reason, and it's because it's really hard to find these people, right?”
Prat Moghe May 24, 2017 ▶ 5:13 Big Data as a Service // Prat Moghe, Cazena (FirstMark's Data Driven)
Dec 19, 2017 neutral
Disclosure
Goldman's compliance analytics rely on Hadoop and MapReduce batch processing
“So other than search, everything I described is batch processing. We use standard Hadoop. We use MapReduce.”
Mayur Thakur Dec 19, 2017 ▶ 16:48 Surveillance Platform for Banks // Mayur Thakur, Goldman Sachs (FirstMark's Data Driven)
Apr 9, 2018 bearish
Prediction Not checkable as stated
Bob Muglia: Hadoop will not see much incremental investment
“So I think Hadoop is, is, is a past technology. I think it's, although it's still gonna, people will still use it still has a place, I think it's not an area where there's gonna be a lot of incremental additional investment.”
Bob Muglia Apr 9, 2018 ▶ 18:13 Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
May 24, 2021 negative
Opinion
Ghodsi: Hadoop was terrible for machine learning tasks
“The people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.”
Ali Ghodsi May 24, 2021 ▶ 2:40 Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
Oct 24, 2022 negative
Assertion Not checkable as stated
Housley: Moving from Hadoop to cloud data stacks has been very tough
“One of the things, one of the transitions that Joe and I went through, which I think a lot of people in this room went through, was the transition from the Hadoop world, from the previous big data world, into this new, like, cloud-based data engineering snack,…”
Matt Housley Oct 24, 2022 ▶ 1:03 Fundamentals of Data Engineering | Joe Reis and Matt Housley
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.