Hadoop

product on 6 shows · 52 statements across 38 episodes

In Depth the Neon Show the Official SaaStr Podcast the MAD Podcast the a16z Podcast David Senra

52 statements about Hadoop, every show

DAVID SENRA Assertion Partly supported
Söderström: In 2009 Spotify ran ML on world's largest Hadoop cluster
“When I came to Spotify, We already had one of these large data sets of music playlists, and so there were some very talented people there, specifically a guy named Erik Bernadsson, who at the time, this is 2009, was doing machine learning something called coll…”
Gustav Söderström Jun 7, 2026 ▶ 40:46 The Company Apple Couldn't Kill | Spotify Co-CEO Gustav Söderström
NEON SHOW Assertion Not checkable as stated
Bajaria: Early Hulu Adopted Hadoop After SQL Server Failed to Scale
“We were on SQL server before, and that thing would never run with the data that was coming in. And we tried to build it ourselves, failed. And then Hadoop came out. And we were like, why don't we build on Hadoop?”
Viral Bajaria Oct 24, 2025 ▶ 3:25 6sense Did What Google Couldn’t : Targeting Companies, Not People | Viral Bajaria
IN DEPTH Assertion Not checkable as stated
Handy: Enterprise data migration to the cloud truly began in 2020
“Enterprises really started moving to cloud for their data in twenty-twenty. There was like Hadoop stuff going on before that, but that like never really achieved any level of maturity or real penetration. And yes, Redshift and stuff like did exist prior to, an…”
Tristan Handy Dec 19, 2024 ▶ 48:08 Building a $4B data platform: Inside dbt Labs' unconventional path | Tristan Handy (Founder & CEO)
MAD Assertion Not checkable as stated
Housley: Moving from Hadoop to cloud data stacks has been very tough
“One of the things, one of the transitions that Joe and I went through, which I think a lot of people in this room went through, was the transition from the Hadoop world, from the previous big data world, into this new, like, cloud-based data engineering snack,…”
Matt Housley Oct 24, 2022 ▶ 1:03 Fundamentals of Data Engineering | Joe Reis and Matt Housley
SAASTR Insight
Ghodsi: On-prem open-source vendors sell services masquerading as software
“So the business models were not robust in the sense that really, it was really selling services masqueraded as software.”
Ali Ghodsi Dec 7, 2021 ▶ 13:14 The Future of AI, Open Source, and Enterprise SaaS with Databricks CEO Ali Ghodsi
MAD Opinion
Ghodsi: Hadoop was terrible for machine learning tasks
“The people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.”
Ali Ghodsi May 24, 2021 ▶ 2:40 Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
a16z Opinion
Naous: Legacy BI and Hadoop are only practical for executive decisions
“Well, you can say maybe we should use Hadoop or BI. Unfortunately, these are kind of older generation tools. They require specialized skills and abilities to be able to use them. And so they need armies of analysts to use. These are really only affordable for …”
Jad Naus May 16, 2019 ▶ 5:55 Everyone is an Analyst: Opportunities in Operational Analytics
a16z Opinion
Stanek: Hadoop and data warehouses are where data goes to die
“With all the investment in Hadoop and this and Hadoop that, you know, most companies are still data bankrupt. You know, Hadoop or Data Warehouse or whatever is a place where data goes to die”
Roman Stanek Jan 2, 2019 ▶ 7:47 a16z Podcast | Making the Most of the Data That Matters
a16z Insight
Stanek: Business users prefer spreadsheet interfaces over Spark and Hadoop
“Some of the most frequently used kind of data analytics tools extremely basic, because they actually look and feel like sheet of paper, like two-dimensional sheet of paper, and you know, so that's the problem with analytics, that on one hand, we have, you know…”
Roman Stanek Jan 2, 2019 ▶ 20:55 a16z Podcast | Making the Most of the Data That Matters
a16z Assertion Not checkable as stated
Moghe: Spark and Hadoop do not replace existing data warehouses
“Spark doesn't subsume data warehousing. Hadoop doesn't subsume, you know, streaming. So they're just like different technologies for different jobs.”
Prat Moghe Jan 2, 2019 ▶ 23:39 a16z Podcast | Making the Most of the Data That Matters
a16z Disclosure
Matei Zaharia interned at Facebook in 2007 when it had 300 employees
“I was a PhD student at UC Berkeley, and we actually started working with Hadoop users back in 2007. And I did, for example, an internship at Facebook when Facebook was only about 300 people and they were just starting to set up Hadoop.”
Matei Zaharia Jan 2, 2019 ▶ 1:44 a16z Podcast | A Conversation With the Inventor of Spark
a16z Assertion Supported
Legacy Hadoop tools like Hive, Pig, and Mahout now run on Spark
“So in particular you know, one of the things we saw is many of the projects that were built on top of Hadoop, such as Hive, which is a SQL processing at scale and Pig and Mahout for machine learning are starting to run on top of Spark as well, so that users of…”
Matei Zaharia Jan 2, 2019 ▶ 13:59 a16z Podcast | A Conversation With the Inventor of Spark
a16z Insight
Nguyen: Big Data Progress Is Driven by Cheaper Tech, Not Smarter People
“We don't necessarily get smarter over time. It's just that certain technologies get cheaper. They get, they become more available. So machine learning algorithms have always been around. The data that exists that you could collect has always been around. But i…”
Christopher Nguyen Jan 2, 2019 ▶ 7:20 a16z Podcast | Making Sense of Big Data, Machine Learning, and Deep Learning
MAD Prediction Not checkable as stated
Bob Muglia: Hadoop will not see much incremental investment
“So I think Hadoop is, is, is a past technology. I think it's, although it's still gonna, people will still use it still has a place, I think it's not an area where there's gonna be a lot of incremental additional investment.”
Bob Muglia Apr 9, 2018 ▶ 18:13 Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
MAD Disclosure
Goldman's compliance analytics rely on Hadoop and MapReduce batch processing
“So other than search, everything I described is batch processing. We use standard Hadoop. We use MapReduce.”
Mayur Thakur Dec 19, 2017 ▶ 16:48 Surveillance Platform for Banks // Mayur Thakur, Goldman Sachs (FirstMark's Data Driven)
a16z Assertion Supported
Analytics frameworks like Spark and Hadoop assume exclusive resource access
“Most things, Hadoop, Spark, Storm, they think they're running by themselves. And so they compete for resources in really interesting ways.”
Chandra Krintz Jul 15, 2017 ▶ 11:52 Chandra Krintz
MAD Assertion Not checkable as stated
Prat Moghe: Hiring qualified Hadoop DevOps engineers is exceptionally difficult
“Can you actually hire a good Hadoop DevOps engineer? Is it easy? I mean, you saw somebody stand up here saying they're recruiting. There's a reason, and it's because it's really hard to find these people, right?”
Prat Moghe May 24, 2017 ▶ 5:13 Big Data as a Service // Prat Moghe, Cazena (FirstMark's Data Driven)
MAD Assertion Not checkable as stated
Uber transitioned from ETL into Vertica to EL into Hadoop
“We went from an ETL model, where we scraped from, like, the original source, transformed the data and loaded to Vertica, to, like, just an EL model, where we just, like, just copy the data as soon as possible into, like, Hadoop, and all the transformation can …”
Praveen Murugesan Sep 30, 2016 ▶ 6:19 The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
MAD Assertion Not checkable as stated
6sense co-founders built the third largest Hadoop instance worldwide
“They're a Y Combinator company, built the third largest instance of Hadoop in the world, a real time predictive ad serving tool.”
Amanda Kahlow May 23, 2016 ▶ 20:18 Predictive Analytics for B2B Marketing // Amanda Kahlow, 6Sense [FirstMark's Data Driven]
MAD Prediction Not checkable as stated
Scholnick: AI commercialization will replicate the massive enterprise boom of Big Data
“And to me it feels like, Big data. Maybe six or seven years ago where companies were real waking up and realizing we have all these data assets. We need to do something with them. And that led to the rise of Hadoop and the Hadoop vendors and then, you know, a …”
Dan Scholnick Jan 25, 2016 ▶ 9:19 Investing in Data and A.I. // Dan Scholnick, Trinity Ventures (Hosted by FirstMark Capital)
MAD Disclosure
Srivas interviewed 50 Hadoop-using companies before founding MapR
“You know, well, before we started Mapper, I spoke to, like, about 40 or 50 people who were using Hadoop. 50 companies.”
M.C. Srivas Dec 17, 2015 ▶ 4:43 A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
MAD Assertion Not checkable as stated
Srivas: Almost every enterprise now has a Hadoop budget item
“Every company now has a Hadoop budget item. Almost every company.”
M.C. Srivas Dec 17, 2015 ▶ 25:50 A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
MAD Opinion
Groschupf: SQL on top of Hadoop was an unfortunate development
“Well, it's unfortunate, what I think is one of the most unfortunate thing that happened in the Hadoop space is kind of the introduction of SQL on top of Hadoop.”
Stefan Groschupf Dec 17, 2015 ▶ 5:42 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
MAD Prediction Not checkable as stated
Groschupf: Data center operating systems will be the next Hadoop killer
“That's a data center OS, and I really think that's the next Hadoop killer.”
Stefan Groschupf Dec 17, 2015 ▶ 13:21 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
MAD Insight
Deighton: Hadoop's schema-on-read innovation hasn't reached front-end users
“I think one of the great innovations of Hadoop is this idea of schema unread, but that's not realized through to the front end, to the user itself”
Anthony Deighton Sep 14, 2015 ▶ 8:47 Anthony Deighton, Qlik: Top 10 Requirements For Visual Analytics (Hosted by FirstMark Capital)
MAD Assertion Supported
Stoica: Early Hadoop was limited to batch processing
“So at that point, in big data space we there was Hadoop just started, but of course that was, by, back then it was mostly, you know, batch, computation, so you could do historical analysis, but not much more than that.”
Ion Stoica Apr 2, 2015 ▶ 3:03 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Ion Stoica Apr 2, 2015 ▶ 4:08 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Ion Stoica Apr 2, 2015 ▶ 11:23 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Insight
Dix: Ship code to where data lives, not data to code
“This is the key thing that we learned from Hadoop and Google's MapReduce framework, which is You want to ship the code to where the data lives, not the other way around.”
Paul Dix Apr 2, 2015 ▶ 5:16 Paul Dix, InfluxDB // Open-Source Time Series Database // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
IBM Watson used Apache UIMA and did not replace Hadoop
“I wouldn't assert that it replaces Hadoop. In fact, it's based on UEMA. It's an Apache project.”
Michael Karasick Jan 15, 2015 ▶ 23:14 Michael Karasick, IBM Watson // Data Driven #33 // Jan 2015 (Hosted by FirstMark Capital)
MAD Insight
Hadoop was designed for new data problems, not relational database issues
“What we didn't understand at the time, and it's been a pretty common feeling, is Hadoop wasn't built to solve the problem we'd been solving with relational databases. It was designed to solve a new problem, and it turned out that new problem was going to be ve…”
Mike Olson Dec 18, 2014 ▶ 5:10 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
MAD Assertion Not checkable as stated
Hadoop's shared-nothing architecture does not easily translate to OLAP or OLTP workloads
“Hadoop, this big scale-out, shared-nothing architecture, is good at much, but that architecture doesn't easily translate into OLAP or OLTP workloads.”
Mike Olson Dec 18, 2014 ▶ 8:45 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
MAD Prediction Not checkable as stated
Mike Olson predicts MapReduce compute cycles in Hadoop clusters will approach zero
“I think the percentage of cycles spent on MapReduce in Hadoop clusters generally is going to asymptotically approach zero. That's not because there will be less MapReduce happening, but because there will be so much of the other stuff happening.”
Mike Olson Dec 18, 2014 ▶ 10:45 Mike Olson, Cloudera // The Cloudera Story (Hosted by FirstMark Capital)
MAD Assertion Not checkable as stated
AppNexus processes 30 billion daily impressions on a 16-node Hadoop cluster
“We have a 16 node Hadoop cluster currently, and I have on my proposed budget for 2015, a 200 node Hadoop cluster so that we can really get our hands on all that raw data of the thirty billion impressions we're transacting daily.”
Catherine Williams Nov 20, 2014 ▶ 14:48 Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014
MAD Prediction Not checkable as stated
Large enterprise companies will eventually move production workloads onto Hadoop
“I think it will happen.”
Mike Abbott Oct 16, 2014 ▶ 4:01 Mike Abbott, KPCB // Data Driven #30 // Oct 2014 (Hosted by FirstMark Capital)
MAD Prediction Not checkable as stated
Tech startup valuations in mid-2014 are unwarranted and unsustainable
“There's gonna be a lot of you know, broken hearts and tears are gonna fall, because I don't think that, that those valuations are warranted, are sustainable and that's too bad.”
Chris Lynch Jun 26, 2014 ▶ 19:03 Chris Lynch, Atlas Venture // Data Driven #28 // June 2014 (Hosted by FirstMark Capital)
MAD Insight
Static Hadoop clusters in the cloud defeat the purpose of elasticity
“They just run long-running Hadoop clusters, which completely defeat the purpose of You know, how, how you can leverage the cloud to be dynamically adaptable to your workloads and things like that.”
Ashish Thusoo May 27, 2014 ▶ 15:15 Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
MAD Assertion Not checkable as stated
Unmodified on-premise Hadoop distributions fail in the cloud beyond 10 nodes
“If you just take a normal Hadoop distro and try to run it in the cloud, the chances are at 10 nodes it'll work fine as you start growing and, you know, as you start growing and growing and growing further. Things will start breaking because, you know, compute …”
Ashish Thusoo May 27, 2014 ▶ 22:34 Ashish Thusoo, Qubole // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
MAD Assertion Not checkable as stated
Steier: Database engines are being partially replaced by Hadoop
“And in particular, this sort of, the database engine itself is being replaced to a certain degree with things like Hadoop.”
Sandy Steier Mar 3, 2014 ▶ 3:43 Sandy Steier, 1010data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital)
MAD Assertion Supported
Tasso Argyros: Aster Data began developing its architecture in 2005
“We were thinking about this problem back in 2005, right? So that was pre-Hadoop”
Tasso Argyros Mar 3, 2014 ▶ 2:22 Tasso Argyros, Aster Data // Data Driven NYC 24 // February 2014 (Hosted by FirstMark Capital)
MAD Prediction Not checkable as stated
Borgman: Database market will see convergence of relational tech and Hadoop
“This is where the market's going. There's going to be this convergence of, you know, sort of relational database technology and Hadoop, and this is the future, and”
Justin Borgman Dec 5, 2013 ▶ 1:46 Panel discussion // Data Driven NYC #12 // Jan 2013
MAD Assertion Supported
Justin Borgman: Sears is making massive Hadoop investments to consolidate data
“Sears actually, there's been some interesting things written about Sears going in that direction. Which you think of, you know, major retail, you wouldn't think they would be, you know, compared to Facebook, but they are, and they're making huge investments in…”
Justin Borgman Dec 5, 2013 ▶ 40:53 Panel discussion // Data Driven NYC #12 // Jan 2013
MAD Insight
Gislason: Hadoop and NoSQL are poor for quantitative data aggregation
“Actually we found that Hadoop and most, kind of, no, no SQL solutions are not very good for, kind of, quantitative data when you need to aggregate and, kind of, go across these things.”
Hjalmar Gislason (Halmar) Dec 5, 2013 ▶ 11:31 Panel discussion // Data Driven NYC #9 // Nov 2012
MAD Prediction Not checkable as stated
Ping Li: Hadoop will be a definitive platform for big data workloads
“Hadoop I think will be a definitive platform for a lot of big data workloads.”
Ping Li Dec 5, 2013 ▶ 38:13 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
MAD Prediction Not checkable as stated
Matt Ocko predicts billion-dollar startups will solve core Hadoop infrastructure limits
“In each one of these, kind of, criteria, or vectors, or themes, there's a handful of billion dollar startups Ah yet to be, ah, yet to be built.”
Matt Ocko Dec 5, 2013 ▶ 49:20 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
MAD Opinion
Ping Li: The tech market does not need 10 more Hadoop infrastructure startups
“The world doesn't need You know, another 10 companies trying to solve the problems of Hadoop.”
Ping Li Dec 5, 2013 ▶ 50:09 Panel: Big data and VCs (Accel, IA Ventures, Data Collective) // Data Driven NYC #8// Oct 2012
MAD Opinion
Merriman: HBase is more directly competitive with MongoDB than Hadoop
“I think Hadoop, or HBase, which is a Hadoop subproject, that's more of a, that's something that's more of an alternative or competitive with Mongo, where you would look at A versus B”
Dwight Merriman Dec 5, 2013 ▶ 41:07 Fireside chat with Dwight Merriman // Data Driven #8 // Sep 2012 (interviewed by Matt Turck)
MAD Assertion Supported
Facebook built Hive to provide a SQL interface on Hadoop
“Facebook built Hive, right, because they needed a tool to sit on top of Hadoop, you know, to allow their business analysts to kind of sequel interface to this big data platform.”
Todd Papaioannou Dec 5, 2013 ▶ 20:08 Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012
MAD Opinion
Mike Driscoll: Running algorithms via Apache Mahout on Hadoop is too slow
“I think the problem with Mahoot is that anything, it's, many of these things are, if you run in Hadoop, you're slow. You need to be able to run in an environment that's fast”
Mike Driscoll Dec 5, 2013 ▶ 53:02 Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
MAD Prediction Not checkable as stated
Turck predicted the industry would realize Hadoop is extremely complex
“So everybody's gonna realize sooner or later that Hadoop is really complicated to install.”
Matt Turck Dec 5, 2013 ▶ 11:19 Matt Turck, Bloomberg Ventures // Data Driven NYC# 3 // Feb 2012
MAD Assertion Not checkable as stated
Hilary Mason: Hadoop is entirely impractical for real-time products
“All it is is a structure for running queries in parallel against data that you store in a redundant file system, and so it is entirely impractical for doing a real-time product.”
Hilary Mason Dec 5, 2013 ▶ 20:21 Hilary Mason, Bitly // Data Driven NYC #3 // Feb 2012

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.