Apache Hadoop, every mention

185 scenes, the whole family · ← back to Apache Hadoop

tap a year for its mentions
005010100202013201420152016201720182019202020212022202320242025episodesmentions
010202013201420152016201720182019202020212022202320242025episodes it came up in
0051010202013201420152016201720182019202020212022202320242025episodesmentions per episode

every year anyone Matt Turck 80Stefan Groschupf 28Justin Borgman 22Mike Olson 19Tobi Knaup 15M.C. Srivas 14Florian Douetteau 12Ashish Thusoo 11Mike Driscoll 10Kirill Sheynkman 10

Verbatim, from the transcripts: the passages where Apache Hadoop comes up

loading…

Rewriting Success: What InfluxDB 3.0 Teaches About Scaling—and Scrapping—Your Core Tech May 8, 2025 · 1 mention

  • ▶ 20:14 Evan Kaplan Even the most, even the most, you know, the most industrial oriented customers had built something on Hadoop.

Snowflake CEO on Winning the AI Arms Race Apr 10, 2025 · 1 mention

Trino, Iceberg and the Battle for the Lakehouse | Justin Borgman, CEO, Starburst Jan 30, 2025 · 5 mentions

  • ▶ 5:48 Justin Borgman Hadoop was just, uh, gaining momentum as, as the data lake of the day, and we were some of the first to think about using this for analytical data warehousing purposes.
  • ▶ 8:33 Justin Borgman And I'd already seen the power of open source just watching Hadoop's rise and the viral adoption of that technology that I thought, you know what, open source is a really powerful distribution mechanism.
  • ▶ 13:50 Justin Borgman You know, call it maybe 2011, 12 time frame when Hadoop was gaining momentum. 3 times in the scene

Dataiku's Secret to Scaling AI in Global Enterprises | Florian Douetteau, CEO, Dataiku Dec 12, 2024 · 2 mentions

  • ▶ 15:22 Florian Douetteau Whatever you do with data, especially because data was moving from existing databases to, in our case back then, it was Hadoop and starting to move to the cloud. 2 times in the scene

The Single Platform for Everyday AI - Fireside Chat with Florian Douetteau (Dataiku) & Matt Turck Jul 5, 2023 · 4 mentions

A Conversation with Chris Wiggins - Author of "How Data Happened" May 31, 2023 · 1 mention

  • ▶ 20:24 Chris Wiggins Um, then we decided, sorry, then it was decided that we should build our own Hadoop on-premises, um, which was the style of the time.

Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota Jan 30, 2023 · 2 mentions

  • ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system. 2 times in the scene

Fundamentals of Data Engineering | Joe Reis and Matt Housley Oct 24, 2022 · 4 mentions

  • ▶ 1:03 Matt Housley One of the things, one of the transitions that Joe and I went through, which I think a lot of people in this room went through, was the transition from the Hadoop world, from the previous big data world, into this new, like, cloud-based…
  • ▶ 2:50 Joe Reis Or Hadoop. 2 times in the scene
  • ▶ 30:50 Matt Housley Um, I, I think when we, I, I think, and correct me if I misunderstood the question, but I think when we talk about not defining, setting definitions around technology, we mean specifically not saying that data engineering is about Spark,…

Behavioral Data Creation for AI | Snowplow Co-Founder & CEO Alex Dean Oct 24, 2022 · 1 mention

Separating Data Hype From Substance | Fireside Chat with Mode Co-Founder Benn Stancil Oct 10, 2022 · 1 mention

  • ▶ 34:45 Matt Turck Like, I'm referring to, you know, there was like the whole way for like Hadoop and then, uh, you know, cloud vendors at some point, like everybody was saying, well, cloud is going to like solve it all, and then that evolved to like…

Fireside Chat: Felix Van de Maele (Co-Founder & CEO, Collibra) with Matt Turck (Partner, FirstMark) Mar 17, 2022 · 1 mention

Top 10 Trends in AI, Machine Learning and Data for 2022 Oct 27, 2021 · 2 mentions

  • ▶ 2:54 Matt Turck And, uh, at the time this was the big data landscape, like all the cool kids talked about, uh, big data and, uh, how Hadoop was going to conquer the world.
  • ▶ 7:32 Matt Turck Uh, you know, again, like people talked about Hadoop, uh, with bated breath, uh, many years ago.

Fireside Chat: Nick Schrock (Founder & CEO, Elementl) with Matt Turck (Partner, FirstMark) Jun 21, 2021 · 1 mention

Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark) May 24, 2021 · 2 mentions

  • ▶ 2:40 Ali Ghodsi And the people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.
  • ▶ 11:17 Matt Turck Uh, yeah, I had the, um, uh, pleasure and honor of, like, hosting your co-founder and CEO at the time, Stoica in 2015, and the conversation, I rewatched it before this, and the conversation at the time was all about, you know, the, the,…

Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark) Apr 5, 2021 · 5 mentions

  • ▶ 2:24 Dave Burgess And so Yahoo created Hadoop for those that don't know. 2 times in the scene
  • ▶ 7:56 Dave Burgess So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. 3 times in the scene

Fireside Chat: Arjun Narayan (Founder & CEO, Materialize) with Matt Turck (Partner, FirstMark) Mar 15, 2021 · 3 mentions

  • ▶ 9:51 Arjun Narayan There was this big movement of NoSQL and Hadoop saying, you know, you're, and the fundamental message there was, was that the data sets were growing so large and your SQL databases were never going to scale and your SQL date, like if you… 3 times in the scene

Data Observability and Pipelines: OpenLineage and Marquez Feb 1, 2021 · 3 mentions

  • ▶ 10:13 Julien Le Dem Or the SQL and Hadoop types like Hive and Presto.
  • ▶ 18:12 Julien Le Dem Your, um, data infrastructure, and you have an injection, and then you add a storage layer for streaming and for batch processing using things like Kafka or Sree or HDFS, and then you would have stream and batch processing, and usually you…
  • ▶ 21:15 Julien Le Dem So, I mean, the Hive Metastore, for a long time, it was a, it's a de facto standard for data catalog or for worse on top of Hadoop and, um, in this environment.

Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark) Feb 1, 2021 · 1 mention

  • ▶ 22:42 Matt Turck After HPC, um, slash Hadoop, slash Spark, slash Ray, what's the long-term future of parallel compute for data intensive workflows?

How to Resurrect Innovation in the Enterprise // Amr Awadallah, Google (FirstMark's Data Driven NYC) Feb 20, 2020 · 2 mentions

  • ▶ 23:48 Matt Turck So Cloudera was one of the key distributions of Hadoop. 2 times in the scene

How to Answer Data Questions Without Being Miserable // Ahmed Elsamadisi, Narrator (Data Driven NYC) Nov 13, 2019 · 1 mention

  • ▶ 16:58 unnamed speaker Um, are there any major issues with say loading it and reading it in HDFS?

Fireside Chat: Mike Volpi, General Partner, Index Ventures (FirstMark's Data Driven NYC) Nov 13, 2019 · 2 mentions

  • ▶ 7:59 Mike Volpi Uh, Red Hat, uh, the first Hadoop companies all fit into that category.
  • ▶ 14:59 Matt Turck And so you, you've had those very, uh, successful investments in sort of the, the core infrastructure of, uh, you know, big data, I guess, I don't know if anybody still talks about big data, uh, but, um, you know, that, that sort of like,…

Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC) Jun 12, 2019 · 3 mentions

  • ▶ 1:12 FJ Yang And, ah, how many people have worked with what kind of different types of big data technologies like Hadoop, Spark, and Kafka, and others?
  • ▶ 6:44 FJ Yang The other kind of movement that has gotten pretty popular is the open source Hadoop movement. 2 times in the scene

Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC) Mar 19, 2019 · 11 mentions

  • ▶ 0:22 Justin Borgman Hadapt was actually one of the first SQL engines for Hadoop. 3 times in the scene
  • ▶ 1:57 Justin Borgman Hive is a SQL, uh, offering on top of Hadoop. 3 times in the scene
  • ▶ 2:55 Justin Borgman Um, this is, uh, maybe a little hard for some of you guys to see, uh, but it shows at the top, uh, a variety of popular BI tools that you can connect to using ODBC or JDBC drivers, and on the bottom, a variety of data sources, and you'll…
  • ▶ 8:56 Justin Borgman So, ah, one of the great things that the Hadoop era brought to us, ah, with all the billions of dollars of venture capital poured into Hadoop companies, is some great open file formats. 3 times in the scene
  • ▶ 19:40 Justin Borgman So if you're already storing data in, let's say, Oracle or Oracle and Hadoop, and you want to do a join between the two, for example, um, we would push down as much filtering and predicates to the Oracle system as possible to sort of…

Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC) Oct 17, 2018 · 2 mentions

  • ▶ 11:37 Mike Tuchen And for folks that had moved to, um, uh, the premise Hadoop world, and took advantage of creating this concept of a data lake, almost everyone that I'm talking to now is saying, now I'm looking at the cloud, and I'm now trying to decide… 2 times in the scene

The Launch of Dataiku 5 // Florian Douetteau, Dataiku (FirstMark's Data Driven NYC) Sep 17, 2018 · 1 mention

3 Heretical Ideas on the Future of Data // Ajay Kulkarni, TimescaleDB (FirstMark's Data Driven NYC) Sep 17, 2018 · 2 mentions

  • ▶ 3:02 Ajay Kulkarni There's Amazon, who published the Dynoa paper, and, and these really, you know, created the foundation for the, for the, the wave of the first commercial big data systems, which included Hadoop, Cassandra, and Mongo.
  • ▶ 7:21 Ajay Kulkarni Technologies and companies like, you know, Hadoop and Cassandra and Mongo, we believe this new era is going to lead to another wave of foundational companies and technologies, and we believe, we believe time skill is poised to be one of…

Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven) Apr 9, 2018 · 11 mentions

  • ▶ 13:20 Bob Muglia That removed most of, if not all, of the limitations that held back companies that were working with both traditional data warehouses, as well as with big data and Hadoop.
  • ▶ 15:17 Bob Muglia Are, are trying to make Hadoop work, and are struggling with that to analyze machine generated data, and so they come to Snowflake from the machine generated side, and they use us for, for analytics associated with that, and, and those… 3 times in the scene
  • ▶ 17:19 Matt Turck So the Hadoop and Spark ecosystems are friend or foe? 6 times in the scene
  • ▶ 28:45 unnamed speaker So you mentioned a lot about unstructured data and Hadoop.

Surveillance Platform for Banks // Mayur Thakur, Goldman Sachs (FirstMark's Data Driven) Dec 19, 2017 · 1 mention

Where Should Machines Go to Learn? // Auren Hoffman, SafeGraph (FirstMark's Data Driven) Nov 20, 2017 · 1 mention

  • ▶ 11:43 Auren Hoffman So, these are companies like Palantir, they're the BI tools, even things like, you know, Hadoop or Spark, um, you know, basically any, most of these companies are basically, let me take your own data and help you make better decisions with…

Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven) Jul 13, 2017 · 1 mention

  • ▶ 0:38 Sean Kandel So if we look at kind of where there have been major advancements, uh, on kind of the data platform side, you know, it's already been mentioned today, and everyone's well aware of this, technologies like the cloud and Hadoop have made it…

Big Data as a Service // Prat Moghe, Cazena (FirstMark's Data Driven) May 24, 2017 · 5 mentions

The Power of GPU Analytics // Todd Mostak, MapD (FirstMark's Data Driven) Apr 6, 2017 · 3 mentions

  • ▶ 1:05 Todd Mostak People are building massive data lakes, Hadoop clusters, throwing their data on disk, but the issue comes when they try to analyze this data.
  • ▶ 3:45 Todd Mostak So MapD would sit in the middle of a standard data, um, kind of data analytics ecosystem, possibly pulling in from a data warehouse, pulling in from a data lake like Hadoop, alternatively pulling in streaming data.
  • ▶ 15:28 Todd Mostak We have JDBC, we have a Hadoop connector,

Becoming an Internet Company // Raymie Stata, Altiscale [FirstMark's Data Driven] Dec 8, 2016 · 4 mentions

  • ▶ 5:44 Raymie Stata Hadoop and big data platforms kind of got, got birthed, and, and you create that feedback, you know, that instantaneous feedback cycle right there.
  • ▶ 16:49 Raymie Stata Um, and, you know, Hadoop at Yahoo became kind of our standard infrastructure, and we made a large, large investment in that. 3 times in the scene

A Process for Discovery // Hilary Mason, Fast Forward Labs [FirstMark's Data Driven] Dec 8, 2016 · 4 mentions

  • ▶ 9:15 Hilary Mason I know Ramey was up here earlier talking a bit about his work with Hadoop. 4 times in the scene

Making Big Data Accessible Using the Cloud // Ashish Thusoo, Qubole [FirstMark's Data Driven] Nov 9, 2016 · 1 mention

  • ▶ 3:07 Ashish Thusoo So, um, you know, big data, essentially the emergence of these new systems, we hear about systems like Hadoop, Spark, Hive, and so on and so forth.

The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark) Sep 30, 2016 · 6 mentions

  • ▶ 5:49 Praveen Murugesan So we built, like, a Hadoop data lake, which really, ah, the fundamental difference between what we built from Hadoop, ah, versus, like, Vertica was, like, we left all the original data in the source data, data set, databases itself, like,… 5 times in the scene
  • ▶ 8:46 Praveen Murugesan And then, ah, it completely abstracts out, like, the runtime environment, so you don't have to really, like, play with your Hadoop configs and so on.

A Kafka-Powered Real-Time Streaming Platform // Neha Narkhede, Confluent [FirstMark's Data Driven] Jun 16, 2016 · 8 mentions

  • ▶ 5:04 Neha Narkhede If you want to ingest all this data in a streaming fashion, maybe load it up in Hadoop where you want to do analytics about how users are viewing products or buying products. 3 times in the scene
  • ▶ 11:42 Neha Narkhede What happens to the connector that pulls data from that database into Hadoop? 2 times in the scene
  • ▶ 15:46 Neha Narkhede It powers many, many, ah, stream processing business logic, and then it is the source of truth pipeline for Hadoop.
  • ▶ 22:17 unnamed speaker Um, so, Kafka Streaming is, ah, basically a library, and then you provide your platform, which runs Kafka with Kafka Streaming, and by doing that you, you're multiplying clusters for any customer, basically, because the customer needs,… 2 times in the scene

Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven] Jun 16, 2016 · 3 mentions

  • ▶ 1:19 Nitay Joffe For those of you that are not familiar, Astor Data built one of the first sort of enterprise MPP class databases, and in particular, they had a lot of very interesting IP around how you actually can take SQL and MapReduce and actually do…
  • ▶ 8:48 Nitay Joffe And so in the upper left, you see a suite of BI systems, Hadoop, Spark, and so on and so forth. 2 times in the scene

Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven] May 23, 2016 · 1 mention

  • ▶ 10:44 Louis DiModugno Um, so the one that we've got here in the U.S., it's, uh, it's primarily a Cloudera Hadoop stack, uh, that we've, uh, uh, used a blueprint that was essentially blessed by our, our brethren over in French, in France,

Predictive Analytics for B2B Marketing // Amanda Kahlow, 6Sense [FirstMark's Data Driven] May 23, 2016 · 1 mention

  • ▶ 20:18 Amanda Kahlow Um, they're a Y Combinator company, built the third largest instance of Hadoop in the world, a real time predictive ad serving tool.

A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark) Apr 13, 2016 · 6 mentions

  • ▶ 3:13 Haoyuan Li It's a memory speed virtual distributed storage systems, and it sits between the application, application layer, like distributed application layer, like Spark, MapReduce, like HBase, like H two O, we'll have another talk later on about H…
  • ▶ 3:39 Haoyuan Li Or you have other open source storage, like Gluster, like HDFS, like SAF, or you have private cloud storage, like, uh, like OpenStack Swift, or you have traditional storage, like EMC, NetApp, IBM, Huawei.
  • ▶ 5:52 Haoyuan Li You have Spark, you have MapReduce, you have Flink, you have HBase, Presto.
  • ▶ 6:00 Haoyuan Li Like you have Amazon, SIII, Swift, OSS from Alibaba, you have HDFS, Gloucester, and this list goes on and on.
  • ▶ 13:44 Haoyuan Li They run Spark and MapReduce on Alluxio on top of Gloucester, and they, you can use Alluxio to manage memory plus SD, and it's a SaaS company. 2 times in the scene

The Solitude of the Data Team Manager // Florian Douetteau, Dataiku (Hosted by FirstMark) Apr 13, 2016 · 5 mentions

  • ▶ 3:13 Florian Douetteau He starts, he starts, uh, hearing some old stories of companies that, uh, started to install Hadoop, and it was not working, it was too slow, not responsive enough compared to SQL technologies and so on.
  • ▶ 5:35 Florian Douetteau Well, whatever happens, he will need to add new roles within his data teams, like a business data analyst, data scientist, Hadoop engineers, and so on. 2 times in the scene
  • ▶ 20:21 Florian Douetteau Um, because the, the, many companies have the vision of having, like, data, data analysts, largely speaking, moving from a SQL, and possibly SAS world, to a kind of Hadoop-ish, Spark-ish world, on the long term. 2 times in the scene

Using AI to Predict the Performance of Text // Kieran Snyder, Textio (Data Driven NYC / FirstMark) Mar 18, 2016 · 2 mentions

  • ▶ 6:44 Kieran Snyder Um, big data analytics Hadoop engineer might be a job that some people in this room would apply for or hire. 2 times in the scene

A Fireside Chat With MongoDB CTO Eliot Horowitz (Data Driven NYC / FirstMark) Mar 18, 2016 · 1 mention

  • ▶ 17:09 Matt Turck And the last part of that convergence is, uh, possibly, you know, databases and analytics, uh, which sort of happened a little bit in the Hadoop world.

A Fireside Chat with Benchmark General Partner Peter Fenton (Data Driven NYC / FirstMark) Mar 18, 2016 · 3 mentions

  • ▶ 20:32 Peter Fenton There's the packaging model, which we're in at Hortonworks, because there isn't one owner of Hadoop. 2 times in the scene
  • ▶ 23:10 Peter Fenton I think what's happening in the Hadoop ecosystem is you have two companies, Cloudera, at least two, you could argue MapR and Hortonworks that are packagers that have the Red Hat business model that, like Red Hat, by the way, Red Hat is not…

Problem Solving With Geospatial Data // Javier de la Torre, CartoDB (Hosted by FirstMark Capital) Feb 21, 2016 · 2 mentions

  • ▶ 20:55 Javier de la Torre So, um, so CartDB actually can actually talk to different, you know, like, uh, NoSQL stores, you know, like, we, we, we have connectors to Hadoop, or different, you know, like, we have to, we have different, different connectors that we… 2 times in the scene

Large Scale Decision Support Systems // Satya Ramachandran, Neustar (Hosted by FirstMark Capital) Feb 21, 2016 · 3 mentions

Investing in Data and A.I. // Dan Scholnick, Trinity Ventures (Hosted by FirstMark Capital) Jan 25, 2016 · 2 mentions

  • ▶ 9:34 Dan Scholnick And, and, and that led to the rise of, um, Hadoop and, uh, the Hadoop vendors and then, you know, a bunch of very successful companies in the big data space. 2 times in the scene

The Benefits of Fast Business Intelligence // Amir Orad, Sisense (Hosted by FirstMark Capital) Jan 25, 2016 · 4 mentions

  • ▶ 17:23 Matt Turck When, when, when, uh, presumably you go talk to IT buyers or even marketing buyers, uh, this jungle of companies out there, uh, and people have heard all the buzz terms and all the Hadoop and the Spark and 4 times in the scene

10 Commandments for BI in Big Data, Shant Hovsepian, Arcadia Data (Data Driven NYC / FirstMark) Dec 17, 2015 · 5 mentions

  • ▶ 5:37 Shant Hovsepian So, there are lots of, we're very fortunate now, uh, Hadoop, Big Data, Mongo, Cassandra, the industry has really, really, really developed a lot, and we have a lot of amazing native tools that you can use for analytics. 2 times in the scene
  • ▶ 6:56 Shant Hovsepian Thanks to things like Yarn, Mesos, we're seeing a whole new resurgence of operating systems.
  • ▶ 19:49 Shant Hovsepian So, if they're a Hadoop vendor, if they're a cloud server, or if they're using big data on the cloud, you know, they're used to utility billing, or they're used to 2 times in the scene

B2B Big Data Challenges, Nick Mehta, Gainsight (Data Driven NYC / FirstMark Capital) Dec 17, 2015 · 2 mentions

  • ▶ 1:02 Nick Mehta You've, you still have Yarn, you still have Mesosphere, you have MapR, you have Hadoop, but unfortunately you also have salespeople, right, that actually have to act on all this stuff, and that's the biggest challenge, and I'm going to…
  • ▶ 1:02 Nick Mehta You've, you still have Yarn, you still have Mesosphere, you have MapR, you have Hadoop, but unfortunately you also have salespeople, right, that actually have to act on all this stuff, and that's the biggest challenge, and I'm going to…

A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark) Dec 17, 2015 · 29 mentions

  • ▶ 0:16 Matt Turck So MAPAR, um, obviously one of the key, uh, perhaps the key, um, uh, distribution of, of Hadoop, uh, and, and that's such a cornerstone of the big data ecosystem.
  • ▶ 2:09 M.C. Srivas And, uh, let's look at what are they doing, and, you know, and, uh, there was this thing called Hadoop everybody was using, and I had a lot of friends who, 5 times in the scene
  • ▶ 12:19 M.C. Srivas So, about six years ago, uh, 30 or 40, uh, Silicon Valley CTOs and CEOs, kind of, um, pro, uh, pro bono, moved to India and built a system on Hadoop to, uh, basically, with biometric identification to, to, uh, identify every person in…
  • ▶ 16:57 M.C. Srivas Now they have six node Hadoop clusters there. 2 times in the scene
  • ▶ 19:34 Matt Turck So, Hadoop fundamentally is an open source, uh, product. 4 times in the scene
  • ▶ 23:24 M.C. Srivas So you have this funky thing where, you know, Impala is controlled by one company, Yarn is controlled by one company, or, or, uh, now Kudu or something.
  • ▶ 25:07 Matt Turck So, where are we in the, I guess, evolution of Fadoop and Big Data as a mainstream product? 3 times in the scene
  • ▶ 32:17 unnamed speaker I was using Hadoop before that, but moving to MapR was a big upgrade for us. 6 times in the scene
  • ▶ 34:33 unnamed speaker Um, Stefan mentioned that Mesosphere, or in his presentation, he added as the Hadoop killer.
  • ▶ 34:40 M.C. Srivas I think he meant Yarn. 5 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.