Apache Hadoop, every mention
12 scenes (2021), the whole family · ← back to Apache Hadoop
every year 2021 anyone Matt Turck 80Stefan Groschupf 28Justin Borgman 22Mike Olson 19Tobi Knaup 15M.C. Srivas 14Florian Douetteau 12Ashish Thusoo 11Mike Driscoll 10Kirill Sheynkman 10
Verbatim, from the transcripts: the passages where Apache Hadoop comes up
Top 10 Trends in AI, Machine Learning and Data for 2022
- ▶ 2:54 Matt Turck And, uh, at the time this was the big data landscape, like all the cool kids talked about, uh, big data and, uh, how Hadoop was going to conquer the world.
- ▶ 7:32 Matt Turck Uh, you know, again, like people talked about Hadoop, uh, with bated breath, uh, many years ago.
Fireside Chat: Nick Schrock (Founder & CEO, Elementl) with Matt Turck (Partner, FirstMark)
- ▶ 4:38 Nick Schrock You know, Hadoop and then spark and now the cloud data warehouse.
Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
- ▶ 2:40 Ali Ghodsi And the people in Amplab that were doing machine learning, the math folks, they had to use this thing called Hadoop, which was just terrible.
- ▶ 11:17 Matt Turck Uh, yeah, I had the, um, uh, pleasure and honor of, like, hosting your co-founder and CEO at the time, Stoica in 2015, and the conversation, I rewatched it before this, and the conversation at the time was all about, you know, the, the,…
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 2:24 Dave Burgess And so Yahoo created Hadoop for those that don't know. 2 times in the scene
- ▶ 7:56 Dave Burgess So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. 3 times in the scene
Fireside Chat: Arjun Narayan (Founder & CEO, Materialize) with Matt Turck (Partner, FirstMark)
- ▶ 9:51 Arjun Narayan There was this big movement of NoSQL and Hadoop saying, you know, you're, and the fundamental message there was, was that the data sets were growing so large and your SQL databases were never going to scale and your SQL date, like if you… 3 times in the scene
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 10:13 Julien Le Dem Or the SQL and Hadoop types like Hive and Presto.
- ▶ 18:12 Julien Le Dem Your, um, data infrastructure, and you have an injection, and then you add a storage layer for streaming and for batch processing using things like Kafka or Sree or HDFS, and then you would have stream and batch processing, and usually you…
- ▶ 21:15 Julien Le Dem So, I mean, the Hive Metastore, for a long time, it was a, it's a de facto standard for data catalog or for worse on top of Hadoop and, um, in this environment.
Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
- ▶ 22:42 Matt Turck After HPC, um, slash Hadoop, slash Spark, slash Ray, what's the long-term future of parallel compute for data intensive workflows?