Hadoop Distributed File System, every mention

17 scenes · ← back to Hadoop Distributed File System

tap a year for its mentions
00821542013201420152016201720182019202020212022202320242025episodesmentions
0242013201420152016201720182019202020212022202320242025episodes it came up in
001.52342013201420152016201720182019202020212022202320242025episodesmentions per episode

every year anyone Matt Turck 7Praveen Murugesan 3Haoyuan Li 3Tobi Knaup 2Ion Stoica 2Vance Loiselle 1Todd Papaioannou 1Stefan Groschupf 1Justin Borgman 1Diego Oppenheimer 1

Verbatim, from the transcripts: the passages where Hadoop Distributed File System comes up

loading…

Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy Jan 23, 2025 · 1 mention

  • ▶ 40:15 Ben Rogojan (Seattle Data Guy) You can just have one data store, uh, and then pick which, uh, data compute engine you want to use, which is actually something that Facebook, when I, when I was leaving was already doing, they had kind of an HDFS sort of, uh, under layer…

Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC) Mar 19, 2019 · 1 mention

  • ▶ 10:34 Justin Borgman So, you know, a few years ago, a lot of people were moving data from, let's say, a large database or data warehousing appliance into HDFS, the Hadoop file system, but now many people are moving that into S three.

Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven) Apr 9, 2018 · 1 mention

  • ▶ 17:41 Bob Muglia Uh, Hadoop, Hadoop's history with HDFS very much makes it a storage-based system,

Building an Operating System for AI // Diego Oppenheimer, Algorithmia (FirstMark's Data Driven) Mar 2, 2018 · 1 mention

The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark) Sep 30, 2016 · 3 mentions

  • ▶ 4:28 Praveen Murugesan But that said, what we really created was, like, using HDFS, like, a data lake, where, uh, we basically copied the whole data sets from, like, uh, whatever we get from, like, analytical logs or, like, all our business data sources, too,… 2 times in the scene
  • ▶ 7:18 Praveen Murugesan So, a few things, like I talked about, strict schema management, so we actually built, like, a central schema repository which is used for schema management, and then, ah, we unlocked, like, a whole bunch of new tools with, like, data on…

A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark) Apr 13, 2016 · 3 mentions

  • ▶ 12:53 Haoyuan Li And below Aluxio, they use Aluxio to manage both HDFS and SAF. 3 times in the scene

A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark) Dec 17, 2015 · 5 mentions

  • ▶ 8:29 Matt Turck How does that, uh, differ from the core HDFS? 4 times in the scene
  • ▶ 32:35 unnamed speaker Uh, and so, my question is, I saw that as one of the biggest effects to my organ, I own a software company, for context, uh, so I have a lot of engineers working with things, and being able to work with, uh, the MapR file system was just…

The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer Dec 17, 2015 · 1 mention

  • ▶ 1:16 Stefan Groschupf That's why we looked into the Google papers and implemented MapReduce and the HDFS, et cetera, et cetera.

Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) Apr 2, 2015 · 5 mentions

  • ▶ 4:26 Ion Stoica So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.
  • ▶ 6:25 Matt Turck Uh, can you, uh, maybe help us understand, um, so this HDFS, which is the storage, uh, and MapReduce, which is, uh, the processing, 3 times in the scene
  • ▶ 24:04 Ion Stoica Or, you take the data out of STORM through maybe HDFS and push it in Impala, right?

Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital) Jan 16, 2015 · 1 mention

  • ▶ 18:27 Chris Wiggins There's a lot of use of, um, both S-III and HDFS, so we do a lot of MapReduce jobs to chomp up huge JSON buckets of our own design in S-III via EC-II in order to render it down to a bite-sized data table so we can beat it down with…

Vance Loiselle, Sumo Logic // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital) Sep 22, 2014 · 1 mention

  • ▶ 5:01 Vance Loiselle So, you know, early on we started building a platform like, oh, let's use Hadoop and HDFS and, you know, back then there was no real-time nature.

Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital) Sep 22, 2014 · 2 mentions

  • ▶ 19:40 Tobi Knaup So if you're using Hadoop, you, you would be running HDFS, and HDFS would take care of it. 2 times in the scene

Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012 Dec 5, 2013 · 1 mention

  • ▶ 15:08 Todd Papaioannou One, we have a file system, HDFS, which will scale forever for probably like 99.9% of the companies on the planet.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.