Apache Hadoop, every mention
31 scenes (2016), the whole family · ← back to Apache Hadoop
every year 2016 anyone Matt Turck 80Stefan Groschupf 28Justin Borgman 22Mike Olson 19Tobi Knaup 15M.C. Srivas 14Florian Douetteau 12Ashish Thusoo 11Mike Driscoll 10Kirill Sheynkman 10
Verbatim, from the transcripts: the passages where Apache Hadoop comes up
Becoming an Internet Company // Raymie Stata, Altiscale [FirstMark's Data Driven]
- ▶ 5:44 Raymie Stata Hadoop and big data platforms kind of got, got birthed, and, and you create that feedback, you know, that instantaneous feedback cycle right there.
- ▶ 16:49 Raymie Stata Um, and, you know, Hadoop at Yahoo became kind of our standard infrastructure, and we made a large, large investment in that. 3 times in the scene
A Process for Discovery // Hilary Mason, Fast Forward Labs [FirstMark's Data Driven]
- ▶ 9:15 Hilary Mason I know Ramey was up here earlier talking a bit about his work with Hadoop. 4 times in the scene
Making Big Data Accessible Using the Cloud // Ashish Thusoo, Qubole [FirstMark's Data Driven]
- ▶ 3:07 Ashish Thusoo So, um, you know, big data, essentially the emergence of these new systems, we hear about systems like Hadoop, Spark, Hive, and so on and so forth.
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
- ▶ 5:49 Praveen Murugesan So we built, like, a Hadoop data lake, which really, ah, the fundamental difference between what we built from Hadoop, ah, versus, like, Vertica was, like, we left all the original data in the source data, data set, databases itself, like,… 5 times in the scene
- ▶ 8:46 Praveen Murugesan And then, ah, it completely abstracts out, like, the runtime environment, so you don't have to really, like, play with your Hadoop configs and so on.
A Kafka-Powered Real-Time Streaming Platform // Neha Narkhede, Confluent [FirstMark's Data Driven]
- ▶ 5:04 Neha Narkhede If you want to ingest all this data in a streaming fashion, maybe load it up in Hadoop where you want to do analytics about how users are viewing products or buying products. 3 times in the scene
- ▶ 11:42 Neha Narkhede What happens to the connector that pulls data from that database into Hadoop? 2 times in the scene
- ▶ 15:46 Neha Narkhede It powers many, many, ah, stream processing business logic, and then it is the source of truth pipeline for Hadoop.
- ▶ 22:17 unnamed speaker Um, so, Kafka Streaming is, ah, basically a library, and then you provide your platform, which runs Kafka with Kafka Streaming, and by doing that you, you're multiplying clusters for any customer, basically, because the customer needs,… 2 times in the scene
Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven]
- ▶ 1:19 Nitay Joffe For those of you that are not familiar, Astor Data built one of the first sort of enterprise MPP class databases, and in particular, they had a lot of very interesting IP around how you actually can take SQL and MapReduce and actually do…
- ▶ 8:48 Nitay Joffe And so in the upper left, you see a suite of BI systems, Hadoop, Spark, and so on and so forth. 2 times in the scene
Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven]
- ▶ 10:44 Louis DiModugno Um, so the one that we've got here in the U.S., it's, uh, it's primarily a Cloudera Hadoop stack, uh, that we've, uh, uh, used a blueprint that was essentially blessed by our, our brethren over in French, in France,
Predictive Analytics for B2B Marketing // Amanda Kahlow, 6Sense [FirstMark's Data Driven]
- ▶ 20:18 Amanda Kahlow Um, they're a Y Combinator company, built the third largest instance of Hadoop in the world, a real time predictive ad serving tool.
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 3:13 Haoyuan Li It's a memory speed virtual distributed storage systems, and it sits between the application, application layer, like distributed application layer, like Spark, MapReduce, like HBase, like H two O, we'll have another talk later on about H…
- ▶ 3:39 Haoyuan Li Or you have other open source storage, like Gluster, like HDFS, like SAF, or you have private cloud storage, like, uh, like OpenStack Swift, or you have traditional storage, like EMC, NetApp, IBM, Huawei.
- ▶ 5:52 Haoyuan Li You have Spark, you have MapReduce, you have Flink, you have HBase, Presto.
- ▶ 6:00 Haoyuan Li Like you have Amazon, SIII, Swift, OSS from Alibaba, you have HDFS, Gloucester, and this list goes on and on.
- ▶ 13:44 Haoyuan Li They run Spark and MapReduce on Alluxio on top of Gloucester, and they, you can use Alluxio to manage memory plus SD, and it's a SaaS company. 2 times in the scene
The Solitude of the Data Team Manager // Florian Douetteau, Dataiku (Hosted by FirstMark)
- ▶ 3:13 Florian Douetteau He starts, he starts, uh, hearing some old stories of companies that, uh, started to install Hadoop, and it was not working, it was too slow, not responsive enough compared to SQL technologies and so on.
- ▶ 5:35 Florian Douetteau Well, whatever happens, he will need to add new roles within his data teams, like a business data analyst, data scientist, Hadoop engineers, and so on. 2 times in the scene
- ▶ 20:21 Florian Douetteau Um, because the, the, many companies have the vision of having, like, data, data analysts, largely speaking, moving from a SQL, and possibly SAS world, to a kind of Hadoop-ish, Spark-ish world, on the long term. 2 times in the scene
Using AI to Predict the Performance of Text // Kieran Snyder, Textio (Data Driven NYC / FirstMark)
- ▶ 6:44 Kieran Snyder Um, big data analytics Hadoop engineer might be a job that some people in this room would apply for or hire. 2 times in the scene
A Fireside Chat With MongoDB CTO Eliot Horowitz (Data Driven NYC / FirstMark)
- ▶ 17:09 Matt Turck And the last part of that convergence is, uh, possibly, you know, databases and analytics, uh, which sort of happened a little bit in the Hadoop world.
A Fireside Chat with Benchmark General Partner Peter Fenton (Data Driven NYC / FirstMark)
- ▶ 20:32 Peter Fenton There's the packaging model, which we're in at Hortonworks, because there isn't one owner of Hadoop. 2 times in the scene
- ▶ 23:10 Peter Fenton I think what's happening in the Hadoop ecosystem is you have two companies, Cloudera, at least two, you could argue MapR and Hortonworks that are packagers that have the Red Hat business model that, like Red Hat, by the way, Red Hat is not…
Problem Solving With Geospatial Data // Javier de la Torre, CartoDB (Hosted by FirstMark Capital)
- ▶ 20:55 Javier de la Torre So, um, so CartDB actually can actually talk to different, you know, like, uh, NoSQL stores, you know, like, we, we, we have connectors to Hadoop, or different, you know, like, we have to, we have different, different connectors that we… 2 times in the scene
Large Scale Decision Support Systems // Satya Ramachandran, Neustar (Hosted by FirstMark Capital)
- ▶ 1:51 Satya Ramachandran Jowin Data was a BI system on top of Hadoop.
- ▶ 12:58 Satya Ramachandran The Hadoop systems by hand, and we used to manage them, and so on and so forth. 2 times in the scene
Investing in Data and A.I. // Dan Scholnick, Trinity Ventures (Hosted by FirstMark Capital)
- ▶ 9:34 Dan Scholnick And, and, and that led to the rise of, um, Hadoop and, uh, the Hadoop vendors and then, you know, a bunch of very successful companies in the big data space. 2 times in the scene
The Benefits of Fast Business Intelligence // Amir Orad, Sisense (Hosted by FirstMark Capital)
- ▶ 17:23 Matt Turck When, when, when, uh, presumably you go talk to IT buyers or even marketing buyers, uh, this jungle of companies out there, uh, and people have heard all the buzz terms and all the Hadoop and the Spark and 4 times in the scene