Apache Spark, every mention
23 scenes (2016), the whole family · ← back to Apache Spark
every year 2016 anyone Matt Turck 59Haoyuan Li 13Ali Ghodsi 11Stefan Groschupf 10Praveen Murugesan 10Julien Le Dem 9Tobi Knaup 6Matt Housley 6Christopher Nguyen 6Prat Moghe 5
Verbatim, from the transcripts: the passages where Apache Spark comes up
Making Big Data Accessible Using the Cloud // Ashish Thusoo, Qubole [FirstMark's Data Driven]
- ▶ 3:07 Ashish Thusoo So, um, you know, big data, essentially the emergence of these new systems, we hear about systems like Hadoop, Spark, Hive, and so on and so forth.
Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]
- ▶ 15:56 Nick (Domino Data Lab) Probably the most, well, I don't know whether it's surprising or not, um, I think there's still more hype around Spark than actual value extraction from it. 5 times in the scene
Making On-Demand Delivery Profitable // Jeremy Stanley, Instacart (Data Driven NYC / FirstMark)
- ▶ 17:37 Jeremy Stanley Um, we use Spark, uh, when we have to, uh, try to stay away from having to use those systems as long as we can.
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
- ▶ 4:41 Praveen Murugesan And, ah, on top of HDFS, we basically have, like, Spark, and, ah, Presto, and Hive.
- ▶ 7:18 Praveen Murugesan So, a few things, like I talked about, strict schema management, so we actually built, like, a central schema repository which is used for schema management, and then, ah, we unlocked, like, a whole bunch of new tools with, like, data on… 2 times in the scene
- ▶ 8:19 Praveen Murugesan So, basically, think about it as, like, if you are a first-time Spark developer, you, you, your, your goal is not to learn Spark, like, in-depth and, like, play around with, like, a hundred different knobs. 5 times in the scene
- ▶ 15:24 Praveen Murugesan And, ah, so given that it's in a Hive UDF, you also can use it on Spark. 2 times in the scene
Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven]
- ▶ 8:48 Nitay Joffe And so in the upper left, you see a suite of BI systems, Hadoop, Spark, and so on and so forth. 2 times in the scene
The Path to A.I. Augmented Human Intelligence // Christopher Nguyen, Arimo [FirstMark's Data Driven]
- ▶ 5:02 Christopher Nguyen I did not pre-aggregate this, and this is, we use Spark for in-memory processing, um, but the idea is that we've made it so simple to use that I, as a business user, I have not done any, ah, you know, any kind of SQL.
- ▶ 9:03 Christopher Nguyen If you, ah, if I move too fast through here, just Google Spark Deep Learning Reference Architecture. 4 times in the scene
- ▶ 12:03 Christopher Nguyen One is Spark Only.
Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven]
- ▶ 10:57 Louis DiModugno Um, and with that, we've got R and Python and Spark.
The Journey to Information for Everyone // Prakash Nanduri, Paxata (Hosted by FirstMark)
- ▶ 13:33 Prakash Nanduri The technology around distributed computing and what's going on with in-memory scale-out, Spark, and Luxia, and all the wonderful solutions that, ah, enabling technologies that are coming up are really important.
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 0:51 Haoyuan Li Funding commuter of Apache Spark as well.
- ▶ 2:08 Haoyuan Li The same lab produced Apache Mesos as well as Apache Spark, and we open sourced it second year April, which was around three years ago, on the Apache two-point-one license, and the latest release is version 1.1, which was actually March… 2 times in the scene
- ▶ 5:05 Haoyuan Li Say for example, Sim Lab will produce Apache Spark.
- ▶ 5:52 Haoyuan Li You have Spark, you have MapReduce, you have Flink, you have HBase, Presto.
- ▶ 11:22 Haoyuan Li Basically, they run Spark and Spark SQL on top of Aluxio. 7 times in the scene
- ▶ 17:01 Haoyuan Li people, like, inclined to believe in the project created by this lab, so when I started this project, it's already, like, four years later than other projects from the lab, like, ah, particularly, like, Mesos and Spark.
The Solitude of the Data Team Manager // Florian Douetteau, Dataiku (Hosted by FirstMark)
- ▶ 3:56 Florian Douetteau You might want to do everything in Spark, but you would find out that maybe the geo-analytics team in the company is really into Postgres and PostGIS, and want to do everything in Postgres and PostGIS.
- ▶ 20:21 Florian Douetteau Um, because the, the, many companies have the vision of having, like, data, data analysts, largely speaking, moving from a SQL, and possibly SAS world, to a kind of Hadoop-ish, Spark-ish world, on the long term. 2 times in the scene
Combining Machine Learning With Expert Human Judgement // Eric Colson, Stitch Fix
- ▶ 15:07 Eric Colson We use the usual tools, um, S-Tree and Spark.
The Benefits of Fast Business Intelligence // Amir Orad, Sisense (Hosted by FirstMark Capital)
- ▶ 17:23 Matt Turck When, when, when, uh, presumably you go talk to IT buyers or even marketing buyers, uh, this jungle of companies out there, uh, and people have heard all the buzz terms and all the Hadoop and the Spark and 2 times in the scene