Spark, every mention
11 scenes, the whole family · ← back to Spark
tap a year for its mentions
every year anyone Dave Burgess 3Praveen Murugesan 2Matt Turck 1M.C. Srivas 1Justin Borgman 1Julien Le Dem 1Ion Stoica 1Haoyuan Li 1
Verbatim, from the transcripts: the passages where Spark comes up
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 8:16 Dave Burgess Uh, we have a query platform, which is Presto and Spark SQL. 2 times in the scene
- ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 25:28 Julien Le Dem Spark, uh, built Spark SQL on top of Parquet, uh, Drill, uh, same thing used Parquet initially, um, you had Hive, Impala,
Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
- ▶ 18:41 Justin Borgman That's our target, you know, market, and yes, there is Spark SQL, which is a sub-project, and that's the five percent overlap, but I think anyone who's trying to do high concurrency will usually realize that
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
- ▶ 7:49 Praveen Murugesan So complex data processing tools, so two things that I'm gonna probably talk about today, like, in the interest of time, is like, ah, we built something called a Spark UDK, which is a tool which we started to build mostly with an intent of…
- ▶ 8:15 Praveen Murugesan So, I'll start with, like, uh, UDK, which we call the Uber Developer Kit.
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 11:22 Haoyuan Li Basically, they run Spark and Spark SQL on top of Aluxio.
A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
- ▶ 24:40 M.C. Srivas So, and in Spark SQL.
Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
- ▶ 12:07 Matt Turck And to take just one of the, one of the things you mentioned, um, so maybe contrast Spark streaming with Storm, which also does micro-batching.
- ▶ 20:58 unnamed speaker But, you know, there's been a lot of investment in optimizing Spark SQL, and, and I think a lot of the use cases that people in practice actually use it for are SQL-like workloads. 4 times in the scene
- ▶ 24:15 Ion Stoica So, because you have unification in the case of Spark, well, you get the data from, from, ah, you know, for Spark streaming,