Spark, every mention

11 scenes, the whole family · ← back to Spark

tap a year for its mentions
0041822015201620172018201920202021episodesmentions
0122015201620172018201920202021episodes it came up in
0021422015201620172018201920202021episodesmentions per episode

every year anyone Dave Burgess 3Praveen Murugesan 2Matt Turck 1M.C. Srivas 1Justin Borgman 1Julien Le Dem 1Ion Stoica 1Haoyuan Li 1

Verbatim, from the transcripts: the passages where Spark comes up

loading…

Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark) Apr 5, 2021 · 3 mentions

  • ▶ 8:16 Dave Burgess Uh, we have a query platform, which is Presto and Spark SQL. 2 times in the scene
  • ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.

Data Observability and Pipelines: OpenLineage and Marquez Feb 1, 2021 · 1 mention

  • ▶ 25:28 Julien Le Dem Spark, uh, built Spark SQL on top of Parquet, uh, Drill, uh, same thing used Parquet initially, um, you had Hive, Impala,

Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC) Mar 19, 2019 · 1 mention

  • ▶ 18:41 Justin Borgman That's our target, you know, market, and yes, there is Spark SQL, which is a sub-project, and that's the five percent overlap, but I think anyone who's trying to do high concurrency will usually realize that

The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark) Sep 30, 2016 · 2 mentions

  • ▶ 7:49 Praveen Murugesan So complex data processing tools, so two things that I'm gonna probably talk about today, like, in the interest of time, is like, ah, we built something called a Spark UDK, which is a tool which we started to build mostly with an intent of…
  • ▶ 8:15 Praveen Murugesan So, I'll start with, like, uh, UDK, which we call the Uber Developer Kit.

A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark) Apr 13, 2016 · 1 mention

A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark) Dec 17, 2015 · 1 mention

Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) Apr 2, 2015 · 6 mentions

  • ▶ 12:07 Matt Turck And to take just one of the, one of the things you mentioned, um, so maybe contrast Spark streaming with Storm, which also does micro-batching.
  • ▶ 20:58 unnamed speaker But, you know, there's been a lot of investment in optimizing Spark SQL, and, and I think a lot of the use cases that people in practice actually use it for are SQL-like workloads. 4 times in the scene
  • ▶ 24:15 Ion Stoica So, because you have unification in the case of Spark, well, you get the data from, from, ah, you know, for Spark streaming,
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.