Spark

includes Spark SQL, Spark Streaming, Spark UDK, Spark Data Frames

11 statements across 11 episodes · 4 bullish · 2 bearish · 11 people on the record · first statement Jun 19, 2015 by Joseph Essas · said 15 times in 7 episodes since 2015 · across every show →

Mentions by year, the whole family

brought up most by Dave Burgess (3), Praveen Murugesan (2), Matt Turck (1), M.C. Srivas (1), Justin Borgman (1), Julien Le Dem (1), Ion Stoica (1), Haoyuan Li (1)

tap a year for its mentions
0041822015201620172018201920202021episodesmentions
0122015201620172018201920202021episodes it came up in
0021422015201620172018201920202021episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Spark, oldest first

Jun 19, 2015 neutral
Disclosure
Essas: OpenTable routes event streams through Kafka, Cassandra, and Spark
“All of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data.”
Joseph Essas Jun 19, 2015 ▶ 2:32 Joseph Essas, OpenTable // Mining Diner Talk (Hosted by FirstMark Capital)
Dec 17, 2015 positive
Assertion Not checkable as stated
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Stefan Groschupf Dec 17, 2015 ▶ 10:35 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
Apr 13, 2016 positive
Assertion Not checkable as stated
Barclays accelerated Spark jobs from hours to seconds using Alluxio
“They used Tachyon to accelerate Spark jobs from hours to seconds”
Haoyuan Li Apr 13, 2016 ▶ 12:03 A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
May 23, 2016
Disclosure
AXA US builds its analytics infrastructure on a Cloudera Hadoop stack
“So the one that we've got here in the U.S., it's primarily a Cloudera Hadoop stack that we've used a blueprint that was essentially blessed by our brethren over in French, in France and with that, we've got R and Python and Spark. We use some Dataiku along the…”
Louis DiModugno May 23, 2016 ▶ 10:44 Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven]
Sep 30, 2016
Assertion Not checkable as stated
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Praveen Murugesan Sep 30, 2016 ▶ 10:21 The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
Nov 9, 2016 negative
Opinion
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 15:56 Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]
Jul 13, 2017 neutral
Assertion Not checkable as stated
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Sean Kandel Jul 13, 2017 ▶ 6:41 Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven)
Apr 9, 2018 positive
Opinion
Bob Muglia: Apache Spark scenarios are complementary to Snowflake
“Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake.”
Bob Muglia Apr 9, 2018 ▶ 18:28 Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
Mar 19, 2019 negative
Opinion
Borgman: Apache Spark is not built to support high concurrency workloads
“Anyone who's trying to do high concurrency will usually realize that Spark is not built for high concurrency.”
Justin Borgman Mar 19, 2019 ▶ 18:41 Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
May 24, 2021
Disclosure
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Ali Ghodsi May 24, 2021 ▶ 23:12 Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
Jun 28, 2022 positive
Assertion Not checkable as stated
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Tristan Handy Jun 28, 2022 ▶ 7:34 The Next Layer of the Modern Data Stack | dbt's Tristan Handy
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.