Apache Spark

includes Apache Spark Streaming, Apache Spark MLlib

14 statements across 6 episodes · 6 bullish · 3 bearish · 6 people on the record · first statement Sep 22, 2014 by Tobi Knaup · said 220 times in 55 episodes since 2014 · across every show →

Mentions by year, the whole family

brought up most by Matt Turck (59), Haoyuan Li (13), Ali Ghodsi (11), Stefan Groschupf (10), Praveen Murugesan (10), Julien Le Dem (9), Tobi Knaup (6), Matt Housley (6)

tap a year for its mentions
003086015201420152016201720182019202020212022202320242025episodesmentions
0815201420152016201720182019202020212022202320242025episodes it came up in
0047.5815201420152016201720182019202020212022202320242025episodesmentions per episode
2025 5 mentions in 1 episode
2024 2 mentions in 2 episodes 1 per episode
2023 2 mentions in 2 episodes 1 per episode
2022 8 mentions in 3 episodes 3 per episode
2021 51 mentions in 12 episodes 4 per episode
2019 11 mentions in 3 episodes 4 per episode
2018 12 mentions in 5 episodes 2 per episode
2017 13 mentions in 4 episodes 3 per episode
2016 46 mentions in 12 episodes 4 per episode
2015 57 mentions in 8 episodes 7 per episode
2014 13 mentions in 3 episodes 4 per episode

every mention, scene by scene, with the transcript →

Everything said about Apache Spark, oldest first

Sep 22, 2014
Assertion Supported
Knaup: Apache Spark was the first application built on Apache Mesos
“Spark, ah, there's a little known fact, Spark was actually the first, ah, application that was ever built on top of Mesos, and it was sort of the, you know, the demo app for Mesos in the beginning.”
Tobi Knaup Sep 22, 2014 ▶ 5:42 Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital)
Apr 2, 2015 positive
Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 neutral
Disclosure
Stoica: Databricks offers only cloud services, not a custom Spark distribution
“We don't have our own distribution. We provide only the service.”
Ion Stoica Apr 2, 2015 ▶ 17:30 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Assertion Supported
Stoica: Apache Spark has exceeded 500 active open-source contributors
“We exceeded 500 contributors. It is the most active big data project right now, Spark.”
Ion Stoica Apr 2, 2015 ▶ 16:47 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015
Assertion Supported
Stoica: Apache Spark originally ran on Mesos before adding YARN and standalone support
“As originally was built to run on top of Mesos. Today is working on Yarn, working, you know, standalone, and is working also in addition to HDFS, you know, imports and exports data to many other data sources.”
Ion Stoica Apr 2, 2015 ▶ 7:07 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015
Assertion Supported
Stoica: Databricks still provides the majority of Apache Spark open-source contributions
“A lot of contributions, still the majority of contributions comes from Databricks.”
Ion Stoica Apr 2, 2015 ▶ 17:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Assertion Supported
Stoica: Apache Spark supports all major data workloads with one engine
“While we spark, You can use only one engine and only one API to support all these workloads.”
Ion Stoica Apr 2, 2015 ▶ 12:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Apr 2, 2015 positive
Prediction Not checkable as stated
Stoica predicts big data infrastructure will eventually unify around one system
“Now, I do think that looking forward you are going to see more and more of this unification. This happens in many other industries and technologies. I do think this will happen in big data because it's so much easier if you have only one system than if you nee…”
Ion Stoica Apr 2, 2015 ▶ 21:53 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
Oct 17, 2018 negative
Assertion Not checkable as stated
Early Apache Spark was so unstable that clusters crashed within two hours
“At the time we made that bet, Spark was barely working. I mean, seriously. It was a really promising technology, very complete programming model. We loved the fact it was in memory, but we couldn't find a single customer that could keep their cluster running f…”
Mike Tuchen Oct 17, 2018 ▶ 16:31 Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC)
Feb 1, 2021 negative
Opinion
Apache Spark failed to effectively shrink down to single-node scale
“One of the things that you find with things like Spark is that they really failed to shrink down and, Effectively do computing at the single node scale.”
Wes McKinney Feb 1, 2021 ▶ 24:22 Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
Feb 1, 2021 negative
Assertion Supported
Running Apache Spark on a single node is slower than Pandas
“You can use spark at the single node scale as an alternative to pandas through the koalas interface, but you'll find that for many workloads, it's simply slower than pandas, which is not super impressive.”
Wes McKinney Feb 1, 2021 ▶ 24:30 Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
Feb 17, 2021
Assertion Not checkable as stated
Netflix uses S3, Spark, Presto, and Snowflake for data querying
“We use SG as so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake as our query engines.”
Savin Goyal Feb 17, 2021 ▶ 3:23 Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
Jan 23, 2025 bullish
Prediction Not checkable as stated
Rogojan: Enterprise migration away from Spark and Kafka will take years
“I think initially just picking up Kafka, Spark, those are still so heavily in use that even if they go out in the next three to five years, it's going to take time to migrate and things of that nature.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 22:06 Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.