Apache Spark

product on 14 shows · 25 statements across 13 episodes · said 378 times in 103 episodes since 2014

the MAD Podcast 215 the a16z Podcast 104 the Official SaaStr Podcast 19 Latent Space 13 20VC 10 Lenny's Podcast 5 the Neon Show 4 No Priors 2 the Y Combinator Startup Podcast 1 the Knowledge Project 1 A Product Market Fit Show 1 Top Founders 1 Big Technology 1 TBPN 1

Mentions by year, every show

tap a year for its mentions
005010100202014201520162017201820192020202120222023202420252026episodesmentions
010202014201520162017201820192020202120222023202420252026episodes it came up in
003106202014201520162017201820192020202120222023202420252026episodesmentions per episode

the MAD Podcast 215the a16z Podcast 104the Official SaaStr Podcast 19Latent Space 1320VC 10Lenny's Podcast 5the Neon Show 4No Priors 26 more shows

2026 14 mentions in 6 episodes 2 per episode
2025 37 mentions in 14 episodes 3 per episode
2024 4 mentions in 4 episodes 1 per episode
2023 4 mentions in 3 episodes 1 per episode
2022 9 mentions in 4 episodes 2 per episode
2021 64 mentions in 14 episodes 5 per episode
2019 97 mentions in 19 episodes 5 per episode
2018 13 mentions in 6 episodes 2 per episode
2017 20 mentions in 7 episodes 3 per episode
2016 49 mentions in 14 episodes 4 per episode
2015 54 mentions in 9 episodes 6 per episode
2014 13 mentions in 3 episodes 4 per episode

every mention on every show, scene by scene, with the transcript →

25 statements about Apache Spark, every show

Zaharia: Scaling down from bulk ingestion to serving is easier than scaling up
“It turned out that, you know, it's easier to go from that bod thing that's really good at the scale and ingesting and super low cost and create versions in it that have the speed and features of the, you know, super easy to use, like smaller data for business …”
Matei Zaharia Jun 24, 2026 ▶ 55:55 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
20VC Assertion Contradicted
Databricks had hundreds of thousands of free Spark users when Gabrisco joined
“When I joined, we had hundreds of thousands of spark free users.”
Ron Gabrisko Aug 4, 2025 ▶ 23:21 20Sales: $0-$3.7BN: The Databricks CRO's Playbook to Build the Fastest GTM Engine in SaaS History | How Databricks Beat Snowflake | How To Build a Sales Org of 5,000 and Close $190M Deals with Ron Gabrisko
MAD Prediction Not checkable as stated
Rogojan: Enterprise migration away from Spark and Kafka will take years
“I think initially just picking up Kafka, Spark, those are still so heavily in use that even if they go out in the next three to five years, it's going to take time to migrate and things of that nature.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 22:06 Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
SAASTR Insight
Ghodsi: Bottom-up open-source adoption accelerates enterprise sales cycles
“With open source, you have people in the room that will raise their hand and say, I know what Spark is. I've already used it. I downloaded it on my laptop at home. I know how this Delta, I'm a big fan. I know what this is. So that moves things along much faste…”
Ali Ghodsi Dec 7, 2021 ▶ 7:46 The Future of AI, Open Source, and Enterprise SaaS with Databricks CEO Ali Ghodsi
SAASTR Insight
Ghodsi: Open-source companies must separate corporate brands from project names
“From early on, we said, look technology comes and goes. Spark in 10 years will have probably aged and hopefully we'll come up with other innovations, so let's pick a name that separates the company from the open source technology, and hopefully we can continue…”
Ali Ghodsi Dec 7, 2021 ▶ 10:28 The Future of AI, Open Source, and Enterprise SaaS with Databricks CEO Ali Ghodsi
MAD Assertion Not checkable as stated
Netflix uses S3, Spark, Presto, and Snowflake for data querying
“We use SG as so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake as our query engines.”
Savin Goyal Feb 17, 2021 ▶ 3:23 Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
MAD Opinion
Apache Spark failed to effectively shrink down to single-node scale
“One of the things that you find with things like Spark is that they really failed to shrink down and, Effectively do computing at the single node scale.”
Wes McKinney Feb 1, 2021 ▶ 24:22 Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
MAD Assertion Supported
Running Apache Spark on a single node is slower than Pandas
“You can use spark at the single node scale as an alternative to pandas through the koalas interface, but you'll find that for many workloads, it's simply slower than pandas, which is not super impressive.”
Wes McKinney Feb 1, 2021 ▶ 24:30 Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
a16z Disclosure
Stoica: Apache Spark was created to prove Mesos's value
“So actually, Mesos was the first project we developed. This is what started the stack. And Mesos was by design to support multiple cluster computing framework. We started with Hadoop. And actually, one of the reasons we designed Spark, it's To show that it's m…”
Ion Stoica Jan 2, 2019 ▶ 4:03 a16z Podcast | A New Lab Rises
a16z Assertion Supported
UC Berkeley's AMPLab Is the Birthplace of Apache Spark
“The place where Apache Spark was born, UC Berkeley's Amplab has not just created a major open source software platform, it's spun out more than its share of groundbreaking companies.”
Michael Copeland Jan 2, 2019 ▶ 0:05 a16z Podcast | AMPLab, the Power of Open Source, and the Future of Systems Software
a16z Assertion Supported
Zaharia: Toyota uses Spark to analyze social media feedback on cars
“One of the coolest ones I saw was a talk from Toyota about how they use Spark to improve, you know, to basically look at social media feedback, what people are writing about their cars, and figure out things like, oh, is there a problem with the brakes on the …”
Matei Zaharia Jan 2, 2019 ▶ 6:44 a16z Podcast | A Conversation With the Inventor of Spark
a16z Assertion Supported
Legacy Hadoop tools like Hive, Pig, and Mahout now run on Spark
“So in particular you know, one of the things we saw is many of the projects that were built on top of Hadoop, such as Hive, which is a SQL processing at scale and Pig and Mahout for machine learning are starting to run on top of Spark as well, so that users of…”
Matei Zaharia Jan 2, 2019 ▶ 13:59 a16z Podcast | A Conversation With the Inventor of Spark
a16z Opinion
Zaharia: Third-party integrations are Spark's most valuable asset for users
“So I think even beyond the activity happening in Spark itself, these projects on top and on the side are one of the most valuable things for the users.”
Matei Zaharia Jan 2, 2019 ▶ 14:44 a16z Podcast | A Conversation With the Inventor of Spark
a16z Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31 a16z Podcast | A Conversation With the Inventor of Spark
MAD Assertion Not checkable as stated
Early Apache Spark was so unstable that clusters crashed within two hours
“At the time we made that bet, Spark was barely working. I mean, seriously. It was a really promising technology, very complete programming model. We loved the fact it was in memory, but we couldn't find a single customer that could keep their cluster running f…”
Mike Tuchen Oct 17, 2018 ▶ 16:31 Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC)
a16z Assertion Partly supported
Jordan claims Bag of Little Bootstraps vastly outperforms traditional bootstrap
“Here's the new algorithm, you know, again, implement on Spark. It's that little red box there. It took about a couple hundred seconds to get it, And the answer quality is better than the bootstrap after 15,000 seconds.”
Michael Jordan Jul 28, 2017 ▶ 21:36 Michael Jordan
MAD Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Apache Spark originally ran on Mesos before adding YARN and standalone support
“As originally was built to run on top of Mesos. Today is working on Yarn, working, you know, standalone, and is working also in addition to HDFS, you know, imports and exports data to many other data sources.”
Ion Stoica Apr 2, 2015 ▶ 7:07 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Apache Spark supports all major data workloads with one engine
“While we spark, You can use only one engine and only one API to support all these workloads.”
Ion Stoica Apr 2, 2015 ▶ 12:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Apache Spark has exceeded 500 active open-source contributors
“We exceeded 500 contributors. It is the most active big data project right now, Spark.”
Ion Stoica Apr 2, 2015 ▶ 16:47 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Stoica: Databricks still provides the majority of Apache Spark open-source contributions
“A lot of contributions, still the majority of contributions comes from Databricks.”
Ion Stoica Apr 2, 2015 ▶ 17:00 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Disclosure
Stoica: Databricks offers only cloud services, not a custom Spark distribution
“We don't have our own distribution. We provide only the service.”
Ion Stoica Apr 2, 2015 ▶ 17:30 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Prediction Not checkable as stated
Stoica predicts big data infrastructure will eventually unify around one system
“Now, I do think that looking forward you are going to see more and more of this unification. This happens in many other industries and technologies. I do think this will happen in big data because it's so much easier if you have only one system than if you nee…”
Ion Stoica Apr 2, 2015 ▶ 21:53 Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
MAD Assertion Supported
Knaup: Apache Spark was the first application built on Apache Mesos
“Spark, ah, there's a little known fact, Spark was actually the first, ah, application that was ever built on top of Mesos, and it was sort of the, you know, the demo app for Mesos in the beginning.”
Tobi Knaup Sep 22, 2014 ▶ 5:42 Tobi Knaup, Mesosphere // Data Driven #29 // Sep 2014 (Hosted by FirstMark Capital)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.