Spark

product on 11 shows · 29 statements across 24 episodes · said 15 times in 13 episodes since 2014

TBPN 9 Mixergy 1 Latent Space 1 Sourcery 1 Top Founders 1 the a16z Podcast 1 20VC 1 the Knowledge Project the Startup Ideas Podcast the Green Blueprint the MAD Podcast

Mentions by year, every show

tap a year for its mentions
008815152014201520162017201820192020202120222023202420252026episodesmentions
08152014201520162017201820192020202120222023202420252026episodes it came up in
000.87.51.5152014201520162017201820192020202120222023202420252026episodesmentions per episode

TBPN 9Mixergy 1Top Founders 120VC 1the a16z Podcast 1Latent Space 1Sourcery 1

2026 13 mentions in 11 episodes 1 per episode
2017 1 mention in 1 episode
2014 1 mention in 1 episode

every mention on every show, scene by scene, with the transcript →

29 statements about Spark, every show

Xin: Single HTAP database engines compromise ecosystem compatibility and performance
“This is sort of the holy grail of database engineering is, why not build a single system that can do both of this? But it ends up just being a lot of compromises. And one, I think one of the first issue is that, hey, each, they say Postgres has a massive ecosy…”
Reynold Xin Jun 24, 2026 ▶ 32:30 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
GREEN BLUEPRINT Prediction Open · timeframe Jun 2028
Needham: SPARC reactor targets 100MW thermal power with 10x energy gain
“That'll produce about a hundred megawatts of thermal power. So a lot like 500 to a thousand times what the national ignition facility did, like a substantial amount of power. And his design point is actually a gain of like 10, so 10 times, actually it's 11 mor…”
Rick Needham Jun 18, 2025 ▶ 19:16 The tape that led to a fusion breakthrough
Needham: CFS ARC commercial plant will generate 400MW electric and 1GW thermal
“Spark is the, you know, the basis of our power plant, which we call ARC, and that'll be a 400 megawatt electric plant, a one gigawatt thermal plant”
Rick Needham Jun 18, 2025 ▶ 20:27 The tape that led to a fusion breakthrough
Ury: Early Romantic Chemistry Is Often Anxiety Triggered by Ambiguity
“The second myth is that if you feel the spark, it's definitely a good thing, and actually, sometimes certain people give us the spark because we actually don't know how they feel about us, so sort of that hot, cold feeling makes us wonder, does he like me? Doe…”
Logan Ury Mar 18, 2025 ▶ 5:01 Logan Ury: The Dating Myths You Need to Stop Believing
20VC Prediction Not checkable as stated
The VC Industry Faces a Reckoning With Many Firms Shutting Down
“I think there will be a lot of folks that go away in this cycle. I watched it in the last go round. You know, Spark and USV and all these firms didn't exist before the last cycle. And they really made their names coming out of the O eight crisis. And on the ot…”
Mo Koyfman Aug 8, 2022 ▶ 58:52 Mo Koyfman: The Secret to Winning in Venture; Why Small Funds Outperform Large Funds | 20VC #915 · 20VC with Harry Stebbings
Mason: Missive's comparison pages explicitly state product shortcomings to deter poor-fit users
“It's a honest take. When there are some shortcomings in Missive compared to the other product, we do state it clearly, and we don't want to bring users in who won't like Missive and would actually rather use Spark or, like, a more personal-oriented email app.”
Raphael Mason Jul 30, 2022 ▶ 15:24 I love these guys. 3 co-Founders bootstrap to $2m in ARR. Beautiful business.
MAD Assertion Not checkable as stated
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Tristan Handy Jun 28, 2022 ▶ 7:34 The Next Layer of the Modern Data Stack | dbt's Tristan Handy
MAD Disclosure
Ghodsi: Databricks re-wrote Apache Spark's execution engine in C++
“Two or three years ago, we set out to re-implement all Spark in C++ in what we call the really, really fast, what's called MPP engine, Massive Apparel Processing Engine.”
Ali Ghodsi May 24, 2021 ▶ 23:12 Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
MAD Opinion
Borgman: Apache Spark is not built to support high concurrency workloads
“Anyone who's trying to do high concurrency will usually realize that Spark is not built for high concurrency.”
Justin Borgman Mar 19, 2019 ▶ 18:41 Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
a16z Assertion Supported
Hirasaki: Spark acquisition vaulted Roche forward in gene therapy platform technology
“And then with the Spark acquisition same thing BMS, I'm sorry Roche was a laggard in the space, and now suddenly leaped ahead with not only a marketed gene therapy, but this AAV platform that it could use for delivering other gene therapies.”
Koki Hirasaki Feb 28, 2019 ▶ 5:59 What’s with All the Bio M&A in 2019?: A Quick Take
a16z Assertion Supported
Hirasaki: Roche Allows Acquired Companies Like Spark to Operate Semi-Autonomously
“Some big pharma companies will just tend to fully integrate the company whereas others will like Roche and Spark will allow the acquired companies to operate semi-autonomously.”
Koki Hirasaki Feb 28, 2019 ▶ 16:10 What’s with All the Bio M&A in 2019?: A Quick Take
a16z Opinion
Leibert: Writing Visual Basic is more joyful than writing Go applications
“And frankly, it's more joyful to write Visual Basic than it is to write Go, right? That you actually achieve Business results. I mean, seriously, right? Like, writing a Spark job, you see the output right away, whereas if you write a large Go application, I me…”
Florian Leibert Jan 2, 2019 ▶ 2:21 a16z Podcast | Containing the Monolith -- From Microservices to DevOps
a16z Insight
Stanek: Business users prefer spreadsheet interfaces over Spark and Hadoop
“Some of the most frequently used kind of data analytics tools extremely basic, because they actually look and feel like sheet of paper, like two-dimensional sheet of paper, and you know, so that's the problem with analytics, that on one hand, we have, you know…”
Roman Stanek Jan 2, 2019 ▶ 20:55 a16z Podcast | Making the Most of the Data That Matters
a16z Assertion Not checkable as stated
Moghe: Spark and Hadoop do not replace existing data warehouses
“Spark doesn't subsume data warehousing. Hadoop doesn't subsume, you know, streaming. So they're just like different technologies for different jobs.”
Prat Moghe Jan 2, 2019 ▶ 23:39 a16z Podcast | Making the Most of the Data That Matters
a16z Assertion Not checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei Zaharia Jan 2, 2019 ▶ 0:29 a16z Podcast | A Conversation With the Inventor of Spark
a16z Assertion Not publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei Zaharia Jan 2, 2019 ▶ 9:15 a16z Podcast | A Conversation With the Inventor of Spark
a16z Disclosure
Zaharia: Spark was originally designed to run Netflix Prize recommendation algorithms
“So it's actually one of the applications that I first tried to support in Spark was you know, the recommendation algorithm he was working on.”
Matei Zaharia Jan 2, 2019 ▶ 16:41 a16z Podcast | A Conversation With the Inventor of Spark
a16z What-if
Nguyen: Apache Spark Would Have Failed Earlier Due to Memory Costs
“Now Spark, if it was created six, five, six years before its time would have completely failed because memory was so much more expensive.”
Christopher Nguyen Jan 2, 2019 ▶ 14:45 a16z Podcast | Making Sense of Big Data, Machine Learning, and Deep Learning
MAD Opinion
Bob Muglia: Apache Spark scenarios are complementary to Snowflake
“Spark, I think, is being used very, very broadly for advanced analytics, machine learning, in some cases for streaming data, and those scenarios are all very, very complimentary to Snowflake.”
Bob Muglia Apr 9, 2018 ▶ 18:28 Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven)
a16z Assertion Supported
Analytics frameworks like Spark and Hadoop assume exclusive resource access
“Most things, Hadoop, Spark, Storm, they think they're running by themselves. And so they compete for resources in really interesting ways.”
Chandra Krintz Jul 15, 2017 ▶ 11:52 Chandra Krintz
MAD Assertion Not checkable as stated
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Sean Kandel Jul 13, 2017 ▶ 6:41 Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven)
TOP FOUNDERS Assertion Supported
Particle's first Kickstarter failed after raising $125K on a $250K goal
“We had a 250,000 dollar goal. We raised 125,000, but Kickstarter is all or nothing. So that means we raised zero and that product never came to be.”
Zach Supalla Mar 6, 2017 ▶ 10:10 EP 590 :Particle.io Raises $14M, Passes $5M In Revenue, Helping Usher in IoT Connecting Keurigs to Internet with CEO Zach Supalla
MAD Opinion
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 15:56 Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]
MAD Assertion Not checkable as stated
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Praveen Murugesan Sep 30, 2016 ▶ 10:21 The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
20VC Assertion Not checkable as stated
Turck: Spark Made Real-Time Massive Data Processing Possible
“What those guys are doing would simply not have been possible to do up until, you know, two or three years ago when the emergence of frameworks like Spark made possible, what was an alibi, it became possible to be processing You know, there's massive amounts o…”
Matt Turck May 30, 2016 ▶ 10:53 20VC: Is Big Data Still A Thing with Matt Turck, Managing Director at FirstMark Capital
MAD Disclosure
AXA US builds its analytics infrastructure on a Cloudera Hadoop stack
“So the one that we've got here in the U.S., it's primarily a Cloudera Hadoop stack that we've used a blueprint that was essentially blessed by our brethren over in French, in France and with that, we've got R and Python and Spark. We use some Dataiku along the…”
Louis DiModugno May 23, 2016 ▶ 10:44 Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven]
MAD Assertion Not checkable as stated
Barclays accelerated Spark jobs from hours to seconds using Alluxio
“They used Tachyon to accelerate Spark jobs from hours to seconds”
Haoyuan Li Apr 13, 2016 ▶ 12:03 A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
MAD Assertion Not checkable as stated
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Stefan Groschupf Dec 17, 2015 ▶ 10:35 The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer
MAD Disclosure
Essas: OpenTable routes event streams through Kafka, Cassandra, and Spark
“All of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data.”
Joseph Essas Jun 19, 2015 ▶ 2:32 Joseph Essas, OpenTable // Mining Diner Talk (Hosted by FirstMark Capital)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.