Apache Spark, every mention
195 scenes across 14 shows · ← back to Apache Spark
the MAD Podcast 215
the a16z Podcast 104
the Official SaaStr Podcast 19
Latent Space 13
20VC 10
Lenny's Podcast 5
the Neon Show 4
No Priors 26 more shows
every year every show
the MAD Podcast 215
the a16z Podcast 104
the Official SaaStr Podcast 19
Latent Space 13
20VC 10
Lenny's Podcast 5
the Neon Show 4
No Priors 2
the Y Combinator Startup Podcast 1
the Knowledge Project 1
A Product Market Fit Show 1
Top Founders 1
Big Technology 1
TBPN 1
Verbatim, from the transcripts: passages where Apache Spark comes up on the MAD Podcast, the a16z Podcast, the Official SaaStr Podcast, Latent Space, 20VC
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
- ▶ 1:34 Matei Zaharia We were doing these tutorials and, yeah, just teach people Spark.
- ▶ 11:17 Shawn Wang And I mean, this goes back to Spark, right? 2 times in the scene
- ▶ 26:17 Reynold Xin Probably very similar to how you did Spark, which is you could just start using. 2 times in the scene
- ▶ 32:49 Reynold Xin And Spark, for example, had a massive ecosystem.
- ▶ 55:43 Matei Zaharia Cause that's what Spark was, was for the large scale map reduce like stuff.
Is Unstructured Data The Key To Successful AI Deployments?
- ▶ 9:36 Jitesh Gai Um, there was this thing called Hadoop and Spark,
Why "Boring" Infrastructure is the Best Path to a $60B Company | Manish Jindal, Cloudflare & Arize
- ▶ 12:06 Manish Jindal I mean, typically, historically, and same if you look at the Databricks journey, uh, they've been now around since, you know, from the Spark days, and now they've been around 11 years, uh, and it took them a while to get to where they are.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 31:43 unnamed speaker I can run it on a spark.
AI Needs to Know Why You took THAT decision | Ashu Garg, Investor at Foundation Capital
[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
- ▶ 16:29 Andy Konwinski We can actually identify projects that are more likely to become a Databricks or an Apache Spark or Array or an LM Arena sooner and, and, and with more confidence.
He killed a $100K ARR product & pivoted—then raised $375M. | Viraj Parekh, Co-Founder of Astronomer · PMF Show
- ▶ 17:53 Viraj Parekh Databricks is, I can't even remember how big the last round of funding they raised was, but that's built around Apache Spark.
Ben Horowitz and Ali Ghodsi: How to Run a $100 Billion Business
- ▶ 1:30 Ali Ghodsi Apache Spark became a worldwide sensation, and we could pride ourselves on the number of downloads, uh, of the software. 6 times in the scene
- ▶ 58:49 Ali Ghodsi We created Apache Spark and we made it a worldwide sensation.
20Sales: Scaling Snowflake from $0-$3BN in ARR | Snowflake vs Databricks: My Biggest Lessons | Why Customer Success is BS and What Replaces It with Chris Chris Degnan
- ▶ 35:02 Chris Degnan I would just would have hired more salespeople and focused on new logo acquisition and pounded on product to give me a notebook and a spark connector, and we would have kicked the shit out of Databricks and they would be dead. 2 times in the scene
$46B of hard truths: Why founders fail and why you need to run toward fear | Ben Horowitz (a16z)
- ▶ 23:19 Ben Horowitz I knew, you know, like I knew at the time that, uh, what they had was this thing called spark and they had, you know, the competitor, uh, was something called Hadoop and Hadoop, you know, had very well funded companies already running… 2 times in the scene
- ▶ 39:47 Ben Horowitz If you're going to catch these guys before they take spark and like use it against you.
How 80,000 companies build with AI: Products as organisms and the death of org charts | Asha Sharma
- ▶ 27:17 Asha Sharma We use it to, uh, we use Spark to create prototypes.
20Sales: $0-$3.7BN: The Databricks CRO's Playbook to Build the Fastest GTM Engine in SaaS History | How Databricks Beat Snowflake | How To Build a Sales Org of 5,000 and Close $190M Deals with Ron Gabrisko
The Future of Software Development - Vibe Coding, Prompt Engineering & AI Assistants
- ▶ 21:32 Matt Bornstein One is kind of this, like, kind of backend data eng driven big data systems, you know, Spark, Kadoop sort of thing.
- ▶ 24:51 Jennifer Li If it's, you know, YubiKeys or if it's, um, again, like Databricks on Spark.
Weekly Recap: CEO Affair at Coldplay Concert, Elon Musk's AI Girlfriend, Cognition Acquires Windsurf
- ▶ 9:57 John Coogan And that's kind of what, uh, Databricks did with spark, like sparks and open source project.
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
- ▶ 43:30 Ion Stoica And then we, we, we done here, then it comes Spark, right?
- ▶ 50:31 Ion Stoica Like, um, you know, like Databricks with Spark and, uh, or AnyScale with Ray.
How Palantir built the ultimate founder factory | Nabeel S. Qureshi (founder, writer, ex-Palantir)
- ▶ 28:42 Nabeel S. Qureshi You'd have to debug, you know, spark errors or whatever it was, but basically that process,
Why is everyone cloning Deep Research?
- ▶ 47:55 Shawn Wang Data science would be, like, this is a Spark job or, you know, it's like a Wraith.
From $1M to $3B ARR: Databricks CRO Ron Gabrisko on Scaling a Revenue Rocket Ship
- ▶ 8:31 Ron Gabrisko Databricks founders created spark, right? 4 times in the scene
- ▶ 12:21 Ron Gabrisko If you look at the strategy of the company, it was like, make Databricks the best place to manage Spark. 2 times in the scene
Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy
- ▶ 20:58 Matt Turck Spark or Kafka or like any of those frameworks, what would you recommend next once I have my language, I have my SQL, uh, what do I do next? 3 times in the scene
- ▶ 27:12 Ben Rogojan (Seattle Data Guy) Like there, there are cases where maybe they don't know what a data warehouse is, or, you know, when you say data pipeline, maybe they, they don't understand, or when you reference Spark, they're like, I, you know, why do I care what Spark…
- ▶ 40:31 Ben Rogojan (Seattle Data Guy) And then from there, you could just say like, I'm going to use Spark.
AI at ZoomInfo: Superpowering GTM teams | Ali Dasdan, CTO, ZoomInfo
- ▶ 16:26 Ali Dasdan All these connections are there that we are extracting out of that, and that is done through, you know, different technologies, you know, either spark code or data flow that GCP has, and all kinds of basic processing.
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 1:04 Jonathan Frankle I, as much as I would love to understand how Spark works,
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Vector databases and the $8 trillion open source market | Bob van Luijt, CEO of Weaviate
- ▶ 18:54 Bob van Luijt The, on a serious note, why this becomes so interesting, at some point we saw, like, it started for Weaviate in an uptick that we saw, so community contributions to, um, our Spark connector.
No Priors Ep. 11 | With Matei Zaharia, CTO of Databricks
- ▶ 0:59 Matei Zaharia We also started open source projects like most notably Apache Spark, which, you know, was essentially, you know, the first version of it was my PhD thesis.
- ▶ 25:58 Elad Gil When, when you were working on, um, Spark for, for, uh, your PhD, did you think you'd become a founder?
Build Fast APIs Faster Over Data at Scale | Tinybird Founder & CEO Jorge Gomez Sancha
- ▶ 23:55 Jorge Gomez Sancha Um, could be in a data warehouse, or it could be in spark or something like that.
Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota
- ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system.
Fundamentals of Data Engineering | Joe Reis and Matt Housley
- ▶ 2:40 Matt Housley Yeah, exactly, and, and I'll kind of skip a bullet point and then go back, but like this one right here, what we kept hearing a lot is that, you know, data engineering is Spark, or data engineering is Kafka.
- ▶ 14:38 Matt Housley Or, or we get, like, things like, well, it's really all about Spark. 2 times in the scene
- ▶ 30:50 Matt Housley Um, I, I think when we, I, I think, and correct me if I misunderstood the question, but I think when we talk about not defining, setting definitions around technology, we mean specifically not saying that data engineering is about Spark,… 3 times in the scene
A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
- ▶ 2:56 Gleb Mezhanskiy For example, your Airflow orchestrator scheduler is broken, or your cluster, like Spark cluster is underwater and backlogged, or your vendor that you use to buy data ships to something which is, ah, of low quality.
The Next Layer of the Modern Data Stack | dbt's Tristan Handy
- ▶ 7:34 Tristan Handy Um, and that's, it's, again, I, sometimes, like, people get defensive, the data engineers in the audience, this is not a diatribe against data engineers, it's just that there are actually two orders of magnitude more human beings on the…
Lessons Learned Building a $2 Billion Company from Scratch with Neo4j CEO & Co-Founder Emil Eifrem
- ▶ 13:16 Emil Eifrem And then over time, we've invested in the ecosystem to create more on-ramps onto NeoPJ from other platforms like Spark or Confluent, you know, Apache Kafka.
The Future of AI, Open Source, and Enterprise SaaS with Databricks CEO Ali Ghodsi
- ▶ 3:29 Nithya Ruff And I'm so glad to speak to Databricks because Databricks not only started out, um, in the AMP labs of, uh, UC Berkeley, um, based on the Apache Spark, but they continue to innovate, continue to open source their projects, uh, like Delta…
- ▶ 4:44 Ali Ghodsi Um, at that time, Spark wasn't very well known. 3 times in the scene
- ▶ 9:58 Ali Ghodsi You know, we were engineers, so we actually made a long list of different names we would pick, and we scored them, and, you know, we did a bunch of research, and, ah, but one of the things was, um, many of the, or several of the founders… 5 times in the scene
- ▶ 11:51 Nithya Ruff So another thing you did right, right from the start was, um, you went with a managed service or, you know, Spark as SAS.
- ▶ 19:17 Ali Ghodsi We, we hosted our own spark service in the cloud and we competed well with the cloud vendors.
- ▶ 23:11 Ali Ghodsi So when we started with Spark, that's just a way to get the data.
Top 10 Trends in AI, Machine Learning and Data for 2022
- ▶ 7:43 Matt Turck And then the spark came along and that was like another whole thing that took several years.
Fireside Chat: Zhamak Dehghani (Founder, Data Mesh) with Matt Turck (Partner, FirstMark)
- ▶ 19:01 Zhamak Dehghani So a lot of people still use a spark or beam or, you know, whatever.
Fireside Chat: Abe Gong (Founder & CEO, Superconductive) with Matt Turck (Partner, FirstMark)
- ▶ 11:50 Abe Gong Even in other places, like within, um, typed data frames in Spark or, you know, your choice of data warehouses, being able to do sets, ranges, regular expressions, uh, distributions, uh, correlations, uh, things like that start to be…
- ▶ 17:14 Abe Gong Uh, we also do Spark data frames, um, and then SQL, uh, through SQL alchemy, uh, which also implies a whole slew of different SQL dialects.
Fireside Chat: Nick Schrock (Founder & CEO, Elementl) with Matt Turck (Partner, FirstMark)
- ▶ 4:38 Nick Schrock You know, Hadoop and then spark and now the cloud data warehouse.
- ▶ 26:29 Nick Schrock And then actually, you know, I've been really impressed with the development of Spark over the last few years.
Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
- ▶ 0:18 Matt Turck So Amplabs, Spark, and Databricks, how did it all start?
- ▶ 6:30 Matt Turck So to close on, on, on that, um, you know, chapter of the early years, um, how did you go from this academic, uh, very popular open source project, uh, which was Spark to 2 times in the scene
- ▶ 11:17 Matt Turck Uh, yeah, I had the, um, uh, pleasure and honor of, like, hosting your co-founder and CEO at the time, Stoica in 2015, and the conversation, I rewatched it before this, and the conversation at the time was all about, you know, the, the,… 4 times in the scene
- ▶ 22:58 Ali Ghodsi So in the past, when someone wanted to do SQL or warehousing on Databricks, we would offer them Spark. 4 times in the scene
- ▶ 26:03 Ali Ghodsi When we had spark and the founders were discussing, what should the name of the company be? 4 times in the scene
- ▶ 28:48 Ali Ghodsi Of course, some of these projects, when they get older, like spark, they move into the maintenance side. 2 times in the scene
- ▶ 33:25 Ali Ghodsi And if we were just doing spark, like we were on this show in, that would have been probably 10, five percent because, you know, over time, these technologies become mature and, you know, the excitement around them, uh, wanes.
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 5:58 Dave Burgess And so we query Intelli, uh, to Presto, to Hive, uh, to Spark SQL, uh, to MySQL, but you can do it with other engines too.
- ▶ 7:56 Dave Burgess So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. 2 times in the scene
- ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.
Fireside Chat: Bindu Reddy (Founder & CEO, Abacus.AI) with Matt Turck (Partner, FirstMark)
- ▶ 12:02 Bindu Reddy Uh, Kubernetes, um, you know, Spark, um, Redis.
- ▶ 28:26 Bindu Reddy The other thing I think from a data science perspective, the thing which has really, really born the test of time has been Spark. 2 times in the scene
Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
- ▶ 3:26 Savin Goyal Uh, so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake, uh, as our query engines. 2 times in the scene
Introducing Kedro
- ▶ 11:59 Kedro Product Manager Um, the data catalog supports multiple integrations of pandas, spark, desk,
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 7:35 Julien Le Dem You know, so I talked to Wes, uh, obviously, uh, we like spark contributors, um, GBT and all the, the air flow and all the very, um, popular frameworks to schedule and process data. 6 times in the scene
- ▶ 22:42 Julien Le Dem I think right now, today, Spark is one of the big projects that people are using. 3 times in the scene
Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
- ▶ 10:31 Matt Turck And then on the other side, you have, uh, the world, uh, of the machine learning and data analysis tools, which is like Spark and NumPy and, and so on and so forth.
- ▶ 21:50 Wes McKinney Spark, uh, Spark supports Arrow as a, as an interchange format, and it's used heavily in the interface with Python and R, for example. 2 times in the scene
- ▶ 22:42 Matt Turck After HPC, um, slash Hadoop, slash Spark, slash Ray, what's the long-term future of parallel compute for data intensive workflows? 3 times in the scene
Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark)
- ▶ 7:24 Alok Gupta We use Python and Databricks to, uh, and Spark to pull data in, build models. 2 times in the scene
20VC: Databricks CEO, Ali Ghodsi on The 3 Phases of Startup Growth, How to Evaluate Risk and Downside Scenario Planning & Who, What and When To Hire When Scaling Your Go-To-Market
- ▶ 0:31 Harry Stebbings And prior to Databricks, Ali was one of the original creators of open source project, Apache Spark, and ideas from his research have been applied to Apache Mesos and Apache Hadoop.
How to Build an Open Source Business
- ▶ 13:57 Peter Levine Databricks and Spark is an example.
Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Fireside Chat: Solmaz Shahalizadeh, VP of Data Science & Engineering at Shopify (Data Driven NYC)
- ▶ 8:37 Solmaz Shahalizadeh So, um, we basically moved from, uh, using Vertica, uh, to building a ETL tool in-house using, uh, Spark, uh, and, uh, Python. 4 times in the scene
Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
- ▶ 9:35 Justin Borgman That means, you know, Hadoop's various projects can read this data, Spark can read this data, and of course, Presto can read this data as well, um, but because it's an open file format,
- ▶ 17:54 unnamed speaker I was wondering, given that Spark is also separated from storage, um, how does Starburst stay dependent with companies like Databricks? 4 times in the scene
a16z Podcast | Containing the Monolith -- From Microservices to DevOps
- ▶ 1:51 Florian Leibert But of course, also data scientists will install notebooks, they'll install Spark, they'll install Hadoop, and then use it. 2 times in the scene
a16z Podcast | AI, from 'Toy' Problems to Practical Application
- ▶ 16:18 Joe Spisak There are a lot of big data SIs, and they've been using Hadoop and Spark, and they've been dabbling in machine learning and advanced analytics.
a16z Podcast | A New Lab Rises
- ▶ 2:31 Peter Levine One, of course, is Spark, which is the basis of Databricks. 4 times in the scene
- ▶ 11:19 Sonal Chokshi Like the precursor to Hadoop and then Spark.
- ▶ 22:16 Ion Stoica There are projects which hope to become a strong artifact to be used across in, in, in industry like Spark or Mesos or, uh, Tachyon.
- ▶ 26:16 Ion Stoica It's about, it's actually, it's also related to Spark.
a16z Podcast | The Changing Culture of Open Source
- ▶ 26:18 Sonal Chokshi Does this then leave the domain of, like, the really big visionary projects that come out of academia, like the way Spark came out of Amplab, Apache Spark, where you have an entire new set of companies being built?
a16z Podcast | The Storage Renaissance
- ▶ 0:28 Sonal Chokshi Which came out of the UC Berkeley Amp Lab, the birthplace of other industry-defining technologies such as Spark and Mesos.
- ▶ 13:51 Mike Matchett So Spark and in-memory approaches to machine learning really accelerate the opportunity to create and apply machine learning algorithms to just about every facet of human existence, not to overstate the case, but, uh, there really is a…
a16z Podcast | The Product Edge in Machine Learning Startups
- ▶ 12:04 AJ Shankar Spark, I think, is the go-toe one for now. 2 times in the scene
a16z Podcast | Software Programs the World
- ▶ 5:26 Marc Andreessen So for example, we've seen the rise of, in that category, we've seen the rise of Hadoop, and now the rise of Spark for distributed data processing.
a16z Podcast | Selling to Developers & Open Source Business Models
- ▶ 17:30 Peter Levine You know, one of the companies that were invested in, Arimo, formerly Adetow, builds machine learning and big predictive big data applications that sit on top of Spark and Hadoop installations. 2 times in the scene
a16z Podcast | Making the Most of the Data That Matters
- ▶ 17:46 Steven Sinofsky The University of California, Berkeley, well, has a, a whole variety of some of the leading technologies that like Spark has come out of there.
- ▶ 21:05 Roman Stanek Uh, extremely basic, because they actually look and feel like sheet of paper, like two-dimensional sheet of paper, and, uh, you know, so that's, that's the problem with analytics, that on one hand, we have, you know, very complex systems… 2 times in the scene