Jan 2, 2019 · 19m · a16z

a16z Podcast | A Conversation With the Inventor of Spark

Matei Zaharia · 13m spoken Sonal Chokshi · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, host Sonal speaks with Apache Spark inventor and Databricks CTO Matei Zaharia about Spark's origins at UC Berkeley, its advantages over legacy tools like MapReduce, and its rapid growth across open-source communities and enterprise cloud analytics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 24.4% of the talking time here. How this is scored →

The host as informed peer 4.0 Guest teaching 4.2 Guest disagreement 1.6 The host pushing back 3.8
05100:0010:000:56–4:40 · The host as informed peer 4/10 Evolution from MapReduce and Facebook Data Challenges Sonal asks about MapReduce limitations and Facebook's data growth, attempting to frame Spark's value around rapid feature testing. Matei clarifies that Spark was built for iterative, ad-hoc data exploration rather than fast deployment cycles.4:40–8:38 · The host as informed peer 5/10 The Role of Ease of Use in Data Accessibility Sonal pushes back on whether interface usability matters for technical data experts and whether IBM's backing is purely a cloud bet. Matei politely corrects these assumptions, highlighting the need for non-expert data access and IBM's broader enterprise footprint.8:38–13:30 · The host as informed peer 4/10 Balancing Open Source Community and Corporate Interests Sonal probes the dynamics of managing large open-source projects and asks for specific factors behind Spark's growth. Matei details the importance of low contribution barriers, testing infrastructure, and community support.13:30–17:04 · The host as informed peer 3/10 The Growing Open Source Ecosystem Around Spark Sonal explores the broader software ecosystem surrounding Spark and steers the conversation toward Matei's background as an inventor. She interrupts playfully during the Netflix Challenge anecdote to ask whether the second-place team received a cash reward.17:04–19:03 · The host as informed peer 4/10 Commercializing Spark and Databricks' Cloud Business Model Sonal articulates the historical conflict between open-source software and corporate commercialization models. Matei explains how Databricks avoids this tension by delivering Spark as a managed cloud service.0:56–4:40 · Guest teaching 4/10 Evolution from MapReduce and Facebook Data Challenges Sonal asks about MapReduce limitations and Facebook's data growth, attempting to frame Spark's value around rapid feature testing. Matei clarifies that Spark was built for iterative, ad-hoc data exploration rather than fast deployment cycles.4:40–8:38 · Guest teaching 6/10 The Role of Ease of Use in Data Accessibility Sonal pushes back on whether interface usability matters for technical data experts and whether IBM's backing is purely a cloud bet. Matei politely corrects these assumptions, highlighting the need for non-expert data access and IBM's broader enterprise footprint.8:38–13:30 · Guest teaching 4/10 Balancing Open Source Community and Corporate Interests Sonal probes the dynamics of managing large open-source projects and asks for specific factors behind Spark's growth. Matei details the importance of low contribution barriers, testing infrastructure, and community support.13:30–17:04 · Guest teaching 3/10 The Growing Open Source Ecosystem Around Spark Sonal explores the broader software ecosystem surrounding Spark and steers the conversation toward Matei's background as an inventor. She interrupts playfully during the Netflix Challenge anecdote to ask whether the second-place team received a cash reward.17:04–19:03 · Guest teaching 4/10 Commercializing Spark and Databricks' Cloud Business Model Sonal articulates the historical conflict between open-source software and corporate commercialization models. Matei explains how Databricks avoids this tension by delivering Spark as a managed cloud service.0:56–4:40 · Guest disagreement 2/10 Evolution from MapReduce and Facebook Data Challenges Sonal asks about MapReduce limitations and Facebook's data growth, attempting to frame Spark's value around rapid feature testing. Matei clarifies that Spark was built for iterative, ad-hoc data exploration rather than fast deployment cycles.4:40–8:38 · Guest disagreement 3/10 The Role of Ease of Use in Data Accessibility Sonal pushes back on whether interface usability matters for technical data experts and whether IBM's backing is purely a cloud bet. Matei politely corrects these assumptions, highlighting the need for non-expert data access and IBM's broader enterprise footprint.8:38–13:30 · Guest disagreement 1/10 Balancing Open Source Community and Corporate Interests Sonal probes the dynamics of managing large open-source projects and asks for specific factors behind Spark's growth. Matei details the importance of low contribution barriers, testing infrastructure, and community support.13:30–17:04 · Guest disagreement 1/10 The Growing Open Source Ecosystem Around Spark Sonal explores the broader software ecosystem surrounding Spark and steers the conversation toward Matei's background as an inventor. She interrupts playfully during the Netflix Challenge anecdote to ask whether the second-place team received a cash reward.17:04–19:03 · Guest disagreement 1/10 Commercializing Spark and Databricks' Cloud Business Model Sonal articulates the historical conflict between open-source software and corporate commercialization models. Matei explains how Databricks avoids this tension by delivering Spark as a managed cloud service.0:56–4:40 · The host pushing back 3/10 Evolution from MapReduce and Facebook Data Challenges Sonal asks about MapReduce limitations and Facebook's data growth, attempting to frame Spark's value around rapid feature testing. Matei clarifies that Spark was built for iterative, ad-hoc data exploration rather than fast deployment cycles.4:40–8:38 · The host pushing back 5/10 The Role of Ease of Use in Data Accessibility Sonal pushes back on whether interface usability matters for technical data experts and whether IBM's backing is purely a cloud bet. Matei politely corrects these assumptions, highlighting the need for non-expert data access and IBM's broader enterprise footprint.8:38–13:30 · The host pushing back 4/10 Balancing Open Source Community and Corporate Interests Sonal probes the dynamics of managing large open-source projects and asks for specific factors behind Spark's growth. Matei details the importance of low contribution barriers, testing infrastructure, and community support.13:30–17:04 · The host pushing back 3/10 The Growing Open Source Ecosystem Around Spark Sonal explores the broader software ecosystem surrounding Spark and steers the conversation toward Matei's background as an inventor. She interrupts playfully during the Netflix Challenge anecdote to ask whether the second-place team received a cash reward.17:04–19:03 · The host pushing back 4/10 Commercializing Spark and Databricks' Cloud Business Model Sonal articulates the historical conflict between open-source software and corporate commercialization models. Matei explains how Databricks avoids this tension by delivering Spark as a managed cloud service.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 33.8% · guest 66.2%0:00 · the host 33.8% · guest 66.2%3:00 · the host 21.1% · guest 78.9%3:00 · the host 21.1% · guest 78.9%6:00 · the host 30.4% · guest 69.6%6:00 · the host 30.4% · guest 69.6%9:00 · the host 23.2% · guest 76.8%9:00 · the host 23.2% · guest 76.8%12:00 · the host 15.2% · guest 84.8%12:00 · the host 15.2% · guest 84.8%15:00 · the host 23.5% · guest 76.5%15:00 · the host 23.5% · guest 76.5%18:00 · the host 21.2% · guest 78.8%18:00 · the host 21.2% · guest 78.8%
Sharpest disagreement ▶ 4:40 Rejecting host premise on insider data needs

Matei directly rejects Sonal's suggestion that technical experts don't need easy tools by pointing out that experts want non-experts to handle their own queries.

Hardest push from the host ▶ 4:40 Challenging the necessity of usability

Sonal challenges the core pitch of Spark's ease-of-use by asking whether technical power users actually care about simple interfaces.

Biggest teaching moment ▶ 7:02 Correcting host misinterpretation of real-time processing

Matei corrects Sonal's assumption that Toyota processes social media in real-time, explaining it is used for deep offline clustering of qualitative issues.

The host holds their own ▶ 17:04 Framing open-source commercialization challenges

Sonal displays deep industry knowledge by highlighting the historical business tension between open-source community value and corporate monetization.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Evolution from MapReduce and Facebook Data Challenges 4423 Sonal asks about MapReduce limitations and Facebook's data growth, attempting to frame Spark's value around rapid feature testing. Matei clarifies that Spark was built for iterative, ad-hoc data exploration rather than fast deployment cycles.
The Role of Ease of Use in Data Accessibility 5635 Sonal pushes back on whether interface usability matters for technical data experts and whether IBM's backing is purely a cloud bet. Matei politely corrects these assumptions, highlighting the need for non-expert data access and IBM's broader enterprise footprint.
Balancing Open Source Community and Corporate Interests 4414 Sonal probes the dynamics of managing large open-source projects and asks for specific factors behind Spark's growth. Matei details the importance of low contribution barriers, testing infrastructure, and community support.
The Growing Open Source Ecosystem Around Spark 3313 Sonal explores the broader software ecosystem surrounding Spark and steers the conversation toward Matei's background as an inventor. She interrupts playfully during the Netflix Challenge anecdote to ask whether the second-place team received a cash reward.
Commercializing Spark and Databricks' Cloud Business Model 4414 Sonal articulates the historical conflict between open-source software and corporate commercialization models. Matei explains how Databricks avoids this tension by delivering Spark as a managed cloud service.

Statements from this episode (13)

Assertion Not checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei Zaharia Jan 2, 2019 ▶ 0:29
Disclosure
Matei Zaharia interned at Facebook in 2007 when it had 300 employees
“I was a PhD student at UC Berkeley, and we actually started working with Hadoop users back in 2007. And I did, for example, an internship at Facebook when Facebook was only about 300 people and they were just starting to set up Hadoop.”
Matei Zaharia Jan 2, 2019 ▶ 1:44
Assertion Supported
Zaharia: MapReduce was created by Google for nightly web indexing
“MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web.”
Matei Zaharia Jan 2, 2019 ▶ 3:49
Insight
Zaharia: Ad hoc data work requires iterative processing over batch runs
“When you work with data, you want to ask multiple questions repeatedly when you're doing ad hoc exploration of the data, as opposed to, you know, when you have a certain application that, you know, okay, I'm just going to run this every night.”
Matei Zaharia Jan 2, 2019 ▶ 4:27
Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei Zaharia Jan 2, 2019 ▶ 4:55
Assertion Supported
Zaharia: Toyota uses Spark to analyze social media feedback on cars
“One of the coolest ones I saw was a talk from Toyota about how they use Spark to improve, you know, to basically look at social media feedback, what people are writing about their cars, and figure out things like, oh, is there a problem with the brakes on the …”
Matei Zaharia Jan 2, 2019 ▶ 6:44
Assertion Not publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei Zaharia Jan 2, 2019 ▶ 9:15
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei Zaharia Jan 2, 2019 ▶ 12:09
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei Zaharia Jan 2, 2019 ▶ 13:10
Assertion Supported
Legacy Hadoop tools like Hive, Pig, and Mahout now run on Spark
“So in particular you know, one of the things we saw is many of the projects that were built on top of Hadoop, such as Hive, which is a SQL processing at scale and Pig and Mahout for machine learning are starting to run on top of Spark as well, so that users of…”
Matei Zaharia Jan 2, 2019 ▶ 13:59
Opinion
Zaharia: Third-party integrations are Spark's most valuable asset for users
“So I think even beyond the activity happening in Spark itself, these projects on top and on the side are one of the most valuable things for the users.”
Matei Zaharia Jan 2, 2019 ▶ 14:44
Disclosure
Zaharia: Spark was originally designed to run Netflix Prize recommendation algorithms
“So it's actually one of the applications that I first tried to support in Spark was you know, the recommendation algorithm he was working on.”
Matei Zaharia Jan 2, 2019 ▶ 16:41
Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.