Apr 2, 2015 · 27m · mad

Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)

Ion Stoica · 19m spoken Matt Turck · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this DataDrivenNYC session hosted by Matt Turck, computer scientist and Databricks co-founder Ion Stoica discusses the creation of Apache Spark, its architectural advantages over traditional Hadoop MapReduce, and Databricks' enterprise cloud platform.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.2% of the talking time here. How this is scored →

Matt as informed peer 1.8 Guest teaching 3.6 Guest disagreement 0.3 Matt pushing back 0.3
05100:0010:0020:001:12–3:42 · Matt as informed peer 1/10 Ion Stoica's Academic Background and Conviva Matt introduces Ion Stoica and sets up the narrative of his transition from academia to entrepreneurship. Ion monologues on his background at CMU and Berkeley, recounting how building Conviva revealed early real-time data processing bottlenecks with Hadoop.3:42–6:08 · Matt as informed peer 0/10 The Origin Story of Apache Spark at UC Berkeley In a continuous monologue without host intervention, Ion details the genesis of Spark at UC Berkeley's RAD Lab. He explains why iterative machine learning algorithms were prohibitively slow on Hadoop MapReduce due to HDFS write and read overhead between iterations.6:08–9:36 · Matt as informed peer 4/10 Comparing Apache Spark and the Hadoop Stack Architecture Matt demonstrates technical understanding by asking whether Spark acts as a replacement processing engine for MapReduce over HDFS. Ion builds on this, outlining the three-layer Hadoop stack and detailing Spark's three core advantages: speed, rich API operators, and workload unification.9:36–12:17 · Matt as informed peer 2/10 Unified Data Pipelines and Multi-Workload Capabilities Matt prompts Ion to elaborate on how Spark simplifies traditional data pipelines. Ion explains how micro-batching and in-memory state sharing allow a single unified engine to replace fragmented specialized tools like Storm, Impala, GraphLab, and Mahout.12:17–14:23 · Matt as informed peer 5/10 Comparing Spark Streaming and Apache Storm Matt demonstrates technical depth by asking Ion to contrast Spark Streaming with Apache Storm's approach. Ion clarifies the architectural differences, highlighting micro-batch fault tolerance while candidly conceding that Spark Streaming is not suitable for ultra-low latency sub-millisecond use cases like high-frequency financial trading.14:23–16:30 · Matt as informed peer 2/10 Databricks Cloud Platform Features and Capabilities Matt transitions to Databricks as a commercial enterprise. Ion describes the cloud platform's value proposition in simplifying big data workflows on AWS through managed clusters, notebooks, and production pipelines.16:30–18:43 · Matt as informed peer 3/10 Databricks and the Apache Spark Open Source Community Matt asks about the commercial dynamic between Databricks and the broader open source Spark community. Ion explains their open platform strategy, highlighting that expanding Spark adoption globally benefits their cloud service while preventing ecosystem fragmentation.18:43–20:43 · Matt as informed peer 1/10 Audience Q&A: Secondary Indexes and In-Memory Search Matt opens the floor to audience Q&A, where audience member Justin asks about secondary index support in Spark. Ion details active UC Berkeley research projects like Succinct that enable querying directly over compressed data without decompression.20:43–24:43 · Matt as informed peer 1/10 Audience Q&A: Spark SQL and Data Warehouse Coexistence Justin asks whether Spark SQL will coexist with or replace traditional data warehouses. Ion explains how sharing in-memory data objects across SQL and machine learning libraries avoids redundant data movement.24:43–26:57 · Matt as informed peer 1/10 Audience Q&A: SaaS Strategy, Competition, and Cloud Growth An audience member asks about an investor claim that Databricks is a 'SaaS killer'. Ion addresses the question by describing Databricks' focus on cloud-native deployments on AWS and leveraging data gravity.26:57–27:02 · Matt as informed peer 0/10 Conclusion and Audience Applause Matt wraps up the interview session with closing thanks to Ion Stoica, followed by audience applause.1:12–3:42 · Guest teaching 3/10 Ion Stoica's Academic Background and Conviva Matt introduces Ion Stoica and sets up the narrative of his transition from academia to entrepreneurship. Ion monologues on his background at CMU and Berkeley, recounting how building Conviva revealed early real-time data processing bottlenecks with Hadoop.3:42–6:08 · Guest teaching 4/10 The Origin Story of Apache Spark at UC Berkeley In a continuous monologue without host intervention, Ion details the genesis of Spark at UC Berkeley's RAD Lab. He explains why iterative machine learning algorithms were prohibitively slow on Hadoop MapReduce due to HDFS write and read overhead between iterations.6:08–9:36 · Guest teaching 5/10 Comparing Apache Spark and the Hadoop Stack Architecture Matt demonstrates technical understanding by asking whether Spark acts as a replacement processing engine for MapReduce over HDFS. Ion builds on this, outlining the three-layer Hadoop stack and detailing Spark's three core advantages: speed, rich API operators, and workload unification.9:36–12:17 · Guest teaching 5/10 Unified Data Pipelines and Multi-Workload Capabilities Matt prompts Ion to elaborate on how Spark simplifies traditional data pipelines. Ion explains how micro-batching and in-memory state sharing allow a single unified engine to replace fragmented specialized tools like Storm, Impala, GraphLab, and Mahout.12:17–14:23 · Guest teaching 5/10 Comparing Spark Streaming and Apache Storm Matt demonstrates technical depth by asking Ion to contrast Spark Streaming with Apache Storm's approach. Ion clarifies the architectural differences, highlighting micro-batch fault tolerance while candidly conceding that Spark Streaming is not suitable for ultra-low latency sub-millisecond use cases like high-frequency financial trading.14:23–16:30 · Guest teaching 3/10 Databricks Cloud Platform Features and Capabilities Matt transitions to Databricks as a commercial enterprise. Ion describes the cloud platform's value proposition in simplifying big data workflows on AWS through managed clusters, notebooks, and production pipelines.16:30–18:43 · Guest teaching 3/10 Databricks and the Apache Spark Open Source Community Matt asks about the commercial dynamic between Databricks and the broader open source Spark community. Ion explains their open platform strategy, highlighting that expanding Spark adoption globally benefits their cloud service while preventing ecosystem fragmentation.18:43–20:43 · Guest teaching 4/10 Audience Q&A: Secondary Indexes and In-Memory Search Matt opens the floor to audience Q&A, where audience member Justin asks about secondary index support in Spark. Ion details active UC Berkeley research projects like Succinct that enable querying directly over compressed data without decompression.20:43–24:43 · Guest teaching 5/10 Audience Q&A: Spark SQL and Data Warehouse Coexistence Justin asks whether Spark SQL will coexist with or replace traditional data warehouses. Ion explains how sharing in-memory data objects across SQL and machine learning libraries avoids redundant data movement.24:43–26:57 · Guest teaching 3/10 Audience Q&A: SaaS Strategy, Competition, and Cloud Growth An audience member asks about an investor claim that Databricks is a 'SaaS killer'. Ion addresses the question by describing Databricks' focus on cloud-native deployments on AWS and leveraging data gravity.26:57–27:02 · Guest teaching 0/10 Conclusion and Audience Applause Matt wraps up the interview session with closing thanks to Ion Stoica, followed by audience applause.1:12–3:42 · Guest disagreement 0/10 Ion Stoica's Academic Background and Conviva Matt introduces Ion Stoica and sets up the narrative of his transition from academia to entrepreneurship. Ion monologues on his background at CMU and Berkeley, recounting how building Conviva revealed early real-time data processing bottlenecks with Hadoop.3:42–6:08 · Guest disagreement 0/10 The Origin Story of Apache Spark at UC Berkeley In a continuous monologue without host intervention, Ion details the genesis of Spark at UC Berkeley's RAD Lab. He explains why iterative machine learning algorithms were prohibitively slow on Hadoop MapReduce due to HDFS write and read overhead between iterations.6:08–9:36 · Guest disagreement 0/10 Comparing Apache Spark and the Hadoop Stack Architecture Matt demonstrates technical understanding by asking whether Spark acts as a replacement processing engine for MapReduce over HDFS. Ion builds on this, outlining the three-layer Hadoop stack and detailing Spark's three core advantages: speed, rich API operators, and workload unification.9:36–12:17 · Guest disagreement 0/10 Unified Data Pipelines and Multi-Workload Capabilities Matt prompts Ion to elaborate on how Spark simplifies traditional data pipelines. Ion explains how micro-batching and in-memory state sharing allow a single unified engine to replace fragmented specialized tools like Storm, Impala, GraphLab, and Mahout.12:17–14:23 · Guest disagreement 1/10 Comparing Spark Streaming and Apache Storm Matt demonstrates technical depth by asking Ion to contrast Spark Streaming with Apache Storm's approach. Ion clarifies the architectural differences, highlighting micro-batch fault tolerance while candidly conceding that Spark Streaming is not suitable for ultra-low latency sub-millisecond use cases like high-frequency financial trading.14:23–16:30 · Guest disagreement 0/10 Databricks Cloud Platform Features and Capabilities Matt transitions to Databricks as a commercial enterprise. Ion describes the cloud platform's value proposition in simplifying big data workflows on AWS through managed clusters, notebooks, and production pipelines.16:30–18:43 · Guest disagreement 0/10 Databricks and the Apache Spark Open Source Community Matt asks about the commercial dynamic between Databricks and the broader open source Spark community. Ion explains their open platform strategy, highlighting that expanding Spark adoption globally benefits their cloud service while preventing ecosystem fragmentation.18:43–20:43 · Guest disagreement 0/10 Audience Q&A: Secondary Indexes and In-Memory Search Matt opens the floor to audience Q&A, where audience member Justin asks about secondary index support in Spark. Ion details active UC Berkeley research projects like Succinct that enable querying directly over compressed data without decompression.20:43–24:43 · Guest disagreement 1/10 Audience Q&A: Spark SQL and Data Warehouse Coexistence Justin asks whether Spark SQL will coexist with or replace traditional data warehouses. Ion explains how sharing in-memory data objects across SQL and machine learning libraries avoids redundant data movement.24:43–26:57 · Guest disagreement 1/10 Audience Q&A: SaaS Strategy, Competition, and Cloud Growth An audience member asks about an investor claim that Databricks is a 'SaaS killer'. Ion addresses the question by describing Databricks' focus on cloud-native deployments on AWS and leveraging data gravity.26:57–27:02 · Guest disagreement 0/10 Conclusion and Audience Applause Matt wraps up the interview session with closing thanks to Ion Stoica, followed by audience applause.1:12–3:42 · Matt pushing back 0/10 Ion Stoica's Academic Background and Conviva Matt introduces Ion Stoica and sets up the narrative of his transition from academia to entrepreneurship. Ion monologues on his background at CMU and Berkeley, recounting how building Conviva revealed early real-time data processing bottlenecks with Hadoop.3:42–6:08 · Matt pushing back 0/10 The Origin Story of Apache Spark at UC Berkeley In a continuous monologue without host intervention, Ion details the genesis of Spark at UC Berkeley's RAD Lab. He explains why iterative machine learning algorithms were prohibitively slow on Hadoop MapReduce due to HDFS write and read overhead between iterations.6:08–9:36 · Matt pushing back 1/10 Comparing Apache Spark and the Hadoop Stack Architecture Matt demonstrates technical understanding by asking whether Spark acts as a replacement processing engine for MapReduce over HDFS. Ion builds on this, outlining the three-layer Hadoop stack and detailing Spark's three core advantages: speed, rich API operators, and workload unification.9:36–12:17 · Matt pushing back 0/10 Unified Data Pipelines and Multi-Workload Capabilities Matt prompts Ion to elaborate on how Spark simplifies traditional data pipelines. Ion explains how micro-batching and in-memory state sharing allow a single unified engine to replace fragmented specialized tools like Storm, Impala, GraphLab, and Mahout.12:17–14:23 · Matt pushing back 2/10 Comparing Spark Streaming and Apache Storm Matt demonstrates technical depth by asking Ion to contrast Spark Streaming with Apache Storm's approach. Ion clarifies the architectural differences, highlighting micro-batch fault tolerance while candidly conceding that Spark Streaming is not suitable for ultra-low latency sub-millisecond use cases like high-frequency financial trading.14:23–16:30 · Matt pushing back 0/10 Databricks Cloud Platform Features and Capabilities Matt transitions to Databricks as a commercial enterprise. Ion describes the cloud platform's value proposition in simplifying big data workflows on AWS through managed clusters, notebooks, and production pipelines.16:30–18:43 · Matt pushing back 0/10 Databricks and the Apache Spark Open Source Community Matt asks about the commercial dynamic between Databricks and the broader open source Spark community. Ion explains their open platform strategy, highlighting that expanding Spark adoption globally benefits their cloud service while preventing ecosystem fragmentation.18:43–20:43 · Matt pushing back 0/10 Audience Q&A: Secondary Indexes and In-Memory Search Matt opens the floor to audience Q&A, where audience member Justin asks about secondary index support in Spark. Ion details active UC Berkeley research projects like Succinct that enable querying directly over compressed data without decompression.20:43–24:43 · Matt pushing back 0/10 Audience Q&A: Spark SQL and Data Warehouse Coexistence Justin asks whether Spark SQL will coexist with or replace traditional data warehouses. Ion explains how sharing in-memory data objects across SQL and machine learning libraries avoids redundant data movement.24:43–26:57 · Matt pushing back 0/10 Audience Q&A: SaaS Strategy, Competition, and Cloud Growth An audience member asks about an investor claim that Databricks is a 'SaaS killer'. Ion addresses the question by describing Databricks' focus on cloud-native deployments on AWS and leveraging data gravity.26:57–27:02 · Matt pushing back 0/10 Conclusion and Audience Applause Matt wraps up the interview session with closing thanks to Ion Stoica, followed by audience applause.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 37.4% · guest 62.6%0:00 · Matt 37.4% · guest 62.6%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 25.5% · guest 74.5%6:00 · Matt 25.5% · guest 74.5%9:00 · Matt 4.6% · guest 95.4%9:00 · Matt 4.6% · guest 95.4%12:00 · Matt 14.3% · guest 85.7%12:00 · Matt 14.3% · guest 85.7%15:00 · Matt 10.4% · guest 89.6%15:00 · Matt 10.4% · guest 89.6%18:00 · Matt 14.4% · guest 85.6%18:00 · Matt 14.4% · guest 85.6%21:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%24:00 · Matt 4.4% · guest 95.6%24:00 · Matt 4.4% · guest 95.6%27:00 · Matt 68.7% · guest 31.3%27:00 · Matt 68.7% · guest 31.3%
Sharpest disagreement ▶ 13:40 Frank acknowledgment of technical limitations

Ion firmly pushes back against overhyped expectations by explicitly stating that Spark Streaming micro-batching cannot achieve millisecond latencies and is unsuitable for financial trading.

Hardest push from Matt ▶ 12:17 Host presses on Spark Streaming vs Storm comparison

Matt refuses to accept a high-level summary and specifically demands a technical contrast between Spark Streaming and Apache Storm regarding micro-batching.

Biggest teaching moment ▶ 6:39 Ion maps out the 3-tier Hadoop stack

Ion provides a structured breakdown of the Hadoop ecosystem into storage, resource management, and execution layers to precise where Spark fits.

Matt holds his own ▶ 6:08 Matt demonstrates firm grasp of big data stack architecture

Matt accurately frames the distinction between HDFS storage and MapReduce execution, asking precise questions about where Spark replaces components.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Ion Stoica's Academic Background and Conviva 1300 Matt introduces Ion Stoica and sets up the narrative of his transition from academia to entrepreneurship. Ion monologues on his background at CMU and Berkeley, recounting how building Conviva revealed early real-time data processing bottlenecks with Hadoop.
The Origin Story of Apache Spark at UC Berkeley 0400 In a continuous monologue without host intervention, Ion details the genesis of Spark at UC Berkeley's RAD Lab. He explains why iterative machine learning algorithms were prohibitively slow on Hadoop MapReduce due to HDFS write and read overhead between iterations.
Comparing Apache Spark and the Hadoop Stack Architecture 4501 Matt demonstrates technical understanding by asking whether Spark acts as a replacement processing engine for MapReduce over HDFS. Ion builds on this, outlining the three-layer Hadoop stack and detailing Spark's three core advantages: speed, rich API operators, and workload unification.
Unified Data Pipelines and Multi-Workload Capabilities 2500 Matt prompts Ion to elaborate on how Spark simplifies traditional data pipelines. Ion explains how micro-batching and in-memory state sharing allow a single unified engine to replace fragmented specialized tools like Storm, Impala, GraphLab, and Mahout.
Comparing Spark Streaming and Apache Storm 5512 Matt demonstrates technical depth by asking Ion to contrast Spark Streaming with Apache Storm's approach. Ion clarifies the architectural differences, highlighting micro-batch fault tolerance while candidly conceding that Spark Streaming is not suitable for ultra-low latency sub-millisecond use cases like high-frequency financial trading.
Databricks Cloud Platform Features and Capabilities 2300 Matt transitions to Databricks as a commercial enterprise. Ion describes the cloud platform's value proposition in simplifying big data workflows on AWS through managed clusters, notebooks, and production pipelines.
Databricks and the Apache Spark Open Source Community 3300 Matt asks about the commercial dynamic between Databricks and the broader open source Spark community. Ion explains their open platform strategy, highlighting that expanding Spark adoption globally benefits their cloud service while preventing ecosystem fragmentation.
Audience Q&A: Secondary Indexes and In-Memory Search 1400 Matt opens the floor to audience Q&A, where audience member Justin asks about secondary index support in Spark. Ion details active UC Berkeley research projects like Succinct that enable querying directly over compressed data without decompression.
Audience Q&A: Spark SQL and Data Warehouse Coexistence 1510 Justin asks whether Spark SQL will coexist with or replace traditional data warehouses. Ion explains how sharing in-memory data objects across SQL and machine learning libraries avoids redundant data movement.
Audience Q&A: SaaS Strategy, Competition, and Cloud Growth 1310 An audience member asks about an investor claim that Databricks is a 'SaaS killer'. Ion addresses the question by describing Databricks' focus on cloud-native deployments on AWS and leveraging data gravity.
Conclusion and Audience Applause 0000 Matt wraps up the interview session with closing thanks to Ion Stoica, followed by audience applause.

Statements from this episode (14)

Assertion Supported
Stoica: Early Hadoop was limited to batch processing
“So at that point, in big data space we there was Hadoop just started, but of course that was, by, back then it was mostly, you know, batch, computation, so you could do historical analysis, but not much more than that.”
Ion Stoica Apr 2, 2015 ▶ 3:03
Assertion Supported
Stoica: Hadoop's HDFS read/write cycle crippled early iterative machine learning
“If you look at the machine learning, it's, fundamentally, it's an iterative algorithm, and every iteration is turned into a Hadoop job. So between the iteration, you write the data and read the data from HDFS, so that's why it's very slow.”
Ion Stoica Apr 2, 2015 ▶ 4:08
Disclosure
Stoica: Apache Spark was created for iterative machine learning and interactive queries
“And Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support machine learning, as well as interactive computation, right…”
Ion Stoica Apr 2, 2015 ▶ 5:45
Assertion Supported
Stoica: Apache Spark originally ran on Mesos before adding YARN and standalone support
“As originally was built to run on top of Mesos. Today is working on Yarn, working, you know, standalone, and is working also in addition to HDFS, you know, imports and exports data to many other data sources.”
Ion Stoica Apr 2, 2015 ▶ 7:07
Assertion Supported
Stoica: Apache Spark won the TerraSort benchmark processing data out of memory
“Just October last year, we had this we won this kind of TerraSort benchmark. And in those, in that benchmark, the data, it's not in memory. Right? It's SSDs and so forth.”
Ion Stoica Apr 2, 2015 ▶ 8:04
Opinion
Stoica: Hadoop remains a very great batch processing engine
“Hadoop is still a very great, ah, batch engine.”
Ion Stoica Apr 2, 2015 ▶ 11:23
Assertion Supported
Stoica: Apache Spark supports all major data workloads with one engine
“While we spark, You can use only one engine and only one API to support all these workloads.”
Ion Stoica Apr 2, 2015 ▶ 12:00
Assertion Not checkable as stated
Stoica: Spark Streaming and Storm have roughly similar throughput
“They are roughly, you know, in terms of the throughput, they are roughly similar.”
Ion Stoica Apr 2, 2015 ▶ 12:51
Assertion Not checkable as stated
Stoica: Spark Streaming micro-batching is unfit for high-frequency trading latency
“It's very hard to, you know, if you want millisecond latency from the time you, ah, the data entered in the system until you get the result, it's very hard to get. You can get latencies of several hundreds of milliseconds, but milliseconds, very hard. Ah, you,…”
Ion Stoica Apr 2, 2015 ▶ 13:40
Assertion Supported
Stoica: Apache Spark has exceeded 500 active open-source contributors
“We exceeded 500 contributors. It is the most active big data project right now, Spark.”
Ion Stoica Apr 2, 2015 ▶ 16:47
Assertion Supported
Stoica: Databricks still provides the majority of Apache Spark open-source contributions
“A lot of contributions, still the majority of contributions comes from Databricks.”
Ion Stoica Apr 2, 2015 ▶ 17:00
Disclosure
Stoica: Databricks offers only cloud services, not a custom Spark distribution
“We don't have our own distribution. We provide only the service.”
Ion Stoica Apr 2, 2015 ▶ 17:30
Assertion Supported
Stoica: Berkeley's Succinct enables query processing on compressed data
“There is a related project that, ah, Berkeley is called, ah, succinct, ah, which actually go even more than, you know, beyond that. It's, ah it's a project that allows you to, you know, provide allow you to have, you know, query processing on the compressed da…”
Ion Stoica Apr 2, 2015 ▶ 19:31
Prediction Not checkable as stated
Stoica predicts big data infrastructure will eventually unify around one system
“Now, I do think that looking forward you are going to see more and more of this unification. This happens in many other industries and technologies. I do think this will happen in big data because it's so much easier if you have only one system than if you nee…”
Ion Stoica Apr 2, 2015 ▶ 21:53
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.