Mar 15, 2021 · 24m · mad

Fireside Chat: Arjun Narayan (Founder & CEO, Materialize) with Matt Turck (Partner, FirstMark)

Arjun Narayan · 19m spoken Matt Turck · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC fireside chat, host Matt Turck interviews Arjun Narayan, Founder and CEO of Materialize, about the evolution of streaming data architectures, the power of SQL-based real-time analytics, and the core technology driving Materialize.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.1% of the talking time here. How this is scored →

Matt as informed peer 2.6 Guest teaching 5.4 Guest disagreement 1.3 Matt pushing back 0.3
05100:0010:0020:000:10–2:42 · Matt as informed peer 3/10 Understanding Streaming Data vs. Batch Processing Matt opens with a broad question on streaming and demonstrates basic industry knowledge by pointing out how the expected timeline for streaming adoption lagged behind industry hype. Arjun politely thanks Matt and explains the fundamental difference between batch processing and streaming using real-time financial use cases.2:42–4:47 · Matt as informed peer 2/10 Demystifying Apache Kafka as a Core Infrastructure Component Matt acts as an audience surrogate, asking Arjun to define Apache Kafka for non-technical listeners. Arjun provides a clear explanation of message brokers, pub-sub architecture, and how Kafka enables real-time microservices without direct coordination.4:47–6:48 · Matt as informed peer 2/10 The Need for Streaming Databases and Incremental Computation Matt asks why a streaming database is necessary given existing tools. Arjun educates Matt on why legacy pull-based batch analytics engines cannot scale to millisecond-level incremental recomputation.6:48–11:53 · Matt as informed peer 4/10 The Resurgence and Power of SQL in Modern Data Stacks Matt shows good domain understanding by outlining the historic industry cycle from SQL to NoSQL back to SQL. Arjun validates this context, strongly criticizing tech pitches that force users to throw away SQL, while explaining how NewSQL and cloud data warehouses restored SQL dominance.11:53–16:20 · Matt as informed peer 3/10 Timely Dataflow and the Engine Driving Materialize Matt mentions Timely Dataflow by name and asks about its technical origin. Arjun gives a detailed breakdown of Frank McSherry's research and contrasts Timely Dataflow with older stream engines like Apache Storm using an engine vs car analogy.16:20–19:33 · Matt as informed peer 2/10 Unpacking Materialized Views, Trigger Logic, and Update Granularity Matt asks questions directly from his notes regarding materialized views, trigger logic, and granularity. Arjun clarifies the mechanics of pushing computation upon data changes rather than query request time.19:33–21:37 · Matt as informed peer 3/10 Integrating with dbt to Bridge Batch and Streaming Workflows Matt connects Materialize's dbt integration with a previous event speaker. Arjun playfully offers a slightly provocative reframe of dbt as 'GitHub for SQL' while detailing how dbt bridges batch and streaming workflows.21:37–23:42 · Matt as informed peer 2/10 Future Roadmap: Hosted Cloud, Tiered Storage, and High Availability Matt prompts for the future roadmap and cleanly wraps up the session. Arjun explains upcoming developments in hosted cloud services, tiered S3 storage, and database replication.0:10–2:42 · Guest teaching 5/10 Understanding Streaming Data vs. Batch Processing Matt opens with a broad question on streaming and demonstrates basic industry knowledge by pointing out how the expected timeline for streaming adoption lagged behind industry hype. Arjun politely thanks Matt and explains the fundamental difference between batch processing and streaming using real-time financial use cases.2:42–4:47 · Guest teaching 6/10 Demystifying Apache Kafka as a Core Infrastructure Component Matt acts as an audience surrogate, asking Arjun to define Apache Kafka for non-technical listeners. Arjun provides a clear explanation of message brokers, pub-sub architecture, and how Kafka enables real-time microservices without direct coordination.4:47–6:48 · Guest teaching 6/10 The Need for Streaming Databases and Incremental Computation Matt asks why a streaming database is necessary given existing tools. Arjun educates Matt on why legacy pull-based batch analytics engines cannot scale to millisecond-level incremental recomputation.6:48–11:53 · Guest teaching 5/10 The Resurgence and Power of SQL in Modern Data Stacks Matt shows good domain understanding by outlining the historic industry cycle from SQL to NoSQL back to SQL. Arjun validates this context, strongly criticizing tech pitches that force users to throw away SQL, while explaining how NewSQL and cloud data warehouses restored SQL dominance.11:53–16:20 · Guest teaching 6/10 Timely Dataflow and the Engine Driving Materialize Matt mentions Timely Dataflow by name and asks about its technical origin. Arjun gives a detailed breakdown of Frank McSherry's research and contrasts Timely Dataflow with older stream engines like Apache Storm using an engine vs car analogy.16:20–19:33 · Guest teaching 6/10 Unpacking Materialized Views, Trigger Logic, and Update Granularity Matt asks questions directly from his notes regarding materialized views, trigger logic, and granularity. Arjun clarifies the mechanics of pushing computation upon data changes rather than query request time.19:33–21:37 · Guest teaching 5/10 Integrating with dbt to Bridge Batch and Streaming Workflows Matt connects Materialize's dbt integration with a previous event speaker. Arjun playfully offers a slightly provocative reframe of dbt as 'GitHub for SQL' while detailing how dbt bridges batch and streaming workflows.21:37–23:42 · Guest teaching 4/10 Future Roadmap: Hosted Cloud, Tiered Storage, and High Availability Matt prompts for the future roadmap and cleanly wraps up the session. Arjun explains upcoming developments in hosted cloud services, tiered S3 storage, and database replication.0:10–2:42 · Guest disagreement 1/10 Understanding Streaming Data vs. Batch Processing Matt opens with a broad question on streaming and demonstrates basic industry knowledge by pointing out how the expected timeline for streaming adoption lagged behind industry hype. Arjun politely thanks Matt and explains the fundamental difference between batch processing and streaming using real-time financial use cases.2:42–4:47 · Guest disagreement 1/10 Demystifying Apache Kafka as a Core Infrastructure Component Matt acts as an audience surrogate, asking Arjun to define Apache Kafka for non-technical listeners. Arjun provides a clear explanation of message brokers, pub-sub architecture, and how Kafka enables real-time microservices without direct coordination.4:47–6:48 · Guest disagreement 1/10 The Need for Streaming Databases and Incremental Computation Matt asks why a streaming database is necessary given existing tools. Arjun educates Matt on why legacy pull-based batch analytics engines cannot scale to millisecond-level incremental recomputation.6:48–11:53 · Guest disagreement 2/10 The Resurgence and Power of SQL in Modern Data Stacks Matt shows good domain understanding by outlining the historic industry cycle from SQL to NoSQL back to SQL. Arjun validates this context, strongly criticizing tech pitches that force users to throw away SQL, while explaining how NewSQL and cloud data warehouses restored SQL dominance.11:53–16:20 · Guest disagreement 1/10 Timely Dataflow and the Engine Driving Materialize Matt mentions Timely Dataflow by name and asks about its technical origin. Arjun gives a detailed breakdown of Frank McSherry's research and contrasts Timely Dataflow with older stream engines like Apache Storm using an engine vs car analogy.16:20–19:33 · Guest disagreement 1/10 Unpacking Materialized Views, Trigger Logic, and Update Granularity Matt asks questions directly from his notes regarding materialized views, trigger logic, and granularity. Arjun clarifies the mechanics of pushing computation upon data changes rather than query request time.19:33–21:37 · Guest disagreement 2/10 Integrating with dbt to Bridge Batch and Streaming Workflows Matt connects Materialize's dbt integration with a previous event speaker. Arjun playfully offers a slightly provocative reframe of dbt as 'GitHub for SQL' while detailing how dbt bridges batch and streaming workflows.21:37–23:42 · Guest disagreement 1/10 Future Roadmap: Hosted Cloud, Tiered Storage, and High Availability Matt prompts for the future roadmap and cleanly wraps up the session. Arjun explains upcoming developments in hosted cloud services, tiered S3 storage, and database replication.0:10–2:42 · Matt pushing back 1/10 Understanding Streaming Data vs. Batch Processing Matt opens with a broad question on streaming and demonstrates basic industry knowledge by pointing out how the expected timeline for streaming adoption lagged behind industry hype. Arjun politely thanks Matt and explains the fundamental difference between batch processing and streaming using real-time financial use cases.2:42–4:47 · Matt pushing back 0/10 Demystifying Apache Kafka as a Core Infrastructure Component Matt acts as an audience surrogate, asking Arjun to define Apache Kafka for non-technical listeners. Arjun provides a clear explanation of message brokers, pub-sub architecture, and how Kafka enables real-time microservices without direct coordination.4:47–6:48 · Matt pushing back 0/10 The Need for Streaming Databases and Incremental Computation Matt asks why a streaming database is necessary given existing tools. Arjun educates Matt on why legacy pull-based batch analytics engines cannot scale to millisecond-level incremental recomputation.6:48–11:53 · Matt pushing back 1/10 The Resurgence and Power of SQL in Modern Data Stacks Matt shows good domain understanding by outlining the historic industry cycle from SQL to NoSQL back to SQL. Arjun validates this context, strongly criticizing tech pitches that force users to throw away SQL, while explaining how NewSQL and cloud data warehouses restored SQL dominance.11:53–16:20 · Matt pushing back 0/10 Timely Dataflow and the Engine Driving Materialize Matt mentions Timely Dataflow by name and asks about its technical origin. Arjun gives a detailed breakdown of Frank McSherry's research and contrasts Timely Dataflow with older stream engines like Apache Storm using an engine vs car analogy.16:20–19:33 · Matt pushing back 0/10 Unpacking Materialized Views, Trigger Logic, and Update Granularity Matt asks questions directly from his notes regarding materialized views, trigger logic, and granularity. Arjun clarifies the mechanics of pushing computation upon data changes rather than query request time.19:33–21:37 · Matt pushing back 0/10 Integrating with dbt to Bridge Batch and Streaming Workflows Matt connects Materialize's dbt integration with a previous event speaker. Arjun playfully offers a slightly provocative reframe of dbt as 'GitHub for SQL' while detailing how dbt bridges batch and streaming workflows.21:37–23:42 · Matt pushing back 0/10 Future Roadmap: Hosted Cloud, Tiered Storage, and High Availability Matt prompts for the future roadmap and cleanly wraps up the session. Arjun explains upcoming developments in hosted cloud services, tiered S3 storage, and database replication.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 28.8% · guest 71.2%0:00 · Matt 28.8% · guest 71.2%3:00 · Matt 1.8% · guest 98.2%3:00 · Matt 1.8% · guest 98.2%6:00 · Matt 7.6% · guest 92.4%6:00 · Matt 7.6% · guest 92.4%9:00 · Matt 13.5% · guest 86.5%9:00 · Matt 13.5% · guest 86.5%12:00 · Matt 4% · guest 96%12:00 · Matt 4% · guest 96%15:00 · Matt 10.6% · guest 89.4%15:00 · Matt 10.6% · guest 89.4%18:00 · Matt 12.4% · guest 87.6%18:00 · Matt 12.4% · guest 87.6%21:00 · Matt 12.7% · guest 87.3%21:00 · Matt 12.7% · guest 87.3%24:00 · Matt 100% · guest 0%24:00 · Matt 100% · guest 0%
Sharpest disagreement ▶ 8:20 Dismissal of non-SQL rewrite efforts

Arjun forcefully rejects the premise of database tools that require rewriting systems from scratch, stating that such efforts are 'largely doomed'.

Hardest push from Matt ▶ 1:47 Challenging the streaming hype timeline

Matt politely challenges the guest by noting that despite annual claims of streaming becoming dominant, adoption took much longer than expected.

Biggest teaching moment ▶ 13:40 Explaining fundamental trade-off breakthroughs in Timely Dataflow

Arjun educates the host on how previous stream processors forced trade-offs between complex batch computations and low latency, whereas Timely Dataflow solved both.

Matt holds his own ▶ 9:20 Framing the NoSQL to SQL historical trajectory

Matt displays strong industry context by accurately summarizing the multi-year trajectory from relational databases to NoSQL hype and the subsequent return to SQL.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Understanding Streaming Data vs. Batch Processing 3511 Matt opens with a broad question on streaming and demonstrates basic industry knowledge by pointing out how the expected timeline for streaming adoption lagged behind industry hype. Arjun politely thanks Matt and explains the fundamental difference between batch processing and streaming using real-time financial use cases.
Demystifying Apache Kafka as a Core Infrastructure Component 2610 Matt acts as an audience surrogate, asking Arjun to define Apache Kafka for non-technical listeners. Arjun provides a clear explanation of message brokers, pub-sub architecture, and how Kafka enables real-time microservices without direct coordination.
The Need for Streaming Databases and Incremental Computation 2610 Matt asks why a streaming database is necessary given existing tools. Arjun educates Matt on why legacy pull-based batch analytics engines cannot scale to millisecond-level incremental recomputation.
The Resurgence and Power of SQL in Modern Data Stacks 4521 Matt shows good domain understanding by outlining the historic industry cycle from SQL to NoSQL back to SQL. Arjun validates this context, strongly criticizing tech pitches that force users to throw away SQL, while explaining how NewSQL and cloud data warehouses restored SQL dominance.
Timely Dataflow and the Engine Driving Materialize 3610 Matt mentions Timely Dataflow by name and asks about its technical origin. Arjun gives a detailed breakdown of Frank McSherry's research and contrasts Timely Dataflow with older stream engines like Apache Storm using an engine vs car analogy.
Unpacking Materialized Views, Trigger Logic, and Update Granularity 2610 Matt asks questions directly from his notes regarding materialized views, trigger logic, and granularity. Arjun clarifies the mechanics of pushing computation upon data changes rather than query request time.
Integrating with dbt to Bridge Batch and Streaming Workflows 3520 Matt connects Materialize's dbt integration with a previous event speaker. Arjun playfully offers a slightly provocative reframe of dbt as 'GitHub for SQL' while detailing how dbt bridges batch and streaming workflows.
Future Roadmap: Hosted Cloud, Tiered Storage, and High Availability 2410 Matt prompts for the future roadmap and cleanly wraps up the session. Arjun explains upcoming developments in hosted cloud services, tiered S3 storage, and database replication.

Statements from this episode (9)

Assertion Not checkable as stated
Arjun Narayan: Streaming data processing has expanded beyond niche financial applications
“And what we've been seeing is that streaming has started to become over the, over several decades and particularly in the past few years, much, much more broadly applicable. Beyond those small niche use cases as more and more businesses and use users benefit f…”
Arjun Narayan Mar 15, 2021 ▶ 1:24
Insight
Narayan: Apache Kafka is essential for building and operating microservices
“It has been to, in my opinion, a key enabler of microservices. I think it's pretty difficult to build and operate a decentralized set of microservices without first adopting something like Kafka in your organization to just move the data between all of these v…”
Arjun Narayan Mar 15, 2021 ▶ 4:25
Assertion Not checkable as stated
Arjun Narayan: Standard analytics databases add latency by lacking incremental computation
“So, so stopping everything and recomputing from scratch is not really a framework that scales to these lower and lower latencies, which is why fundamentally a lot of analytics databases today including some of the more famous ones, they would prefer it if you …”
Arjun Narayan Mar 15, 2021 ▶ 5:45
Insight
Narayan: Software requiring code rewrites in new languages is largely doomed
“Any pitch. I'm generally very skeptical where you can tell folks, you can have all these great new benefits of low latency or whatever it is, but you got to start all over from scratch, right? You have to throw everything out there and you're going to rebuild …”
Arjun Narayan Mar 15, 2021 ▶ 7:37
Opinion
Arjun Narayan: Snowflake succeeded by wrapping cloud architecture in familiar SQL
“Snowflake is a fantastic example of, under the hood, they, you know, they have a very modern microservice architecture, but it's all sort of very neatly wrapped up with a bow on top, such that it looks like a SQL database, and that's much, much more attractive…”
Arjun Narayan Mar 15, 2021 ▶ 11:20
Assertion Supported
Narayan: Timely Dataflow was first stream processor to match batch processing capabilities
“It was sort of the first, what I would describe as The very first stream processor that could do everything that batch processors could do.”
Arjun Narayan Mar 15, 2021 ▶ 13:13
Assertion Not checkable as stated
Narayan: Materialize provides 99.9% of batch database functionality
“And we really think materialize the product is the first database that gives you know, I don't want to say literally all the functionality, but to the morally speaking, you know, 99.9% of the functionality that you can write in a batch database.”
Arjun Narayan Mar 15, 2021 ▶ 18:13
Opinion
Narayan: dbt serves as GitHub for SQL code
“I think of it as like GitHub for all of your SQL, right?”
Arjun Narayan Mar 15, 2021 ▶ 19:58
Prediction Not checkable as stated
Narayan: Streaming will remain niche until tooling matches batch systems
“Until streaming gets to the same level of tooling, and a large part of that tooling is dbt as existing batch systems you know, it will still remain a fairly niche technology.”
Arjun Narayan Mar 15, 2021 ▶ 21:17
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.