Jun 28, 2022 · 26m · mad

The Next Layer of the Modern Data Stack | dbt's Tristan Handy

Tristan Handy · 21m spoken Matt Turck · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat from Data Driven NYC, Matt Turck interviews dbt Labs Founder & CEO Tristan Handy about the evolution of the modern data stack, the role of dbt in bringing software engineering rigor to data transformation, and the future of semantic layers and polyglot data processing.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 8.2% of the talking time here. How this is scored →

Matt as informed peer 2.7 Guest teaching 4.3 Guest disagreement 0.6 Matt pushing back 1.0
05100:0010:0020:000:08–3:02 · Matt as informed peer 1/10 Company Founding and Distributed Work Culture Host opens with standard operational questions about founding dates and distributed work. Guest explains dbt's remote culture based on the GitLab handbook.3:02–8:21 · Matt as informed peer 2/10 Origins and Paradigm Shift of the Modern Data Stack Host prompts guest to outline the steps of the data journey. Guest provides an in-depth explanation of the modern data stack paradigm shift using a factory electrification analogy.8:21–10:39 · Matt as informed peer 2/10 Understanding Data Transformation Through Real-World Examples Host asks for a practical example of data transformation to ground the discussion for listeners. Guest educates the host using a detailed green onions unit economics calculation example.10:39–15:55 · Matt as informed peer 5/10 Core Functionality of dbt and Abstraction Frameworks Host offers an insightful comparison between dbt and web frameworks like Ruby on Rails. Guest validates the comparison and explains how software engineering patterns like Git apply to data.15:55–18:57 · Matt as informed peer 3/10 Transitioning from Open-Source dbt Core to dbt Cloud Host asks how dbt manages the boundary between open source dbt Core and paid dbt Cloud. Guest outlines the separation between stateless SQL compilation and operational scheduling engines.18:57–23:58 · Matt as informed peer 2/10 The Semantic Layer and Ecosystem Vision Host asks where dbt's multi-year strategic roadmap leads. Guest elaborates on the semantic layer and uses an Apple App Store analogy to frame dbt as ecosystem infrastructure.23:58–26:25 · Matt as informed peer 4/10 Audience Q&A: Polyglot Languages and Machine Learning Workflows Q&A challenges whether dbt's SQL-first focus contradicts higher abstraction layers. Guest pushes back on the 'SQL maximalism' framing and clarifies the platform's boundaries regarding machine learning.0:08–3:02 · Guest teaching 2/10 Company Founding and Distributed Work Culture Host opens with standard operational questions about founding dates and distributed work. Guest explains dbt's remote culture based on the GitLab handbook.3:02–8:21 · Guest teaching 5/10 Origins and Paradigm Shift of the Modern Data Stack Host prompts guest to outline the steps of the data journey. Guest provides an in-depth explanation of the modern data stack paradigm shift using a factory electrification analogy.8:21–10:39 · Guest teaching 5/10 Understanding Data Transformation Through Real-World Examples Host asks for a practical example of data transformation to ground the discussion for listeners. Guest educates the host using a detailed green onions unit economics calculation example.10:39–15:55 · Guest teaching 4/10 Core Functionality of dbt and Abstraction Frameworks Host offers an insightful comparison between dbt and web frameworks like Ruby on Rails. Guest validates the comparison and explains how software engineering patterns like Git apply to data.15:55–18:57 · Guest teaching 4/10 Transitioning from Open-Source dbt Core to dbt Cloud Host asks how dbt manages the boundary between open source dbt Core and paid dbt Cloud. Guest outlines the separation between stateless SQL compilation and operational scheduling engines.18:57–23:58 · Guest teaching 5/10 The Semantic Layer and Ecosystem Vision Host asks where dbt's multi-year strategic roadmap leads. Guest elaborates on the semantic layer and uses an Apple App Store analogy to frame dbt as ecosystem infrastructure.23:58–26:25 · Guest teaching 5/10 Audience Q&A: Polyglot Languages and Machine Learning Workflows Q&A challenges whether dbt's SQL-first focus contradicts higher abstraction layers. Guest pushes back on the 'SQL maximalism' framing and clarifies the platform's boundaries regarding machine learning.0:08–3:02 · Guest disagreement 0/10 Company Founding and Distributed Work Culture Host opens with standard operational questions about founding dates and distributed work. Guest explains dbt's remote culture based on the GitLab handbook.3:02–8:21 · Guest disagreement 1/10 Origins and Paradigm Shift of the Modern Data Stack Host prompts guest to outline the steps of the data journey. Guest provides an in-depth explanation of the modern data stack paradigm shift using a factory electrification analogy.8:21–10:39 · Guest disagreement 0/10 Understanding Data Transformation Through Real-World Examples Host asks for a practical example of data transformation to ground the discussion for listeners. Guest educates the host using a detailed green onions unit economics calculation example.10:39–15:55 · Guest disagreement 0/10 Core Functionality of dbt and Abstraction Frameworks Host offers an insightful comparison between dbt and web frameworks like Ruby on Rails. Guest validates the comparison and explains how software engineering patterns like Git apply to data.15:55–18:57 · Guest disagreement 0/10 Transitioning from Open-Source dbt Core to dbt Cloud Host asks how dbt manages the boundary between open source dbt Core and paid dbt Cloud. Guest outlines the separation between stateless SQL compilation and operational scheduling engines.18:57–23:58 · Guest disagreement 0/10 The Semantic Layer and Ecosystem Vision Host asks where dbt's multi-year strategic roadmap leads. Guest elaborates on the semantic layer and uses an Apple App Store analogy to frame dbt as ecosystem infrastructure.23:58–26:25 · Guest disagreement 3/10 Audience Q&A: Polyglot Languages and Machine Learning Workflows Q&A challenges whether dbt's SQL-first focus contradicts higher abstraction layers. Guest pushes back on the 'SQL maximalism' framing and clarifies the platform's boundaries regarding machine learning.0:08–3:02 · Matt pushing back 0/10 Company Founding and Distributed Work Culture Host opens with standard operational questions about founding dates and distributed work. Guest explains dbt's remote culture based on the GitLab handbook.3:02–8:21 · Matt pushing back 0/10 Origins and Paradigm Shift of the Modern Data Stack Host prompts guest to outline the steps of the data journey. Guest provides an in-depth explanation of the modern data stack paradigm shift using a factory electrification analogy.8:21–10:39 · Matt pushing back 1/10 Understanding Data Transformation Through Real-World Examples Host asks for a practical example of data transformation to ground the discussion for listeners. Guest educates the host using a detailed green onions unit economics calculation example.10:39–15:55 · Matt pushing back 1/10 Core Functionality of dbt and Abstraction Frameworks Host offers an insightful comparison between dbt and web frameworks like Ruby on Rails. Guest validates the comparison and explains how software engineering patterns like Git apply to data.15:55–18:57 · Matt pushing back 1/10 Transitioning from Open-Source dbt Core to dbt Cloud Host asks how dbt manages the boundary between open source dbt Core and paid dbt Cloud. Guest outlines the separation between stateless SQL compilation and operational scheduling engines.18:57–23:58 · Matt pushing back 0/10 The Semantic Layer and Ecosystem Vision Host asks where dbt's multi-year strategic roadmap leads. Guest elaborates on the semantic layer and uses an Apple App Store analogy to frame dbt as ecosystem infrastructure.23:58–26:25 · Matt pushing back 4/10 Audience Q&A: Polyglot Languages and Machine Learning Workflows Q&A challenges whether dbt's SQL-first focus contradicts higher abstraction layers. Guest pushes back on the 'SQL maximalism' framing and clarifies the platform's boundaries regarding machine learning.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 12.6% · guest 87.4%0:00 · Matt 12.6% · guest 87.4%3:00 · Matt 4.3% · guest 95.7%3:00 · Matt 4.3% · guest 95.7%6:00 · Matt 3.5% · guest 96.5%6:00 · Matt 3.5% · guest 96.5%9:00 · Matt 9.7% · guest 90.3%9:00 · Matt 9.7% · guest 90.3%12:00 · Matt 8.1% · guest 91.9%12:00 · Matt 8.1% · guest 91.9%15:00 · Matt 22.1% · guest 77.9%15:00 · Matt 22.1% · guest 77.9%18:00 · Matt 3% · guest 97%18:00 · Matt 3% · guest 97%21:00 · Matt 4.2% · guest 95.8%21:00 · Matt 4.2% · guest 95.8%24:00 · Matt 5.4% · guest 94.6%24:00 · Matt 5.4% · guest 94.6%
Sharpest disagreement ▶ 23:58 Rejecting SQL maximalism label

The guest directly disputes the premise of the Q&A question, rejecting the claim that dbt adheres to SQL maximalism and re-framing their philosophy around persona preference and bringing code to data.

Hardest push from Matt ▶ 23:58 Challenging SQL-first vs abstraction contradiction

The questioner challenges dbt's stance, arguing that pushing SQL-first transformations inherently conflicts with building higher-level abstraction layers.

Biggest teaching moment ▶ 6:15 Factory electrification paradigm shift analogy

The guest educates the host on how technological infrastructure shifts operate, using a 30-year historical parallel of factory electrification layout changes to explain Redshift and ELT.

Matt holds his own ▶ 12:11 Host framing dbt through Ruby on Rails abstraction

The host demonstrates deep domain familiarity by introducing a Ruby on Rails software framework analogy to explain dbt's high-leverage abstraction layer.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Company Founding and Distributed Work Culture 1200 Host opens with standard operational questions about founding dates and distributed work. Guest explains dbt's remote culture based on the GitLab handbook.
Origins and Paradigm Shift of the Modern Data Stack 2510 Host prompts guest to outline the steps of the data journey. Guest provides an in-depth explanation of the modern data stack paradigm shift using a factory electrification analogy.
Understanding Data Transformation Through Real-World Examples 2501 Host asks for a practical example of data transformation to ground the discussion for listeners. Guest educates the host using a detailed green onions unit economics calculation example.
Core Functionality of dbt and Abstraction Frameworks 5401 Host offers an insightful comparison between dbt and web frameworks like Ruby on Rails. Guest validates the comparison and explains how software engineering patterns like Git apply to data.
Transitioning from Open-Source dbt Core to dbt Cloud 3401 Host asks how dbt manages the boundary between open source dbt Core and paid dbt Cloud. Guest outlines the separation between stateless SQL compilation and operational scheduling engines.
The Semantic Layer and Ecosystem Vision 2500 Host asks where dbt's multi-year strategic roadmap leads. Guest elaborates on the semantic layer and uses an Apple App Store analogy to frame dbt as ecosystem infrastructure.
Audience Q&A: Polyglot Languages and Machine Learning Workflows 4534 Q&A challenges whether dbt's SQL-first focus contradicts higher abstraction layers. Guest pushes back on the 'SQL maximalism' framing and clarifies the platform's boundaries regarding machine learning.

Statements from this episode (9)

Prediction Not checkable as stated
Tristan Handy: The GitLab handbook will be used for decades
“I'm sure that people will continue to use GitLab for a long time, but I think people will continue to use the GitLab handbook for, like, decades and decades.”
Tristan Handy Jun 28, 2022 ▶ 1:11
Insight
Handy: The original modern data stack was four core layers
“The original modern data stack was four layers. It was data ingestion. How do you get your data from all of your different upstream systems? It was data storage or warehousing, and how do you actually store and compute data? It was transformation. And then it …”
Tristan Handy Jun 28, 2022 ▶ 5:17
Assertion Not checkable as stated
Handy: 100x more people write SQL than Spark or Scala
“There are actually two orders of magnitude more human beings on the planet that can write SQL than can write Spark or Scala or whatever.”
Tristan Handy Jun 28, 2022 ▶ 7:34
Opinion
Handy: Data profession is two decades behind software engineering
“And as a profession, like data is generally Probably two decades behind software engineering in terms of, like, the productivity of practitioners and the level of abstraction and everything.”
Tristan Handy Jun 28, 2022 ▶ 13:47
Assertion Not checkable as stated
Handy: Data teams in 2016 shared SQL files as email attachments
“Back in 2016, the, like, data practitioners were sending each other SQL files as attachments to emails, and that was, like, the way that we worked together.”
Tristan Handy Jun 28, 2022 ▶ 14:37
Assertion Not checkable as stated
Handy: Early-stage VCs doubted data practitioners would learn Git in 2016
“Early stage VCs that I spoke to back in 2016 told me that, like, it wasn't at all clear that data practitioners actually wanted to learn Git.”
Tristan Handy Jun 28, 2022 ▶ 14:51
Disclosure
Handy: dbt Labs aims to power ecosystem infrastructure, not own cataloging
“As we move forward, we're not looking to, like, own cataloging or own whatever, these different categories. We're looking to be the infrastructure that powers this ecosystem, because it turns out that you don't actually want to connect to, like, four different…”
Tristan Handy Jun 28, 2022 ▶ 22:41
Prediction Not checkable as stated
Handy: Data processing will be polyglot with non-SQL abstractions within five years
“We really do think that the future of data processing is polyglot, and I think that if you look in five years, you will find more robust abstractions on top of data, and even in the DBT ecosystem, than SQL.”
Tristan Handy Jun 28, 2022 ▶ 24:51
Prediction Not checkable as stated
Handy: Gap between modern data stack and machine learning will vanish by 2027
“Again, if you look in five years, I think that this distinction will have been sanded over and will not be salient anymore.”
Tristan Handy Jun 28, 2022 ▶ 26:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.