Dec 11, 2020 · 1h 3m · mad

Fireside Chat: Jeremiah Lowin (Prefect), Tristan Handy (dbt) with Matt Turck (Partner, FirstMark)

Jeremiah Lowin · 25m spoken Tristan Handy · 19m spoken Matt Turck · 9m spoken Jack Cohen · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Host Matt Turck brings together Tristan Handy (Founder/CEO of dbt) and Jeremiah Lowin (Founder/CEO of Prefect) for a Data Driven NYC panel discussing the evolution of the modern data stack, ELT data transformation, workflow orchestration, open-source commercialization, and emerging trends in operational analytics and data governance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.1% of the talking time here. How this is scored →

Matt as informed peer 2.7 Guest teaching 3.6 Guest disagreement 0.5 Matt pushing back 1.0
05100:0015:0030:0045:001:00:002:17–4:42 · Matt as informed peer 0/10 Jeremiah's Perspective on Modern Data Stack Interoperability Matt does not speak in this segment monologue where Jeremiah expands on Tristan's definition of the modern data stack. Host scores are set to zero.4:42–9:02 · Matt as informed peer 5/10 Data Engineering versus Data Science Workflows Matt proposes a binary taxonomy of analytics versus machine learning stacks. Jeremiah politely reframes Matt's model around job status versus data transformation, playfully teasing Matt about his landscape diagram.9:02–11:45 · Matt as informed peer 2/10 The Impact and Evolution of Cloud Data Warehouses Matt asks a introductory question about cloud data warehouses. Tristan educates the room on how Redshift democratized OLAP technology by making it available for $160 per month.11:45–14:08 · Matt as informed peer 1/10 Comparing Data Warehouses and Data Lakes Matt asks the guests to explain the difference between a data lake and a data warehouse. Tristan and Jeremiah explain compute coupling and schema timing analogies.14:08–17:55 · Matt as informed peer 3/10 The Evolution of Data Transformation and ELT Matt asks Tristan to explain the shift from ETL to ELT and requests a practical example. Tristan details how SQL standards evolved and gives a clear breakdown of amortizing Stripe subscription revenue.17:55–20:25 · Matt as informed peer 2/10 Defining the Role and Skill Set of the Data Analyst Matt asks Tristan to define the role and technical skill set of a modern data analyst. Tristan frames analysts as business problem solvers who adopt technology out of necessity.20:25–22:53 · Matt as informed peer 5/10 The Philosophy, Origins, and Growth of dbt Matt demonstrates clear familiarity with dbt's core philosophy of bringing software engineering principles to analysts and highlights their recent Series B funding.22:53–25:41 · Matt as informed peer 3/10 Prefect's Origin Story and Core Mission Matt draws parallels between Tristan and Jeremiah's founder journeys. Jeremiah describes building Prefect to solve his own data science and risk management pain points.25:41–29:02 · Matt as informed peer 1/10 Workflow Automation and the Concept of Negative Engineering Matt prompts Jeremiah to explain workflow automation. Jeremiah introduces his signature concept of negative engineering and positions Prefect as defensive risk management software.29:02–31:29 · Matt as informed peer 4/10 Differentiating Prefect from Apache Airflow Matt asks Jeremiah to compare Prefect directly with Apache Airflow, noting Jeremiah's blog post. Jeremiah explains his background as an Airflow maintainer and why Airflow could not address these new needs.31:29–35:33 · Matt as informed peer 0/10 Audience Q&A: Eliminating Time-Wasting Tasks for Data Teams Jack presents an audience question about time-wasting tasks. Jeremiah highlights incident management while Tristan emphasizes analysts getting cross-functionally blocked. Matt does not host this segment.35:33–43:40 · Matt as informed peer 4/10 Building and Monetizing Open Source Data Companies Matt probes open source monetization strategies. Tristan gently calls out Jeremiah's podcast confidence, while Jeremiah gives a strong thesis on why hosted open source is a flawed business model.43:40–46:46 · Matt as informed peer 3/10 Audience Q&A: Integrating dbt with LookML and Cube.js The panel answers technical audience questions about dbt integration with LookML and GraphQL usage. Matt poses the GraphQL query, prompting Jeremiah to weigh API flexibility against database performance.46:46–55:21 · Matt as informed peer 5/10 Audience Q&A: The Evolving Role of Data Engineers Matt challenges whether data engineers will be automated away and raises future trends like streaming and governance. Tristan points to reverse ETL and operational analytics as the next frontier.55:21–58:02 · Matt as informed peer 1/10 Audience Q&A: Data Masking, PII, and Regulatory Compliance Jack reads an audience question about PII data masking. Tristan discusses warehouse data retention risks and Jeremiah explains Prefect's hybrid metadata-only architecture for compliance.58:02–59:50 · Matt as informed peer 4/10 Audience Q&A: Balancing Data Democratization with Governance Matt asks about balancing data democratization with quality governance. Tristan reframes the premise by drawing an analogy to modern software engineering CI/CD pipelines.2:17–4:42 · Guest teaching 2/10 Jeremiah's Perspective on Modern Data Stack Interoperability Matt does not speak in this segment monologue where Jeremiah expands on Tristan's definition of the modern data stack. Host scores are set to zero.4:42–9:02 · Guest teaching 4/10 Data Engineering versus Data Science Workflows Matt proposes a binary taxonomy of analytics versus machine learning stacks. Jeremiah politely reframes Matt's model around job status versus data transformation, playfully teasing Matt about his landscape diagram.9:02–11:45 · Guest teaching 5/10 The Impact and Evolution of Cloud Data Warehouses Matt asks a introductory question about cloud data warehouses. Tristan educates the room on how Redshift democratized OLAP technology by making it available for $160 per month.11:45–14:08 · Guest teaching 4/10 Comparing Data Warehouses and Data Lakes Matt asks the guests to explain the difference between a data lake and a data warehouse. Tristan and Jeremiah explain compute coupling and schema timing analogies.14:08–17:55 · Guest teaching 4/10 The Evolution of Data Transformation and ELT Matt asks Tristan to explain the shift from ETL to ELT and requests a practical example. Tristan details how SQL standards evolved and gives a clear breakdown of amortizing Stripe subscription revenue.17:55–20:25 · Guest teaching 3/10 Defining the Role and Skill Set of the Data Analyst Matt asks Tristan to define the role and technical skill set of a modern data analyst. Tristan frames analysts as business problem solvers who adopt technology out of necessity.20:25–22:53 · Guest teaching 2/10 The Philosophy, Origins, and Growth of dbt Matt demonstrates clear familiarity with dbt's core philosophy of bringing software engineering principles to analysts and highlights their recent Series B funding.22:53–25:41 · Guest teaching 2/10 Prefect's Origin Story and Core Mission Matt draws parallels between Tristan and Jeremiah's founder journeys. Jeremiah describes building Prefect to solve his own data science and risk management pain points.25:41–29:02 · Guest teaching 5/10 Workflow Automation and the Concept of Negative Engineering Matt prompts Jeremiah to explain workflow automation. Jeremiah introduces his signature concept of negative engineering and positions Prefect as defensive risk management software.29:02–31:29 · Guest teaching 4/10 Differentiating Prefect from Apache Airflow Matt asks Jeremiah to compare Prefect directly with Apache Airflow, noting Jeremiah's blog post. Jeremiah explains his background as an Airflow maintainer and why Airflow could not address these new needs.31:29–35:33 · Guest teaching 3/10 Audience Q&A: Eliminating Time-Wasting Tasks for Data Teams Jack presents an audience question about time-wasting tasks. Jeremiah highlights incident management while Tristan emphasizes analysts getting cross-functionally blocked. Matt does not host this segment.35:33–43:40 · Guest teaching 5/10 Building and Monetizing Open Source Data Companies Matt probes open source monetization strategies. Tristan gently calls out Jeremiah's podcast confidence, while Jeremiah gives a strong thesis on why hosted open source is a flawed business model.43:40–46:46 · Guest teaching 3/10 Audience Q&A: Integrating dbt with LookML and Cube.js The panel answers technical audience questions about dbt integration with LookML and GraphQL usage. Matt poses the GraphQL query, prompting Jeremiah to weigh API flexibility against database performance.46:46–55:21 · Guest teaching 4/10 Audience Q&A: The Evolving Role of Data Engineers Matt challenges whether data engineers will be automated away and raises future trends like streaming and governance. Tristan points to reverse ETL and operational analytics as the next frontier.55:21–58:02 · Guest teaching 3/10 Audience Q&A: Data Masking, PII, and Regulatory Compliance Jack reads an audience question about PII data masking. Tristan discusses warehouse data retention risks and Jeremiah explains Prefect's hybrid metadata-only architecture for compliance.58:02–59:50 · Guest teaching 5/10 Audience Q&A: Balancing Data Democratization with Governance Matt asks about balancing data democratization with quality governance. Tristan reframes the premise by drawing an analogy to modern software engineering CI/CD pipelines.2:17–4:42 · Guest disagreement 0/10 Jeremiah's Perspective on Modern Data Stack Interoperability Matt does not speak in this segment monologue where Jeremiah expands on Tristan's definition of the modern data stack. Host scores are set to zero.4:42–9:02 · Guest disagreement 3/10 Data Engineering versus Data Science Workflows Matt proposes a binary taxonomy of analytics versus machine learning stacks. Jeremiah politely reframes Matt's model around job status versus data transformation, playfully teasing Matt about his landscape diagram.9:02–11:45 · Guest disagreement 0/10 The Impact and Evolution of Cloud Data Warehouses Matt asks a introductory question about cloud data warehouses. Tristan educates the room on how Redshift democratized OLAP technology by making it available for $160 per month.11:45–14:08 · Guest disagreement 0/10 Comparing Data Warehouses and Data Lakes Matt asks the guests to explain the difference between a data lake and a data warehouse. Tristan and Jeremiah explain compute coupling and schema timing analogies.14:08–17:55 · Guest disagreement 0/10 The Evolution of Data Transformation and ELT Matt asks Tristan to explain the shift from ETL to ELT and requests a practical example. Tristan details how SQL standards evolved and gives a clear breakdown of amortizing Stripe subscription revenue.17:55–20:25 · Guest disagreement 0/10 Defining the Role and Skill Set of the Data Analyst Matt asks Tristan to define the role and technical skill set of a modern data analyst. Tristan frames analysts as business problem solvers who adopt technology out of necessity.20:25–22:53 · Guest disagreement 0/10 The Philosophy, Origins, and Growth of dbt Matt demonstrates clear familiarity with dbt's core philosophy of bringing software engineering principles to analysts and highlights their recent Series B funding.22:53–25:41 · Guest disagreement 0/10 Prefect's Origin Story and Core Mission Matt draws parallels between Tristan and Jeremiah's founder journeys. Jeremiah describes building Prefect to solve his own data science and risk management pain points.25:41–29:02 · Guest disagreement 0/10 Workflow Automation and the Concept of Negative Engineering Matt prompts Jeremiah to explain workflow automation. Jeremiah introduces his signature concept of negative engineering and positions Prefect as defensive risk management software.29:02–31:29 · Guest disagreement 1/10 Differentiating Prefect from Apache Airflow Matt asks Jeremiah to compare Prefect directly with Apache Airflow, noting Jeremiah's blog post. Jeremiah explains his background as an Airflow maintainer and why Airflow could not address these new needs.31:29–35:33 · Guest disagreement 0/10 Audience Q&A: Eliminating Time-Wasting Tasks for Data Teams Jack presents an audience question about time-wasting tasks. Jeremiah highlights incident management while Tristan emphasizes analysts getting cross-functionally blocked. Matt does not host this segment.35:33–43:40 · Guest disagreement 3/10 Building and Monetizing Open Source Data Companies Matt probes open source monetization strategies. Tristan gently calls out Jeremiah's podcast confidence, while Jeremiah gives a strong thesis on why hosted open source is a flawed business model.43:40–46:46 · Guest disagreement 0/10 Audience Q&A: Integrating dbt with LookML and Cube.js The panel answers technical audience questions about dbt integration with LookML and GraphQL usage. Matt poses the GraphQL query, prompting Jeremiah to weigh API flexibility against database performance.46:46–55:21 · Guest disagreement 0/10 Audience Q&A: The Evolving Role of Data Engineers Matt challenges whether data engineers will be automated away and raises future trends like streaming and governance. Tristan points to reverse ETL and operational analytics as the next frontier.55:21–58:02 · Guest disagreement 0/10 Audience Q&A: Data Masking, PII, and Regulatory Compliance Jack reads an audience question about PII data masking. Tristan discusses warehouse data retention risks and Jeremiah explains Prefect's hybrid metadata-only architecture for compliance.58:02–59:50 · Guest disagreement 1/10 Audience Q&A: Balancing Data Democratization with Governance Matt asks about balancing data democratization with quality governance. Tristan reframes the premise by drawing an analogy to modern software engineering CI/CD pipelines.2:17–4:42 · Matt pushing back 0/10 Jeremiah's Perspective on Modern Data Stack Interoperability Matt does not speak in this segment monologue where Jeremiah expands on Tristan's definition of the modern data stack. Host scores are set to zero.4:42–9:02 · Matt pushing back 4/10 Data Engineering versus Data Science Workflows Matt proposes a binary taxonomy of analytics versus machine learning stacks. Jeremiah politely reframes Matt's model around job status versus data transformation, playfully teasing Matt about his landscape diagram.9:02–11:45 · Matt pushing back 1/10 The Impact and Evolution of Cloud Data Warehouses Matt asks a introductory question about cloud data warehouses. Tristan educates the room on how Redshift democratized OLAP technology by making it available for $160 per month.11:45–14:08 · Matt pushing back 0/10 Comparing Data Warehouses and Data Lakes Matt asks the guests to explain the difference between a data lake and a data warehouse. Tristan and Jeremiah explain compute coupling and schema timing analogies.14:08–17:55 · Matt pushing back 1/10 The Evolution of Data Transformation and ELT Matt asks Tristan to explain the shift from ETL to ELT and requests a practical example. Tristan details how SQL standards evolved and gives a clear breakdown of amortizing Stripe subscription revenue.17:55–20:25 · Matt pushing back 1/10 Defining the Role and Skill Set of the Data Analyst Matt asks Tristan to define the role and technical skill set of a modern data analyst. Tristan frames analysts as business problem solvers who adopt technology out of necessity.20:25–22:53 · Matt pushing back 1/10 The Philosophy, Origins, and Growth of dbt Matt demonstrates clear familiarity with dbt's core philosophy of bringing software engineering principles to analysts and highlights their recent Series B funding.22:53–25:41 · Matt pushing back 0/10 Prefect's Origin Story and Core Mission Matt draws parallels between Tristan and Jeremiah's founder journeys. Jeremiah describes building Prefect to solve his own data science and risk management pain points.25:41–29:02 · Matt pushing back 0/10 Workflow Automation and the Concept of Negative Engineering Matt prompts Jeremiah to explain workflow automation. Jeremiah introduces his signature concept of negative engineering and positions Prefect as defensive risk management software.29:02–31:29 · Matt pushing back 1/10 Differentiating Prefect from Apache Airflow Matt asks Jeremiah to compare Prefect directly with Apache Airflow, noting Jeremiah's blog post. Jeremiah explains his background as an Airflow maintainer and why Airflow could not address these new needs.31:29–35:33 · Matt pushing back 0/10 Audience Q&A: Eliminating Time-Wasting Tasks for Data Teams Jack presents an audience question about time-wasting tasks. Jeremiah highlights incident management while Tristan emphasizes analysts getting cross-functionally blocked. Matt does not host this segment.35:33–43:40 · Matt pushing back 3/10 Building and Monetizing Open Source Data Companies Matt probes open source monetization strategies. Tristan gently calls out Jeremiah's podcast confidence, while Jeremiah gives a strong thesis on why hosted open source is a flawed business model.43:40–46:46 · Matt pushing back 0/10 Audience Q&A: Integrating dbt with LookML and Cube.js The panel answers technical audience questions about dbt integration with LookML and GraphQL usage. Matt poses the GraphQL query, prompting Jeremiah to weigh API flexibility against database performance.46:46–55:21 · Matt pushing back 3/10 Audience Q&A: The Evolving Role of Data Engineers Matt challenges whether data engineers will be automated away and raises future trends like streaming and governance. Tristan points to reverse ETL and operational analytics as the next frontier.55:21–58:02 · Matt pushing back 0/10 Audience Q&A: Data Masking, PII, and Regulatory Compliance Jack reads an audience question about PII data masking. Tristan discusses warehouse data retention risks and Jeremiah explains Prefect's hybrid metadata-only architecture for compliance.58:02–59:50 · Matt pushing back 1/10 Audience Q&A: Balancing Data Democratization with Governance Matt asks about balancing data democratization with quality governance. Tristan reframes the premise by drawing an analogy to modern software engineering CI/CD pipelines.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 23.6% · guest 76.4%0:00 · Matt 23.6% · guest 76.4%3:00 · Matt 29.7% · guest 70.3%3:00 · Matt 29.7% · guest 70.3%6:00 · Matt 2.4% · guest 97.6%6:00 · Matt 2.4% · guest 97.6%9:00 · Matt 29.9% · guest 70.1%9:00 · Matt 29.9% · guest 70.1%12:00 · Matt 18.2% · guest 81.8%12:00 · Matt 18.2% · guest 81.8%15:00 · Matt 5.8% · guest 94.2%15:00 · Matt 5.8% · guest 94.2%18:00 · Matt 32.9% · guest 67.1%18:00 · Matt 32.9% · guest 67.1%21:00 · Matt 20.4% · guest 79.6%21:00 · Matt 20.4% · guest 79.6%24:00 · Matt 12.8% · guest 87.2%24:00 · Matt 12.8% · guest 87.2%27:00 · Matt 11.9% · guest 88.1%27:00 · Matt 11.9% · guest 88.1%30:00 · Matt 8.9% · guest 91.1%30:00 · Matt 8.9% · guest 91.1%33:00 · Matt 15.4% · guest 84.6%33:00 · Matt 15.4% · guest 84.6%36:00 · Matt 5.2% · guest 94.8%36:00 · Matt 5.2% · guest 94.8%39:00 · Matt 6.2% · guest 93.8%39:00 · Matt 6.2% · guest 93.8%42:00 · Matt 1.4% · guest 98.6%42:00 · Matt 1.4% · guest 98.6%45:00 · Matt 31.1% · guest 68.9%45:00 · Matt 31.1% · guest 68.9%48:00 · Matt 32.5% · guest 67.5%48:00 · Matt 32.5% · guest 67.5%51:00 · Matt 0% · guest 100%51:00 · Matt 0% · guest 100%54:00 · Matt 15.8% · guest 84.2%54:00 · Matt 15.8% · guest 84.2%57:00 · Matt 23.6% · guest 76.4%57:00 · Matt 23.6% · guest 76.4%1:00:00 · Matt 18% · guest 82%1:00:00 · Matt 18% · guest 82%1:03:00 · Matt 86.5% · guest 13.5%1:03:00 · Matt 86.5% · guest 13.5%
Sharpest disagreement ▶ 36:30 Tristan challenges Jeremiah's podcast posture

Tristan playfully calls out Jeremiah regarding his appearance on Invest Like the Best, noting that Jeremiah presented his commercialization thesis as if he had all the answers when reality is much less certain.

Hardest push from Matt ▶ 8:52 Matt defends his two-box landscape taxonomy

When Jeremiah rejects Matt's proposed categorization of analytics versus data science, Matt pushes back directly by asking if Jeremiah believes reality cannot be fitted into structured boxes.

Biggest teaching moment ▶ 41:45 Jeremiah critiques managed open source business models

Jeremiah educates the audience on why selling managed hosting for open source code is a weak business model, arguing that successful commercial entities must offer distinct value beyond running cloud servers.

Matt holds his own ▶ 20:25 Matt articulates dbt's philosophy and recent momentum

Matt displays clear domain expertise by articulating dbt's core ethos of empowering analysts with software engineering habits while correctly referencing their rapid Series B financing timeline.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Jeremiah's Perspective on Modern Data Stack Interoperability 0200 Matt does not speak in this segment monologue where Jeremiah expands on Tristan's definition of the modern data stack. Host scores are set to zero.
Data Engineering versus Data Science Workflows 5434 Matt proposes a binary taxonomy of analytics versus machine learning stacks. Jeremiah politely reframes Matt's model around job status versus data transformation, playfully teasing Matt about his landscape diagram.
The Impact and Evolution of Cloud Data Warehouses 2501 Matt asks a introductory question about cloud data warehouses. Tristan educates the room on how Redshift democratized OLAP technology by making it available for $160 per month.
Comparing Data Warehouses and Data Lakes 1400 Matt asks the guests to explain the difference between a data lake and a data warehouse. Tristan and Jeremiah explain compute coupling and schema timing analogies.
The Evolution of Data Transformation and ELT 3401 Matt asks Tristan to explain the shift from ETL to ELT and requests a practical example. Tristan details how SQL standards evolved and gives a clear breakdown of amortizing Stripe subscription revenue.
Defining the Role and Skill Set of the Data Analyst 2301 Matt asks Tristan to define the role and technical skill set of a modern data analyst. Tristan frames analysts as business problem solvers who adopt technology out of necessity.
The Philosophy, Origins, and Growth of dbt 5201 Matt demonstrates clear familiarity with dbt's core philosophy of bringing software engineering principles to analysts and highlights their recent Series B funding.
Prefect's Origin Story and Core Mission 3200 Matt draws parallels between Tristan and Jeremiah's founder journeys. Jeremiah describes building Prefect to solve his own data science and risk management pain points.
Workflow Automation and the Concept of Negative Engineering 1500 Matt prompts Jeremiah to explain workflow automation. Jeremiah introduces his signature concept of negative engineering and positions Prefect as defensive risk management software.
Differentiating Prefect from Apache Airflow 4411 Matt asks Jeremiah to compare Prefect directly with Apache Airflow, noting Jeremiah's blog post. Jeremiah explains his background as an Airflow maintainer and why Airflow could not address these new needs.
Audience Q&A: Eliminating Time-Wasting Tasks for Data Teams 0300 Jack presents an audience question about time-wasting tasks. Jeremiah highlights incident management while Tristan emphasizes analysts getting cross-functionally blocked. Matt does not host this segment.
Building and Monetizing Open Source Data Companies 4533 Matt probes open source monetization strategies. Tristan gently calls out Jeremiah's podcast confidence, while Jeremiah gives a strong thesis on why hosted open source is a flawed business model.
Audience Q&A: Integrating dbt with LookML and Cube.js 3300 The panel answers technical audience questions about dbt integration with LookML and GraphQL usage. Matt poses the GraphQL query, prompting Jeremiah to weigh API flexibility against database performance.
Audience Q&A: The Evolving Role of Data Engineers 5403 Matt challenges whether data engineers will be automated away and raises future trends like streaming and governance. Tristan points to reverse ETL and operational analytics as the next frontier.
Audience Q&A: Data Masking, PII, and Regulatory Compliance 1300 Jack reads an audience question about PII data masking. Tristan discusses warehouse data retention risks and Jeremiah explains Prefect's hybrid metadata-only architecture for compliance.
Audience Q&A: Balancing Data Democratization with Governance 4511 Matt asks about balancing data democratization with quality governance. Tristan reframes the premise by drawing an analogy to modern software engineering CI/CD pipelines.

Statements from this episode (32)

Insight
Handy defines the modern data stack across four distinct functional layers
“When we talk about the modern data stack, we think about what's really four layers. So there's data ingestion. There's the data warehouse. There's data transformation or like taking all that raw data and like turning it into something valuable. And then there'…”
Tristan Handy Dec 11, 2020 ▶ 0:54
Assertion Not checkable as stated
Handy: Amazon Redshift's launch sparked a rewrite of data stack products
“Those set of technologies have really been completely rebuilt. I think over the past seven years, really, it was like the introduction of Amazon redshift with In, in, in, that sparked kind of a rewrite of all of the products in that space.”
Tristan Handy Dec 11, 2020 ▶ 1:18
Insight
Lowin: Modern tools let analysts query data warehouses directly over static CSVs
“If we think back to BI tools, say five, certainly 10 years ago, you would sort of beg for a CSV and God help you if the data, if the insight you weren't, the insight you were looking for was not in that CSV, you were sort of screwed. But now thanks to, you kno…”
Jeremiah Lowin Dec 11, 2020 ▶ 4:06
Disclosure
Handy: dbt serves batch data workloads, not sub-100ms real-time applications
“There's a lot of data science that's done Internally where like you're doing predictive analysis that, you know, the data latency can be a day or it can be like last quarter's data is fine. And those data science workloads are often run on the exact same stack…”
Tristan Handy Dec 11, 2020 ▶ 7:30
Assertion Partly supported
Handy: Amazon Redshift introduced affordable $160-per-month OLAP databases
“And so Amazon Redshift in, for the first time released an OLAP database, a database designed for large scale query processing that you could purchase for 160 dollars a month.”
Tristan Handy Dec 11, 2020 ▶ 10:39
Insight
Handy: Data lakes offer more flexibility but require more effort
“Where we are today, the data lake can kind of do anything. But it also probably takes more work to do anything. Whereas the data warehouse is, has a more constrained set of use cases, but it is much easier to get up and running for those constraints set of use…”
Tristan Handy Dec 11, 2020 ▶ 12:50
Insight
Jeremiah Lowin: Every data system eventually requires defining a schema
“You're going to have a schema. It's just a question of whether you define it upfront or you figure it out later.”
Jeremiah Lowin Dec 11, 2020 ▶ 13:21
Assertion Not checkable as stated
Handy: Data transformation has shifted inside data warehouses using SQL
“What's happening now or over the past, you know, five or so years, the, Transformation step now happens inside the data warehouse and it happens in SQL. And because of that, it is now accessible to a dramatically larger number of people.”
Tristan Handy Dec 11, 2020 ▶ 15:57
Insight
Handy: Data analysts learn technology to solve problems without identifying as technologists
“I think that a data analyst is somebody who answers business questions with data and they have a, they frequently will have a business or econ degree. They like are interested in solving business problems, but they're also not afraid of technology and they oft…”
Tristan Handy Dec 11, 2020 ▶ 18:28
Disclosure
Handy: dbt's philosophy empowers data analysts to think like software engineers
“Yes, that is completely true.”
Tristan Handy Dec 11, 2020 ▶ 20:25
Insight
Handy: Depending on data engineering backlogs destroys data analyst productivity
“That is like the death of the data analyst as like a productive member of your team. They will get frustrated. They will leave. They, they're, they don't have great career paths. All of these like negative outcomes.”
Tristan Handy Dec 11, 2020 ▶ 21:16
Disclosure
Handy: dbt originated as an internal tool for a consulting firm
“First time analytics, my company that is the maintainers of dbt we started as a consulting business and dbt was the tool that I wanted to be able to do this work.”
Tristan Handy Dec 11, 2020 ▶ 21:39
Assertion Supported
Handy: The dbt community reached over 8,000 members by late 2020
“We've built this community of over 8000 people today who you know, have bought into this as like the way that they want their careers to look.”
Tristan Handy Dec 11, 2020 ▶ 22:01
Prediction Held up
Lowin: PyMC3 developers will revitalize the Theano open-source project
“Because I think the IMC three folks are going to sort of revitalize it, which I think is awesome.”
Jeremiah Lowin Dec 11, 2020 ▶ 24:42
Disclosure
Lowin: Prefect acts as insurance, delivering value when data pipelines break
“That's why we frequently describe our product as an insurance product is because we deliver value mainly when things go wrong. And this will sound crazy maybe to say, but if everything goes the way one of our users expected it to go, they really don't Need our…”
Jeremiah Lowin Dec 11, 2020 ▶ 28:21
Insight
Lowin: Negative engineering is defensive work ensuring code actually runs
“And so we've named this problem, the negative engineering problem, because it's not about what you're trying to achieve. It's about all the defensive work you have to do to make sure that it actually took place, in fact, took place.”
Jeremiah Lowin Dec 11, 2020 ▶ 28:50
What-if
Lowin would have built Prefect within Apache Airflow if possible
“And if I were able to do the things that Prefect does in airflow, I would have.”
Jeremiah Lowin Dec 11, 2020 ▶ 29:43
Assertion Not checkable as stated
Lowin: Prefect's functionality is a superset of Apache Airflow's
“Now it happens that our functionality is a superset of Airflow's functionality.”
Jeremiah Lowin Dec 11, 2020 ▶ 30:03
Insight
Lowin: Saving time during outages offers higher leverage than routine orchestration
“If you imagine saving someone one hour orchestrating something, just getting that code to run versus saving a team one hour in a production crisis incident, the leverage of that hour, it's the same hour, but the leverage, the impact, the emotional burden, the …”
Jeremiah Lowin Dec 11, 2020 ▶ 32:20
Assertion Not checkable as stated
Handy: Data analysts spend 50% of their time blocked by engineering
“Like if you take a team at a like very good, like the data forward organization, a team of 20 data analysts, my guess is that they spend 50% of their time blocked.”
Tristan Handy Dec 11, 2020 ▶ 35:12
Assertion Not checkable as stated
Lowin: Very few open-source projects turn into successful commercial businesses
“The number of successful companies that have actually iterated on this open source, The number is very, very, very small, and it's especially small relative to the denominator, which is much larger because the cost of entering the market is so low.”
Jeremiah Lowin Dec 11, 2020 ▶ 40:12
Opinion
Lowin: Managed hosting for open-source tools is a bad business model
“I think that that is a bad business model. And I sort of, I'm on the record of that. And for a very simple reason, which is that A company that is solving a problem needs to find the correct way to express its knowledge and its solution to that problem.”
Jeremiah Lowin Dec 11, 2020 ▶ 41:52
Opinion
Handy: Batch transformation is not LookML's core focus or innovation
“The part of LookML that does batch based transformation is, is not the primary focus of their product. It's not like the core innovation that they have.”
Tristan Handy Dec 11, 2020 ▶ 44:47
Insight
Lowin: GraphQL is fabulous for exploratory work and well-defined queries
“What we've found is that it's fabulous for exploratory work. Or ironically well defined queries, which is sort of the opposite of the purpose of GraphQL.”
Jeremiah Lowin Dec 11, 2020 ▶ 46:14
Insight
Handy: Data engineers should focus on scalability, not business logic
“Data engineers shouldn't want to be spending their days expressing business logic. Like how do you amortize revenue across multiple periods? They are technologists. They should want to be thinking about platforms and scalability and all of these like hard tech…”
Tristan Handy Dec 11, 2020 ▶ 47:44
Opinion
Handy: Stitch Fix's 2016 blog post best frames the data engineering role
“There's a guy at Stitch Fix who wrote a blog post called Engineers Shouldn't Write ETL back in 2016, and I still think that that is the best point of view on this topic today.”
Tristan Handy Dec 11, 2020 ▶ 48:10
Disclosure
Lowin: Mapping is Prefect's most popular feature
“It's by far the most popular feature of Prefect. And we can see that in our data. And that's because it solves a real problem that I experienced with data science, which I have a bunch of things I need to work with, but I don't know about it until Until runtim…”
Jeremiah Lowin Dec 11, 2020 ▶ 51:13
Prediction Not checkable as stated
Handy: Reverse ETL will automate workflows and multiply the data market opportunity
“There's a class of tools that takes the data that is in your data warehouse and pushes it back to operational systems. And I think that you will, as soon as that starts happening, you can automate the entire process and all of these technologies become, you kn…”
Tristan Handy Dec 11, 2020 ▶ 53:48
Insight
Handy: Once sensitive data lands in a warehouse, it resists elimination
“If you land data in your warehouse, it is very hard to ever have it go away completely.”
Tristan Handy Dec 11, 2020 ▶ 56:04
Assertion Not checkable as stated
Handy: Turnkey solutions for in-flight data masking did not exist in 2020
“The easy solution to this, I don't think by and large exists yet.”
Tristan Handy Dec 11, 2020 ▶ 56:53
Assertion Contradicted
Handy: Software engineers pushing to production grew over 10x in a decade
“The population of software engineers that exist today and push code to production applications. That has, I don't know what, oh, more than 10 X in the past 10 years.”
Tristan Handy Dec 11, 2020 ▶ 58:35
Insight
Handy: CI/CD guardrails enable both software democratization and governance
“What you do is you have mature CICD processes. You have like DevOps workflows that, so you like build these guardrails that create high quality processes around code releases to production where you have your cake and eat it too. You have democratization. You …”
Tristan Handy Dec 11, 2020 ▶ 59:25
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.