Sep 12, 2022 · 22m · mad

A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy

Gleb Mezhanskiy · 17m spoken Matt Turck · 28s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC, Datafold Founder and CEO Gleb Mezhanskiy introduces DataDiff as a crucial third pillar of data quality, demonstrating how comparing datasets value-by-value in CI/CD workflows and cross-database replications prevents costly data regressions.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.4% of the talking time here. How this is scored →

Matt as informed peer 0.4 Guest teaching 6.6 Guest disagreement 1.0 Matt pushing back 0.2
05100:0010:0020:000:08–3:53 · Matt as informed peer 0/10 Audience Poll and Data Quality Pain Points In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero.3:53–7:05 · Matt as informed peer 0/10 Mainstream Approaches: Data Testing vs. Data Monitoring Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero.7:05–13:33 · Matt as informed peer 0/10 Introducing DataDiff as the Third Pillar of Data Quality Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration.13:33–18:11 · Matt as informed peer 0/10 Demo: Open Source DataDiff for Cross-Database Replication Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero.18:11–22:01 · Matt as informed peer 2/10 Q&A Session with Matt Turck Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response.0:08–3:53 · Guest teaching 6/10 Audience Poll and Data Quality Pain Points In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero.3:53–7:05 · Guest teaching 7/10 Mainstream Approaches: Data Testing vs. Data Monitoring Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero.7:05–13:33 · Guest teaching 7/10 Introducing DataDiff as the Third Pillar of Data Quality Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration.13:33–18:11 · Guest teaching 7/10 Demo: Open Source DataDiff for Cross-Database Replication Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero.18:11–22:01 · Guest teaching 6/10 Q&A Session with Matt Turck Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response.0:08–3:53 · Guest disagreement 1/10 Audience Poll and Data Quality Pain Points In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero.3:53–7:05 · Guest disagreement 1/10 Mainstream Approaches: Data Testing vs. Data Monitoring Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero.7:05–13:33 · Guest disagreement 1/10 Introducing DataDiff as the Third Pillar of Data Quality Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration.13:33–18:11 · Guest disagreement 1/10 Demo: Open Source DataDiff for Cross-Database Replication Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero.18:11–22:01 · Guest disagreement 1/10 Q&A Session with Matt Turck Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response.0:08–3:53 · Matt pushing back 0/10 Audience Poll and Data Quality Pain Points In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero.3:53–7:05 · Matt pushing back 0/10 Mainstream Approaches: Data Testing vs. Data Monitoring Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero.7:05–13:33 · Matt pushing back 0/10 Introducing DataDiff as the Third Pillar of Data Quality Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration.13:33–18:11 · Matt pushing back 0/10 Demo: Open Source DataDiff for Cross-Database Replication Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero.18:11–22:01 · Matt pushing back 1/10 Q&A Session with Matt Turck Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 19.5% · guest 80.5%18:00 · Matt 19.5% · guest 80.5%21:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%
Sharpest disagreement ▶ 2:40 Gleb gently challenges audience assumptions about data breakage sources

Gleb playfully pushes back against the audience's belief that third-party vendors and infrastructure cause most errors, arguing instead that data teams themselves introduce most bugs.

Hardest push from Matt ▶ 18:09 Matt Turck intervenes to manage stage time and steer the Q&A

Matt steps in post-presentation to keep the session strictly on schedule and pivot immediately to founder background questions.

Biggest teaching moment ▶ 19:40 Gleb explains automated data lineage generation via database log compilation

Gleb educates an audience member on how DataFold parses query logs from Snowflake and Databricks to automatically generate end-to-end dependency graphs down to the column level.

Matt holds his own ▶ 18:09 Matt Turck smoothly transitions the presentation into a structured founder interview

Matt takes command of the presentation environment to transition from technical demonstration to startup company metrics and founding story.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Audience Poll and Data Quality Pain Points 0610 In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero.
Mainstream Approaches: Data Testing vs. Data Monitoring 0710 Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero.
Introducing DataDiff as the Third Pillar of Data Quality 0710 Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration.
Demo: Open Source DataDiff for Cross-Database Replication 0710 Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero.
Q&A Session with Matt Turck 2611 Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response.

Statements from this episode (4)

Disclosure
Gleb Mezhanskiy: A three-line SQL hotfix once crashed Lyft's data platform
“In my years of data engineer at Lyft, I was unlucky to break down, well, actually blow up the entire data platform by making a three line SQL code hotfix that filtered a little bit more rights than I anticipated.”
Gleb Mezhanskiy Sep 12, 2022 ▶ 2:25
Insight
Gleb Mezhanskiy: Data monitoring scales better than manual data testing
“Data testing is basically us writing unit tests for data are not scalable, whereas data monitoring is scalable because we can just automatically track data in the warehouse and automatically provision machine learning to track data for anomalies.”
Gleb Mezhanskiy Sep 12, 2022 ▶ 5:45
Insight
Gleb Mezhanskiy: Data testing gives better signal-to-noise than data monitoring
“The signal to noise in data testing typically tends to be better because we define exactly what is wrong, what is right, versus in monitoring, it's Quite noisy, because we just, it just tells us when data doesn't conform to the historical properties, which doe…”
Gleb Mezhanskiy Sep 12, 2022 ▶ 5:59
Assertion Not publicly verifiable
Mezhanskiy: Datafold's data-diff benchmarks over 1B rows in 5 minutes
“We run some benchmarks, and you can run this on like a twenty-five million row data set in less than 10 seconds, An over one billion row dataset in about five minutes.”
Gleb Mezhanskiy Sep 12, 2022 ▶ 16:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.