Sep 12, 2022 · 22m · mad
A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At Data Driven NYC, Datafold Founder and CEO Gleb Mezhanskiy introduces DataDiff as a crucial third pillar of data quality, demonstrating how comparing datasets value-by-value in CI/CD workflows and cross-database replications prevents costly data regressions.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Gleb playfully pushes back against the audience's belief that third-party vendors and infrastructure cause most errors, arguing instead that data teams themselves introduce most bugs.
Hardest push from Matt ▶ 18:09 Matt Turck intervenes to manage stage time and steer the Q&AMatt steps in post-presentation to keep the session strictly on schedule and pivot immediately to founder background questions.
Biggest teaching moment ▶ 19:40 Gleb explains automated data lineage generation via database log compilationGleb educates an audience member on how DataFold parses query logs from Snowflake and Databricks to automatically generate end-to-end dependency graphs down to the column level.
Matt holds his own ▶ 18:09 Matt Turck smoothly transitions the presentation into a structured founder interviewMatt takes command of the presentation environment to transition from technical demonstration to startup company metrics and founding story.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Audience Poll and Data Quality Pain Points | 0 | 6 | 1 | 0 | In this opening presentation segment, Gleb conducts an audience poll and shares personal technical anecdotes about breaking data infrastructure at Lyft. The host is not present during this monologue, so host scores are zero. | |
| Mainstream Approaches: Data Testing vs. Data Monitoring | 0 | 7 | 1 | 0 | Gleb provides a detailed breakdown comparing data testing with data monitoring, explaining trade-offs in signal-to-noise and scalability. Because this is a monologue presentation, host scores remain zero. | |
| Introducing DataDiff as the Third Pillar of Data Quality | 0 | 7 | 1 | 0 | Gleb presents DataDiff as a solution and demonstrates how source code changes propagate downstream into BI and reverse ETL tools. The segment is a solo product demonstration. | |
| Demo: Open Source DataDiff for Cross-Database Replication | 0 | 7 | 1 | 0 | Gleb demonstrates open-source DataDiff for cross-database replication, explaining high-speed hashing and binary search algorithms. As a monologue, host scores are set to zero. | |
| Q&A Session with Matt Turck | 2 | 6 | 1 | 1 | Matt Turck facilitates a friendly Q&A session asking standard background questions while audience members inquire about data lineage and discovery. Gleb clearly explains data log parsing mechanisms in response. |