Disclosure certainty 5/5 debate potential 1/5

Gleb Mezhanskiy: A three-line SQL hotfix once crashed Lyft's data platform

Gleb Mezhanskiy · A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy · Sep 12, 2022 · at 2:25

Gleb Mezhanskiy, founder and CEO of Datafold, shares a personal experience from his previous role as a data engineer at Lyft to illustrate how internal code changes break data.

0:00 / 0:16exact quote · 16.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“In my years of data engineer at Lyft, I was unlucky to break down, well, actually blow up the entire data platform by making a three line SQL code hotfix that filtered a little bit more rights than I anticipated.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Gleb Mezhanskiy

Insight
Gleb Mezhanskiy: Data monitoring scales better than manual data testing
“Data testing is basically us writing unit tests for data are not scalable, whereas data monitoring is scalable because we can just automatically track data in the warehouse and automatically provision machine learning to track data for anomalies.”
Gleb Mezhanskiy Sep 12, 2022 ▶ 5:45 A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
Insight
Gleb Mezhanskiy: Data testing gives better signal-to-noise than data monitoring
“The signal to noise in data testing typically tends to be better because we define exactly what is wrong, what is right, versus in monitoring, it's Quite noisy, because we just, it just tells us when data doesn't conform to the historical properties, which doe…”
Gleb Mezhanskiy Sep 12, 2022 ▶ 5:59 A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
Assertion Not publicly verifiable
Mezhanskiy: Datafold's data-diff benchmarks over 1B rows in 5 minutes
“We run some benchmarks, and you can run this on like a twenty-five million row data set in less than 10 seconds, An over one billion row dataset in about five minutes.”
Gleb Mezhanskiy Sep 12, 2022 ▶ 16:37 A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.