Apr 5, 2021 · 33m · mad
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC fireside chat hosted by Matt Turck, Dave Burgess, Head of Data Engineering at Pinterest, details Pinterest's scale, big data infrastructure, real-time streaming architectures, machine learning platforms, and open-source initiatives like Querybook.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25.1% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
In a completely non-combative interview, this represents the mildest pushback where Dave corrects Matt's assumption about using AWS Redshift by clarifying they use Oracle for financial data.
Hardest push from Matt ▶ 19:57 Matt questions reliance on commercial open-source vendorsMatt pushes past the general architecture discussion to ask whether Pinterest depends on third-party commercial vendors like Confluent for Kafka or manages the code bases directly.
Biggest teaching moment ▶ 11:07 Dave explains schema-level PII identificationDave educates the host on how Pinterest handles PII detection at the source schema level rather than relying solely on downstream scanning tools.
Matt holds his own ▶ 2:12 Matt highlights Yahoo's historical importance in big dataMatt displays deep industry familiarity by pointing out Yahoo's massive influence as an incubator for the modern big data ecosystem.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Dave Burgess's Career Path and Background | 4 | 3 | 0 | 0 | Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML. | |
| Announcement and Overview of Querybook | 3 | 4 | 0 | 0 | Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest. | |
| Pinterest Data Engineering Stack and PII Handling | 5 | 5 | 0 | 1 | Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging. | |
| Machine Learning Use Cases and Infrastructure at Pinterest | 4 | 5 | 0 | 0 | Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow. | |
| Real-Time Data Streaming with Kafka, Flink, and Druid | 4 | 5 | 0 | 1 | Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime. | |
| Data Team Organizational Structure at Pinterest | 3 | 4 | 0 | 0 | Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest. | |
| Audience Q&A: Infrastructure, Monitoring, and Career Advice | 3 | 5 | 0 | 0 | Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database. |