Apr 5, 2021 · 33m · mad

Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)

Dave Burgess · 22m spoken Matt Turck · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC fireside chat hosted by Matt Turck, Dave Burgess, Head of Data Engineering at Pinterest, details Pinterest's scale, big data infrastructure, real-time streaming architectures, machine learning platforms, and open-source initiatives like Querybook.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25.1% of the talking time here. How this is scored →

Matt as informed peer 3.7 Guest teaching 4.4 Guest disagreement 0.0 Matt pushing back 0.3
05100:0010:0020:0030:000:10–4:05 · Matt as informed peer 4/10 Dave Burgess's Career Path and Background Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML.4:05–7:32 · Matt as informed peer 3/10 Announcement and Overview of Querybook Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest.7:32–11:40 · Matt as informed peer 5/10 Pinterest Data Engineering Stack and PII Handling Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging.11:40–17:34 · Matt as informed peer 4/10 Machine Learning Use Cases and Infrastructure at Pinterest Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow.17:34–21:47 · Matt as informed peer 4/10 Real-Time Data Streaming with Kafka, Flink, and Druid Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime.21:47–25:54 · Matt as informed peer 3/10 Data Team Organizational Structure at Pinterest Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest.25:54–33:05 · Matt as informed peer 3/10 Audience Q&A: Infrastructure, Monitoring, and Career Advice Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database.0:10–4:05 · Guest teaching 3/10 Dave Burgess's Career Path and Background Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML.4:05–7:32 · Guest teaching 4/10 Announcement and Overview of Querybook Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest.7:32–11:40 · Guest teaching 5/10 Pinterest Data Engineering Stack and PII Handling Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging.11:40–17:34 · Guest teaching 5/10 Machine Learning Use Cases and Infrastructure at Pinterest Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow.17:34–21:47 · Guest teaching 5/10 Real-Time Data Streaming with Kafka, Flink, and Druid Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime.21:47–25:54 · Guest teaching 4/10 Data Team Organizational Structure at Pinterest Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest.25:54–33:05 · Guest teaching 5/10 Audience Q&A: Infrastructure, Monitoring, and Career Advice Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database.0:10–4:05 · Guest disagreement 0/10 Dave Burgess's Career Path and Background Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML.4:05–7:32 · Guest disagreement 0/10 Announcement and Overview of Querybook Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest.7:32–11:40 · Guest disagreement 0/10 Pinterest Data Engineering Stack and PII Handling Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging.11:40–17:34 · Guest disagreement 0/10 Machine Learning Use Cases and Infrastructure at Pinterest Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow.17:34–21:47 · Guest disagreement 0/10 Real-Time Data Streaming with Kafka, Flink, and Druid Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime.21:47–25:54 · Guest disagreement 0/10 Data Team Organizational Structure at Pinterest Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest.25:54–33:05 · Guest disagreement 0/10 Audience Q&A: Infrastructure, Monitoring, and Career Advice Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database.0:10–4:05 · Matt pushing back 0/10 Dave Burgess's Career Path and Background Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML.4:05–7:32 · Matt pushing back 0/10 Announcement and Overview of Querybook Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest.7:32–11:40 · Matt pushing back 1/10 Pinterest Data Engineering Stack and PII Handling Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging.11:40–17:34 · Matt pushing back 0/10 Machine Learning Use Cases and Infrastructure at Pinterest Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow.17:34–21:47 · Matt pushing back 1/10 Real-Time Data Streaming with Kafka, Flink, and Druid Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime.21:47–25:54 · Matt pushing back 0/10 Data Team Organizational Structure at Pinterest Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest.25:54–33:05 · Matt pushing back 0/10 Audience Q&A: Infrastructure, Monitoring, and Career Advice Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 9.2% · guest 90.8%0:00 · Matt 9.2% · guest 90.8%3:00 · Matt 28.5% · guest 71.5%3:00 · Matt 28.5% · guest 71.5%6:00 · Matt 30.7% · guest 69.3%6:00 · Matt 30.7% · guest 69.3%9:00 · Matt 28.9% · guest 71.1%9:00 · Matt 28.9% · guest 71.1%12:00 · Matt 10.5% · guest 89.5%12:00 · Matt 10.5% · guest 89.5%15:00 · Matt 53.3% · guest 46.7%15:00 · Matt 53.3% · guest 46.7%18:00 · Matt 18.3% · guest 81.7%18:00 · Matt 18.3% · guest 81.7%21:00 · Matt 38.5% · guest 61.5%21:00 · Matt 38.5% · guest 61.5%24:00 · Matt 23.2% · guest 76.8%24:00 · Matt 23.2% · guest 76.8%27:00 · Matt 14.7% · guest 85.3%27:00 · Matt 14.7% · guest 85.3%30:00 · Matt 8.6% · guest 91.4%30:00 · Matt 8.6% · guest 91.4%33:00 · Matt 86.5% · guest 13.5%33:00 · Matt 86.5% · guest 13.5%
Sharpest disagreement ▶ 10:22 Dave gently corrects host's assumption about Redshift

In a completely non-combative interview, this represents the mildest pushback where Dave corrects Matt's assumption about using AWS Redshift by clarifying they use Oracle for financial data.

Hardest push from Matt ▶ 19:57 Matt questions reliance on commercial open-source vendors

Matt pushes past the general architecture discussion to ask whether Pinterest depends on third-party commercial vendors like Confluent for Kafka or manages the code bases directly.

Biggest teaching moment ▶ 11:07 Dave explains schema-level PII identification

Dave educates the host on how Pinterest handles PII detection at the source schema level rather than relying solely on downstream scanning tools.

Matt holds his own ▶ 2:12 Matt highlights Yahoo's historical importance in big data

Matt displays deep industry familiarity by pointing out Yahoo's massive influence as an incubator for the modern big data ecosystem.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Dave Burgess's Career Path and Background 4300 Matt demonstrates industry context by highlighting Yahoo's foundational role in the big data ecosystem. Dave warmly agrees and expands on Yahoo's contributions to Hadoop and early ML.
Announcement and Overview of Querybook 3400 Matt asks clarifying questions about Querybook to frame it as a collaboration layer and query execution interface. Dave confirms and explains its role within Pinterest.
Pinterest Data Engineering Stack and PII Handling 5501 Matt demonstrates domain familiarity by prompting a definition of Parquet, referencing author Julien Ledem, and guessing Redshift usage. Dave clarifies they actually use Oracle for financial metrics and details their S3 data lake and PII tagging.
Machine Learning Use Cases and Infrastructure at Pinterest 4500 Matt offers commentary on Pinterest's ad quality and clean user environment before pivoting to ML frameworks. Dave details Pinterest's ML platform for glueing TensorFlow, PyTorch, and MLflow.
Real-Time Data Streaming with Kafka, Flink, and Druid 4501 Matt asks about streaming adoption and presses gently on whether Pinterest relies on commercial vendors like Confluent or manages OSS directly. Dave explains why direct management ensures maximum uptime.
Data Team Organizational Structure at Pinterest 3400 Matt inquires about organizational structure and transitions into rapid-fire ecosystem questions. Dave outlines how platform data engineering supports embedded ML teams across Pinterest.
Audience Q&A: Infrastructure, Monitoring, and Career Advice 3500 Matt moderates audience questions on Druid, monitoring tools, Airflow, and career paths. Dave provides detailed explanations including Pinterest's custom Goku time-series database.

Statements from this episode (14)

Assertion Contradicted
Apache Kafka was built on prior engineering work at Yahoo
“Things like Kafka eventually came out of LinkedIn that was based on work that we had been doing at Yahoo”
Dave Burgess Apr 5, 2021 ▶ 2:38
Assertion Supported
Yahoo used ML and behavioral targeting in ads over 15 years ago
“We were doing machine learning and behavioral targeting back 15 years ago, more than 15 years ago in advertising.”
Dave Burgess Apr 5, 2021 ▶ 2:46
Assertion Not checkable as stated
Pinterest stores more than 400 petabytes of data
“We have more than 400 petabytes of data.”
Dave Burgess Apr 5, 2021 ▶ 8:59
Assertion Not checkable as stated
Pinterest runs about 1,000 parallel A/B experiments at any time
“And so we have about a thousand experiments running in parallel at any point of time.”
Dave Burgess Apr 5, 2021 ▶ 12:15
Assertion Not checkable as stated
Pinterest deploys machine learning across roughly 80 distinct use cases
“So we have about 80 different use cases of machine learning.”
Dave Burgess Apr 5, 2021 ▶ 12:26
Opinion
Pinterest visual search beats Google and Bing in head-to-head tests
“You'll find that Pinterest is usually the one that comes out of top.”
Dave Burgess Apr 5, 2021 ▶ 13:06
Assertion Not checkable as stated
Pinterest platform deploys ML models across thousands of servers within hours
“So we, what we can do is within a few hours, you could do create a model, train a model, and then deploy it to our production service with thousands and thousands of servers by using this MLflow repository.”
Dave Burgess Apr 5, 2021 ▶ 16:30
Assertion Not checkable as stated
Pinterest uses 50 machine learning models in its pin safety pipeline
“And so, there's a, that whole pipeline has got like about 50 different machine learning models in itself.”
Dave Burgess Apr 5, 2021 ▶ 19:01
Disclosure
Pinterest manages open-source data tech in-house to limit downtime
“Yeah, we do. I mean, we're fortunate to have the number of engineers to do that, and one of the reasons why we work directly either with the open source community or the companies that are also working on that is that we want Pinterest to be up all the time, a…”
Dave Burgess Apr 5, 2021 ▶ 20:29
Disclosure
Burgess outlines Pinterest's online serving and batch data engineering stack
“So data engineering at Pinterest comprises the serving part, all the online systems. So that includes the online databases like MySQL and key value stores like RocksDB and HBase, and also includes Druid platform. So we have our online systems. We also have all…”
Dave Burgess Apr 5, 2021 ▶ 22:07
Disclosure
Pinterest embeds machine learning engineers directly within product teams
“With machine learning is actually distributed everywhere. We have Many, many machine learning engineers and many, many use cases. And those mission machine learning engineers are embedded within each organization.”
Dave Burgess Apr 5, 2021 ▶ 22:50
Disclosure
Pinterest wants non-coding data scientists to deploy ML models to production
“And to get to a point where we can have people that are data scientists that don't even code, that can just build models and be able to deploy those to production. So that's really what we want to get to within Pinterest as well.”
Dave Burgess Apr 5, 2021 ▶ 24:22
Assertion Not checkable as stated
Pinterest reduced maintenance costs, latency, and infrastructure expenses by migrating to Druid
“And so we decided to migrate that to Druid. And so the maintenance cost has gone way down. The latency has gone way down. The actual cost of running the infrastructure way down. So really, really happy.”
Dave Burgess Apr 5, 2021 ▶ 27:25
Assertion Not checkable as stated
Pinterest operates over 2,000 production workflows on Apache Airflow
“We're still using the APIs for the existing system so that we could migrate, but we have over 2000 workflows running in production.”
Dave Burgess Apr 5, 2021 ▶ 29:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.