Feb 1, 2021 · 27m · mad

Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark)

Alok Gupta · 18m spoken Matt Turck · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC fireside chat hosted by Matt Turck, DoorDash Head of Data Science & ML Alok Gupta discusses how DoorDash leverages advanced machine learning, robust feature store infrastructure, and specialized governance models to power its multi-sided logistics marketplace. He also shares organizational strategies for scaling data science teams and adapting models to major market disruptions like COVID-19.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25.1% of the talking time here. How this is scored →

Matt as informed peer 3.6 Guest teaching 2.8 Guest disagreement 0.5 Matt pushing back 1.1
05100:0010:0020:000:08–3:33 · Matt as informed peer 4/10 DoorDash Mission and Marketplace Dynamics Matt demonstrates understanding of marketplace dynamics by asking if DoorDash is a three-sided marketplace and later framing it as a centralized software brain allocating resources. Alok reframes food delivery as three-sided but grocery as four-sided with pickers, and reframes the software brain as white-label logistics infrastructure similar to AWS.3:33–8:41 · Matt as informed peer 6/10 ML Tech Stack and Model Lifecycle Matt demonstrates strong preparation by citing DoorDash's engineering blog and naming specific ML libraries like XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch. Alok explains how DoorDash standardized onto LightGBM and PyTorch for centralized platform migration and describes their code-commit model training workflow.8:41–13:04 · Matt as informed peer 4/10 Feature Store Architecture and Industry Insights Matt asks Alok to define features for the audience and prompts a comparison of DoorDash's ML stack with Airbnb and Lyft. Alok educates on feature store architecture and common industry hurdles like offline-online data drift and unifying batch versus real-time predictions.13:04–16:53 · Matt as informed peer 3/10 Data Science Team Growth and Interview Process Matt asks about team scale, organizational structure, and specifically probes how DoorDash evaluates business skills in interviews. Alok details how they prioritize business impact over building complex algorithms and borrow analytics and PM style case interviews.16:53–20:26 · Matt as informed peer 5/10 ML Governance and Machine Learning Council Matt articulates a well-informed query regarding how COVID-19 altered the baseline assumptions of predictive models. Alok details their governance approach via the Machine Learning Council and explains tactical fixes like capping predictions and shortening historical lookback windows.20:26–25:01 · Matt as informed peer 5/10 Audience Q&A on ML Infrastructure and Modeling Matt fields audience questions and pushes Alok to clarify whether DoorDash operates parallel stacks between Snowflake and Databricks/S3. Alok clarifies the distinction between Snowflake as the data warehouse and Databricks running on S3 while pulling data from Snowflake.25:01–27:23 · Matt as informed peer 2/10 Career Guidance for Data Scientists and New Graduates Matt passes along audience questions regarding transitioning from academia to industry and advice for new graduates. Alok offers standard career guidance on adopting a business-first mindset and building a T-shaped technical skill set.27:23–27:37 · Matt as informed peer 0/10 Fireside Chat Conclusion and Transition Brief wrap-up segment where Matt thanks Alok and concludes the chat.0:08–3:33 · Guest teaching 3/10 DoorDash Mission and Marketplace Dynamics Matt demonstrates understanding of marketplace dynamics by asking if DoorDash is a three-sided marketplace and later framing it as a centralized software brain allocating resources. Alok reframes food delivery as three-sided but grocery as four-sided with pickers, and reframes the software brain as white-label logistics infrastructure similar to AWS.3:33–8:41 · Guest teaching 3/10 ML Tech Stack and Model Lifecycle Matt demonstrates strong preparation by citing DoorDash's engineering blog and naming specific ML libraries like XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch. Alok explains how DoorDash standardized onto LightGBM and PyTorch for centralized platform migration and describes their code-commit model training workflow.8:41–13:04 · Guest teaching 4/10 Feature Store Architecture and Industry Insights Matt asks Alok to define features for the audience and prompts a comparison of DoorDash's ML stack with Airbnb and Lyft. Alok educates on feature store architecture and common industry hurdles like offline-online data drift and unifying batch versus real-time predictions.13:04–16:53 · Guest teaching 3/10 Data Science Team Growth and Interview Process Matt asks about team scale, organizational structure, and specifically probes how DoorDash evaluates business skills in interviews. Alok details how they prioritize business impact over building complex algorithms and borrow analytics and PM style case interviews.16:53–20:26 · Guest teaching 3/10 ML Governance and Machine Learning Council Matt articulates a well-informed query regarding how COVID-19 altered the baseline assumptions of predictive models. Alok details their governance approach via the Machine Learning Council and explains tactical fixes like capping predictions and shortening historical lookback windows.20:26–25:01 · Guest teaching 4/10 Audience Q&A on ML Infrastructure and Modeling Matt fields audience questions and pushes Alok to clarify whether DoorDash operates parallel stacks between Snowflake and Databricks/S3. Alok clarifies the distinction between Snowflake as the data warehouse and Databricks running on S3 while pulling data from Snowflake.25:01–27:23 · Guest teaching 2/10 Career Guidance for Data Scientists and New Graduates Matt passes along audience questions regarding transitioning from academia to industry and advice for new graduates. Alok offers standard career guidance on adopting a business-first mindset and building a T-shaped technical skill set.27:23–27:37 · Guest teaching 0/10 Fireside Chat Conclusion and Transition Brief wrap-up segment where Matt thanks Alok and concludes the chat.0:08–3:33 · Guest disagreement 1/10 DoorDash Mission and Marketplace Dynamics Matt demonstrates understanding of marketplace dynamics by asking if DoorDash is a three-sided marketplace and later framing it as a centralized software brain allocating resources. Alok reframes food delivery as three-sided but grocery as four-sided with pickers, and reframes the software brain as white-label logistics infrastructure similar to AWS.3:33–8:41 · Guest disagreement 1/10 ML Tech Stack and Model Lifecycle Matt demonstrates strong preparation by citing DoorDash's engineering blog and naming specific ML libraries like XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch. Alok explains how DoorDash standardized onto LightGBM and PyTorch for centralized platform migration and describes their code-commit model training workflow.8:41–13:04 · Guest disagreement 0/10 Feature Store Architecture and Industry Insights Matt asks Alok to define features for the audience and prompts a comparison of DoorDash's ML stack with Airbnb and Lyft. Alok educates on feature store architecture and common industry hurdles like offline-online data drift and unifying batch versus real-time predictions.13:04–16:53 · Guest disagreement 1/10 Data Science Team Growth and Interview Process Matt asks about team scale, organizational structure, and specifically probes how DoorDash evaluates business skills in interviews. Alok details how they prioritize business impact over building complex algorithms and borrow analytics and PM style case interviews.16:53–20:26 · Guest disagreement 0/10 ML Governance and Machine Learning Council Matt articulates a well-informed query regarding how COVID-19 altered the baseline assumptions of predictive models. Alok details their governance approach via the Machine Learning Council and explains tactical fixes like capping predictions and shortening historical lookback windows.20:26–25:01 · Guest disagreement 1/10 Audience Q&A on ML Infrastructure and Modeling Matt fields audience questions and pushes Alok to clarify whether DoorDash operates parallel stacks between Snowflake and Databricks/S3. Alok clarifies the distinction between Snowflake as the data warehouse and Databricks running on S3 while pulling data from Snowflake.25:01–27:23 · Guest disagreement 0/10 Career Guidance for Data Scientists and New Graduates Matt passes along audience questions regarding transitioning from academia to industry and advice for new graduates. Alok offers standard career guidance on adopting a business-first mindset and building a T-shaped technical skill set.27:23–27:37 · Guest disagreement 0/10 Fireside Chat Conclusion and Transition Brief wrap-up segment where Matt thanks Alok and concludes the chat.0:08–3:33 · Matt pushing back 1/10 DoorDash Mission and Marketplace Dynamics Matt demonstrates understanding of marketplace dynamics by asking if DoorDash is a three-sided marketplace and later framing it as a centralized software brain allocating resources. Alok reframes food delivery as three-sided but grocery as four-sided with pickers, and reframes the software brain as white-label logistics infrastructure similar to AWS.3:33–8:41 · Matt pushing back 2/10 ML Tech Stack and Model Lifecycle Matt demonstrates strong preparation by citing DoorDash's engineering blog and naming specific ML libraries like XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch. Alok explains how DoorDash standardized onto LightGBM and PyTorch for centralized platform migration and describes their code-commit model training workflow.8:41–13:04 · Matt pushing back 1/10 Feature Store Architecture and Industry Insights Matt asks Alok to define features for the audience and prompts a comparison of DoorDash's ML stack with Airbnb and Lyft. Alok educates on feature store architecture and common industry hurdles like offline-online data drift and unifying batch versus real-time predictions.13:04–16:53 · Matt pushing back 1/10 Data Science Team Growth and Interview Process Matt asks about team scale, organizational structure, and specifically probes how DoorDash evaluates business skills in interviews. Alok details how they prioritize business impact over building complex algorithms and borrow analytics and PM style case interviews.16:53–20:26 · Matt pushing back 1/10 ML Governance and Machine Learning Council Matt articulates a well-informed query regarding how COVID-19 altered the baseline assumptions of predictive models. Alok details their governance approach via the Machine Learning Council and explains tactical fixes like capping predictions and shortening historical lookback windows.20:26–25:01 · Matt pushing back 3/10 Audience Q&A on ML Infrastructure and Modeling Matt fields audience questions and pushes Alok to clarify whether DoorDash operates parallel stacks between Snowflake and Databricks/S3. Alok clarifies the distinction between Snowflake as the data warehouse and Databricks running on S3 while pulling data from Snowflake.25:01–27:23 · Matt pushing back 0/10 Career Guidance for Data Scientists and New Graduates Matt passes along audience questions regarding transitioning from academia to industry and advice for new graduates. Alok offers standard career guidance on adopting a business-first mindset and building a T-shaped technical skill set.27:23–27:37 · Matt pushing back 0/10 Fireside Chat Conclusion and Transition Brief wrap-up segment where Matt thanks Alok and concludes the chat.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 37.1% · guest 62.9%0:00 · Matt 37.1% · guest 62.9%3:00 · Matt 36.1% · guest 63.9%3:00 · Matt 36.1% · guest 63.9%6:00 · Matt 10.6% · guest 89.4%6:00 · Matt 10.6% · guest 89.4%9:00 · Matt 11.9% · guest 88.1%9:00 · Matt 11.9% · guest 88.1%12:00 · Matt 9.2% · guest 90.8%12:00 · Matt 9.2% · guest 90.8%15:00 · Matt 17.1% · guest 82.9%15:00 · Matt 17.1% · guest 82.9%18:00 · Matt 28.2% · guest 71.8%18:00 · Matt 28.2% · guest 71.8%21:00 · Matt 46.4% · guest 53.6%21:00 · Matt 46.4% · guest 53.6%24:00 · Matt 28.2% · guest 71.8%24:00 · Matt 28.2% · guest 71.8%27:00 · Matt 43.7% · guest 56.3%27:00 · Matt 43.7% · guest 56.3%
Sharpest disagreement ▶ 21:22 Clarifying data warehouse architecture misconception

Alok politely corrects the premise of an audience question relayed by Matt, clarifying that Databricks runs on S3 rather than directly on top of Snowflake.

Hardest push from Matt ▶ 21:46 Host drilling into parallel data stack layout

Matt presses Alok on whether DoorDash maintains two parallel data stacks for BI and ML, seeking clear structural separation between their Snowflake warehouse and S3/Databricks pipeline.

Biggest teaching moment ▶ 11:08 Explaining the holy grail problems of ML infrastructure

Alok educates Matt on the core technical tensions shared across tech giants, such as online versus offline feature drift and reconciling batch versus real-time prediction pipelines.

Matt holds his own ▶ 5:29 Listing specific ML frameworks from DoorDash tech blog

Matt displays thorough technical preparation by naming XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch, pressing Alok on how DoorDash narrowed down its model portfolio.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
DoorDash Mission and Marketplace Dynamics 4311 Matt demonstrates understanding of marketplace dynamics by asking if DoorDash is a three-sided marketplace and later framing it as a centralized software brain allocating resources. Alok reframes food delivery as three-sided but grocery as four-sided with pickers, and reframes the software brain as white-label logistics infrastructure similar to AWS.
ML Tech Stack and Model Lifecycle 6312 Matt demonstrates strong preparation by citing DoorDash's engineering blog and naming specific ML libraries like XGBoost, LightGBM, CatBoost, TensorFlow, and PyTorch. Alok explains how DoorDash standardized onto LightGBM and PyTorch for centralized platform migration and describes their code-commit model training workflow.
Feature Store Architecture and Industry Insights 4401 Matt asks Alok to define features for the audience and prompts a comparison of DoorDash's ML stack with Airbnb and Lyft. Alok educates on feature store architecture and common industry hurdles like offline-online data drift and unifying batch versus real-time predictions.
Data Science Team Growth and Interview Process 3311 Matt asks about team scale, organizational structure, and specifically probes how DoorDash evaluates business skills in interviews. Alok details how they prioritize business impact over building complex algorithms and borrow analytics and PM style case interviews.
ML Governance and Machine Learning Council 5301 Matt articulates a well-informed query regarding how COVID-19 altered the baseline assumptions of predictive models. Alok details their governance approach via the Machine Learning Council and explains tactical fixes like capping predictions and shortening historical lookback windows.
Audience Q&A on ML Infrastructure and Modeling 5413 Matt fields audience questions and pushes Alok to clarify whether DoorDash operates parallel stacks between Snowflake and Databricks/S3. Alok clarifies the distinction between Snowflake as the data warehouse and Databricks running on S3 while pulling data from Snowflake.
Career Guidance for Data Scientists and New Graduates 2200 Matt passes along audience questions regarding transitioning from academia to industry and advice for new graduates. Alok offers standard career guidance on adopting a business-first mindset and building a T-shaped technical skill set.
Fireside Chat Conclusion and Transition 0000 Brief wrap-up segment where Matt thanks Alok and concludes the chat.

Statements from this episode (14)

Disclosure
Grocery expansion turns DoorDash from a three- to four-sided marketplace
“The food delivery is a three-sided. You're right. There's the merchant, the consumer, and the dasher. As we start to branch out into grocery delivery, convenience delivery, it becomes a four-sided where we introduce a picker as well.”
Alok Gupta Feb 1, 2021 ▶ 1:11
Disclosure
DoorDash aims to offer white-label infrastructure services modeled after AWS
“I think we would like to think of ourselves as an infrastructure company. That's right. Where we build products and services to help us in the things we're trying to do, but ultimately we get them to a stage of maturity and robustness through experimenting on …”
Alok Gupta Feb 1, 2021 ▶ 3:02
Insight
Problem framing is the hardest part of the machine learning lifecycle
“I think right at the start of a project is probably, and probably the hardest and most impactful part of the life cycle of the model. I think I prefer to call it sort of data-driven software because it could ultimately be it could be something other than a mod…”
Alok Gupta Feb 1, 2021 ▶ 4:08
Disclosure
DoorDash standardized its core machine learning platform on LightGBM and PyTorch
“We landed on using a framework that enables tree-based models. And we picked light GBM for that after trying a few different packages and also deep learning. And for that, we then used PyTorch. And so we started with those two core libraries.”
Alok Gupta Feb 1, 2021 ▶ 6:29
Disclosure
DoorDash commits raw training scripts instead of pickled machine learning models
“Instead of building a model and pickling it or whatever a conversion package you want to use, we commit the actual training script to our model library.”
Alok Gupta Feb 1, 2021 ▶ 7:42
Disclosure
DoorDash, Airbnb, and Lyft all sought a centralized machine learning stack
“I would say at all three places, there was a desire for a centralized stack so that iterative improvements sort of helped all models.”
Alok Gupta Feb 1, 2021 ▶ 11:13
Insight
Consolidating training and serving data is the feature store holy grail
“And so a way to create a feature store whereby we can consolidate the training data with the prediction serving time data is a common sort of holy grail for all the companies I've been at in their pursuit.”
Alok Gupta Feb 1, 2021 ▶ 11:58
Assertion Not checkable as stated
DoorDash's data science team grew from six to thirty in 18 months
“We, so a year and a half ago when I joined we had six, five or six people on the team. We're now at almost 30 people a year and a half later”
Alok Gupta Feb 1, 2021 ▶ 13:14
Disclosure
DoorDash plans to double its data science team in 2021
“And this year we want to double, so to get to 50 plus people.”
Alok Gupta Feb 1, 2021 ▶ 13:20
Disclosure
Deploying ML models to production at DoorDash requires data science approval
“It is perfectly acceptable for anyone at DoorDash to Build a machine learning model, but if they want to put it into production, they need approval or a review from someone like a data scientist or machine learning engineer.”
Alok Gupta Feb 1, 2021 ▶ 17:45
Disclosure
DoorDash shortened ML model lookback from four weeks to one during COVID
“Whereas in the past for an assignment algorithm, we might have looked back four weeks. We changed it to one week.”
Alok Gupta Feb 1, 2021 ▶ 19:49
Assertion Not checkable as stated
DoorDash's ML feature store is mostly homegrown and uses Apache Flink
“It's mostly homegrown. We're using some technologies like Flint for one of some of our real time feature aggregation, data aggregation, but mostly it's homegrown.”
Alok Gupta Feb 1, 2021 ▶ 20:41
Assertion Not checkable as stated
DoorDash runs Databricks on AWS S3 and ingests data from Snowflake
“We run a Databricks on S-III, but we pull data into Databricks from Snowflake, so we have a connection between the two.”
Alok Gupta Feb 1, 2021 ▶ 21:24
Insight
Training dedicated models for distinct cohorts beats fitting general models
“If there's a strong enough use case for a different cohort or segment like new users, typically it'll make more sense just to train a new model rather than trying to shoehorn it into a general model.”
Alok Gupta Feb 1, 2021 ▶ 22:56
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.