Feb 17, 2021 · 25m · mad

Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)

Savin Goyal · 19m spoken Matt Turck · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC fireside chat, Matt Turck interviews Savin Goyal, Tech Lead for Netflix's ML Infrastructure Team, discussing Netflix's data ecosystem, the creation and architecture of Metaflow, and how open-sourcing the framework empowers data science teams.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 16.7% of the talking time here. How this is scored →

Matt as informed peer 2.4 Guest teaching 3.7 Guest disagreement 0.0 Matt pushing back 0.9
05100:0010:0020:000:08–2:40 · Matt as informed peer 1/10 Machine Learning Organization and Applications at Netflix Matt opens with broad introductory questions about Netflix's data science organization and team size. Savin politely educates him on how machine learning extends beyond recommendations into business data science and explains that his small team of six engineers supports advanced ML practitioners.2:40–5:12 · Matt as informed peer 2/10 Overview of Netflix's Data Stack and Computing Infrastructure Matt asks for a general overview of Netflix's data stack. Savin gives a detailed technical inventory including AWS S3, Spark, Snowflake, Titus, and Mason, with Matt briefly interrupting only to clarify an unfamiliar project name.5:12–8:56 · Matt as informed peer 3/10 Addressing the Productivity Gap with Metaflow Matt relates Metaflow to a prior talk on Kedro before asking for the core motivation behind the framework. Savin articulates the productivity gap data scientists face when forced to handle heavy software engineering and infrastructure concerns.8:56–14:53 · Matt as informed peer 3/10 Metaflow Core Features and Architecture Overview Matt asks probing practical questions regarding local versus cloud execution and how resource demands are specified. Savin explains the mechanics of using decorators to seamlessly scale Python and R code to cloud clusters.14:53–18:00 · Matt as informed peer 3/10 Navigating the MLOps Landscape and Automated Infrastructure Matt directly challenges Savin on the true level of automation provided by Metaflow, asking if it is genuinely a single-button experience or requires engineering labor. Savin clarifies that Metaflow fully automates infrastructure dependencies and cluster offloading.18:00–21:32 · Matt as informed peer 2/10 Reasons Behind Open-Sourcing Metaflow and Community Growth Matt inquires about Netflix's strategic motivation for open-sourcing internal projects. Savin highlights Netflix's culture of freedom and responsibility, along with the benefit of learning from external enterprise use cases.21:32–24:39 · Matt as informed peer 3/10 Ideal Use Cases and Constraints for Metaflow Matt brings in an audience question asking when NOT to use Metaflow and suggests non-AWS usage as a primary constraint. Savin slightly reframes this, explaining that local laptop execution still works off AWS, but non-Python/R codebases or simple scripts are key reasons to avoid it.0:08–2:40 · Guest teaching 3/10 Machine Learning Organization and Applications at Netflix Matt opens with broad introductory questions about Netflix's data science organization and team size. Savin politely educates him on how machine learning extends beyond recommendations into business data science and explains that his small team of six engineers supports advanced ML practitioners.2:40–5:12 · Guest teaching 4/10 Overview of Netflix's Data Stack and Computing Infrastructure Matt asks for a general overview of Netflix's data stack. Savin gives a detailed technical inventory including AWS S3, Spark, Snowflake, Titus, and Mason, with Matt briefly interrupting only to clarify an unfamiliar project name.5:12–8:56 · Guest teaching 5/10 Addressing the Productivity Gap with Metaflow Matt relates Metaflow to a prior talk on Kedro before asking for the core motivation behind the framework. Savin articulates the productivity gap data scientists face when forced to handle heavy software engineering and infrastructure concerns.8:56–14:53 · Guest teaching 4/10 Metaflow Core Features and Architecture Overview Matt asks probing practical questions regarding local versus cloud execution and how resource demands are specified. Savin explains the mechanics of using decorators to seamlessly scale Python and R code to cloud clusters.14:53–18:00 · Guest teaching 3/10 Navigating the MLOps Landscape and Automated Infrastructure Matt directly challenges Savin on the true level of automation provided by Metaflow, asking if it is genuinely a single-button experience or requires engineering labor. Savin clarifies that Metaflow fully automates infrastructure dependencies and cluster offloading.18:00–21:32 · Guest teaching 3/10 Reasons Behind Open-Sourcing Metaflow and Community Growth Matt inquires about Netflix's strategic motivation for open-sourcing internal projects. Savin highlights Netflix's culture of freedom and responsibility, along with the benefit of learning from external enterprise use cases.21:32–24:39 · Guest teaching 4/10 Ideal Use Cases and Constraints for Metaflow Matt brings in an audience question asking when NOT to use Metaflow and suggests non-AWS usage as a primary constraint. Savin slightly reframes this, explaining that local laptop execution still works off AWS, but non-Python/R codebases or simple scripts are key reasons to avoid it.0:08–2:40 · Guest disagreement 0/10 Machine Learning Organization and Applications at Netflix Matt opens with broad introductory questions about Netflix's data science organization and team size. Savin politely educates him on how machine learning extends beyond recommendations into business data science and explains that his small team of six engineers supports advanced ML practitioners.2:40–5:12 · Guest disagreement 0/10 Overview of Netflix's Data Stack and Computing Infrastructure Matt asks for a general overview of Netflix's data stack. Savin gives a detailed technical inventory including AWS S3, Spark, Snowflake, Titus, and Mason, with Matt briefly interrupting only to clarify an unfamiliar project name.5:12–8:56 · Guest disagreement 0/10 Addressing the Productivity Gap with Metaflow Matt relates Metaflow to a prior talk on Kedro before asking for the core motivation behind the framework. Savin articulates the productivity gap data scientists face when forced to handle heavy software engineering and infrastructure concerns.8:56–14:53 · Guest disagreement 0/10 Metaflow Core Features and Architecture Overview Matt asks probing practical questions regarding local versus cloud execution and how resource demands are specified. Savin explains the mechanics of using decorators to seamlessly scale Python and R code to cloud clusters.14:53–18:00 · Guest disagreement 0/10 Navigating the MLOps Landscape and Automated Infrastructure Matt directly challenges Savin on the true level of automation provided by Metaflow, asking if it is genuinely a single-button experience or requires engineering labor. Savin clarifies that Metaflow fully automates infrastructure dependencies and cluster offloading.18:00–21:32 · Guest disagreement 0/10 Reasons Behind Open-Sourcing Metaflow and Community Growth Matt inquires about Netflix's strategic motivation for open-sourcing internal projects. Savin highlights Netflix's culture of freedom and responsibility, along with the benefit of learning from external enterprise use cases.21:32–24:39 · Guest disagreement 0/10 Ideal Use Cases and Constraints for Metaflow Matt brings in an audience question asking when NOT to use Metaflow and suggests non-AWS usage as a primary constraint. Savin slightly reframes this, explaining that local laptop execution still works off AWS, but non-Python/R codebases or simple scripts are key reasons to avoid it.0:08–2:40 · Matt pushing back 0/10 Machine Learning Organization and Applications at Netflix Matt opens with broad introductory questions about Netflix's data science organization and team size. Savin politely educates him on how machine learning extends beyond recommendations into business data science and explains that his small team of six engineers supports advanced ML practitioners.2:40–5:12 · Matt pushing back 1/10 Overview of Netflix's Data Stack and Computing Infrastructure Matt asks for a general overview of Netflix's data stack. Savin gives a detailed technical inventory including AWS S3, Spark, Snowflake, Titus, and Mason, with Matt briefly interrupting only to clarify an unfamiliar project name.5:12–8:56 · Matt pushing back 0/10 Addressing the Productivity Gap with Metaflow Matt relates Metaflow to a prior talk on Kedro before asking for the core motivation behind the framework. Savin articulates the productivity gap data scientists face when forced to handle heavy software engineering and infrastructure concerns.8:56–14:53 · Matt pushing back 2/10 Metaflow Core Features and Architecture Overview Matt asks probing practical questions regarding local versus cloud execution and how resource demands are specified. Savin explains the mechanics of using decorators to seamlessly scale Python and R code to cloud clusters.14:53–18:00 · Matt pushing back 2/10 Navigating the MLOps Landscape and Automated Infrastructure Matt directly challenges Savin on the true level of automation provided by Metaflow, asking if it is genuinely a single-button experience or requires engineering labor. Savin clarifies that Metaflow fully automates infrastructure dependencies and cluster offloading.18:00–21:32 · Matt pushing back 0/10 Reasons Behind Open-Sourcing Metaflow and Community Growth Matt inquires about Netflix's strategic motivation for open-sourcing internal projects. Savin highlights Netflix's culture of freedom and responsibility, along with the benefit of learning from external enterprise use cases.21:32–24:39 · Matt pushing back 1/10 Ideal Use Cases and Constraints for Metaflow Matt brings in an audience question asking when NOT to use Metaflow and suggests non-AWS usage as a primary constraint. Savin slightly reframes this, explaining that local laptop execution still works off AWS, but non-Python/R codebases or simple scripts are key reasons to avoid it.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 28% · guest 72%0:00 · Matt 28% · guest 72%3:00 · Matt 20.6% · guest 79.4%3:00 · Matt 20.6% · guest 79.4%6:00 · Matt 1.7% · guest 98.3%6:00 · Matt 1.7% · guest 98.3%9:00 · Matt 5.3% · guest 94.7%9:00 · Matt 5.3% · guest 94.7%12:00 · Matt 5.5% · guest 94.5%12:00 · Matt 5.5% · guest 94.5%15:00 · Matt 18.4% · guest 81.6%15:00 · Matt 18.4% · guest 81.6%18:00 · Matt 8.9% · guest 91.1%18:00 · Matt 8.9% · guest 91.1%21:00 · Matt 19.1% · guest 80.9%21:00 · Matt 19.1% · guest 80.9%24:00 · Matt 61.6% · guest 38.4%24:00 · Matt 61.6% · guest 38.4%
Sharpest disagreement ▶ 23:48 Polite reframe of AWS as absolute requirement

When Matt suggests that not using AWS is the main reason to avoid Metaflow, Savin mildly reframes the premise by pointing out that non-AWS users can still gain value locally on laptops.

Hardest push from Matt ▶ 16:45 Pressing on actual automation vs engineering overhead

Matt refuses to accept a general explanation of features, explicitly challenging Savin to specify if Metaflow is truly one-click automated or requires days of engineer tweaking.

Biggest teaching moment ▶ 12:04 Correcting assumption on automatic GPU detection

When Matt asks how Metaflow automatically senses that a workflow needs GPUs or extra memory, Savin corrects the assumption by explaining that users must explicitly tag code with decorators.

Matt holds his own ▶ 5:08 Connecting Metaflow to prior industry tools

Matt demonstrates domain context by comparing Metaflow's design approach directly with Kedro and existing scheduling paradigms.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Machine Learning Organization and Applications at Netflix 1300 Matt opens with broad introductory questions about Netflix's data science organization and team size. Savin politely educates him on how machine learning extends beyond recommendations into business data science and explains that his small team of six engineers supports advanced ML practitioners.
Overview of Netflix's Data Stack and Computing Infrastructure 2401 Matt asks for a general overview of Netflix's data stack. Savin gives a detailed technical inventory including AWS S3, Spark, Snowflake, Titus, and Mason, with Matt briefly interrupting only to clarify an unfamiliar project name.
Addressing the Productivity Gap with Metaflow 3500 Matt relates Metaflow to a prior talk on Kedro before asking for the core motivation behind the framework. Savin articulates the productivity gap data scientists face when forced to handle heavy software engineering and infrastructure concerns.
Metaflow Core Features and Architecture Overview 3402 Matt asks probing practical questions regarding local versus cloud execution and how resource demands are specified. Savin explains the mechanics of using decorators to seamlessly scale Python and R code to cloud clusters.
Navigating the MLOps Landscape and Automated Infrastructure 3302 Matt directly challenges Savin on the true level of automation provided by Metaflow, asking if it is genuinely a single-button experience or requires engineering labor. Savin clarifies that Metaflow fully automates infrastructure dependencies and cluster offloading.
Reasons Behind Open-Sourcing Metaflow and Community Growth 2300 Matt inquires about Netflix's strategic motivation for open-sourcing internal projects. Savin highlights Netflix's culture of freedom and responsibility, along with the benefit of learning from external enterprise use cases.
Ideal Use Cases and Constraints for Metaflow 3401 Matt brings in an audience question asking when NOT to use Metaflow and suggests non-AWS usage as a primary constraint. Savin slightly reframes this, explaining that local laptop execution still works off AWS, but non-Python/R codebases or simple scripts are key reasons to avoid it.

Statements from this episode (7)

Disclosure
Netflix's machine learning infrastructure team consists of six engineers
“My team has six engineers right now.”
Savin Goyal Feb 17, 2021 ▶ 2:28
Assertion Not checkable as stated
Netflix uses S3, Spark, Presto, and Snowflake for data querying
“We use SG as so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake as our query engines.”
Savin Goyal Feb 17, 2021 ▶ 3:23
Assertion Not checkable as stated
Netflix uses open-source Titus for container orchestration
“In terms of compute our container orchestration platform is called Titus, which is yet another open source project.”
Savin Goyal Feb 17, 2021 ▶ 3:55
Assertion Not checkable as stated
Netflix drives ETL and ML pipelines using internal scheduler Meson
“We have a workflow scheduler called Mason. Mason, which is a program scheduler that's being used to drive all of our ETL as well as machine learning pipelines.”
Savin Goyal Feb 17, 2021 ▶ 4:29
Disclosure
Netflix's Metaflow is not intended to be a workflow orchestrator
“Metaflow does not intend to be a workflow orchestrator.”
Savin Goyal Feb 17, 2021 ▶ 13:46
Disclosure
Netflix data scientists are free to choose their own tools
“Netflix has this really interesting corporate culture of freedom and responsibility. Which means that our data science teams, they are essentially free to use whatever tooling that works best for them.”
Savin Goyal Feb 17, 2021 ▶ 19:51
Assertion Not checkable as stated
Most Metaflow users adopt it to avoid building ML infra teams
“A big majority of our users are essentially companies who have made serious investments in machine learning. But for one reason or the other, they don't want to invest too much into machine learning infrastructure per se.”
Savin Goyal Feb 17, 2021 ▶ 22:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.