Feb 17, 2021 · 24m · mad

Introducing Kedro

Kedro Product Manager · 19m spoken Matt Turck · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Yetunde Dada, Principal Product Manager at QuantumBlack (McKinsey & Company), introduces Kedro, an open-source Python framework designed to bring software engineering best practices to data science and machine learning pipelines. Through detailed architectural explanations and a live code demonstration, she shows how Kedro resolves scalability bottlenecks, enforces modularity, and integrates into modern MLOps ecosystems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 9.2% of the talking time here. How this is scored →

Matt as informed peer 0.9 Guest teaching 5.0 Guest disagreement 0.1 Matt pushing back 0.3
05100:0010:0020:000:08–3:26 · Matt as informed peer 0/10 Background on QuantumBlack and McKinsey Monologue presentation segment setting up QuantumBlack's background and explaining the gap between proof-of-concept scripts and scalable production machine learning code. The host does not speak, resulting in 0s for host metrics.3:26–7:00 · Matt as informed peer 0/10 Introduction to the Kedro Open-Source Framework Guest continues monologue introducing Kedro's architecture, software engineering concepts, and its distinction from orchestrators like Airflow or Dagster. The host remains silent throughout the segment.7:00–10:59 · Matt as informed peer 0/10 Kedro Demo: Project Template Setup Guest presents a live terminal demonstration of Kedro project templates derived from cookiecutter data science. The monologue format leaves host metrics at 0.10:59–13:06 · Matt as informed peer 0/10 Kedro Demo: Configuration and Data Catalog Guest details configuration settings and the data catalog features in Kedro. The segment is entirely guest-led with no host participation.13:06–16:13 · Matt as informed peer 0/10 Kedro Demo: Pipeline Visualization with Kedro-Viz Guest demonstrates pipeline execution and Kedro-Viz UI, with brief interjections from the host acknowledging the transition back to questions. Host engagement remains purely receptive.16:13–19:08 · Matt as informed peer 2/10 Presentation Wrap-up & AWS SageMaker Comparison Host opens Q&A by asking about SageMaker integration, product roadmap, and team composition. The dynamic is collaborative and inquisitive without tension.19:08–24:21 · Matt as informed peer 4/10 Q&A: QuantumBlack Business Model and Audience Questions Host moderates audience questions regarding data catalogs, ETL tool comparisons, and internal QuantumBlack projects. Host demonstrates subject knowledge by contextualizing Great Expectations for the audience.0:08–3:26 · Guest teaching 5/10 Background on QuantumBlack and McKinsey Monologue presentation segment setting up QuantumBlack's background and explaining the gap between proof-of-concept scripts and scalable production machine learning code. The host does not speak, resulting in 0s for host metrics.3:26–7:00 · Guest teaching 6/10 Introduction to the Kedro Open-Source Framework Guest continues monologue introducing Kedro's architecture, software engineering concepts, and its distinction from orchestrators like Airflow or Dagster. The host remains silent throughout the segment.7:00–10:59 · Guest teaching 5/10 Kedro Demo: Project Template Setup Guest presents a live terminal demonstration of Kedro project templates derived from cookiecutter data science. The monologue format leaves host metrics at 0.10:59–13:06 · Guest teaching 5/10 Kedro Demo: Configuration and Data Catalog Guest details configuration settings and the data catalog features in Kedro. The segment is entirely guest-led with no host participation.13:06–16:13 · Guest teaching 5/10 Kedro Demo: Pipeline Visualization with Kedro-Viz Guest demonstrates pipeline execution and Kedro-Viz UI, with brief interjections from the host acknowledging the transition back to questions. Host engagement remains purely receptive.16:13–19:08 · Guest teaching 5/10 Presentation Wrap-up & AWS SageMaker Comparison Host opens Q&A by asking about SageMaker integration, product roadmap, and team composition. The dynamic is collaborative and inquisitive without tension.19:08–24:21 · Guest teaching 4/10 Q&A: QuantumBlack Business Model and Audience Questions Host moderates audience questions regarding data catalogs, ETL tool comparisons, and internal QuantumBlack projects. Host demonstrates subject knowledge by contextualizing Great Expectations for the audience.0:08–3:26 · Guest disagreement 0/10 Background on QuantumBlack and McKinsey Monologue presentation segment setting up QuantumBlack's background and explaining the gap between proof-of-concept scripts and scalable production machine learning code. The host does not speak, resulting in 0s for host metrics.3:26–7:00 · Guest disagreement 0/10 Introduction to the Kedro Open-Source Framework Guest continues monologue introducing Kedro's architecture, software engineering concepts, and its distinction from orchestrators like Airflow or Dagster. The host remains silent throughout the segment.7:00–10:59 · Guest disagreement 0/10 Kedro Demo: Project Template Setup Guest presents a live terminal demonstration of Kedro project templates derived from cookiecutter data science. The monologue format leaves host metrics at 0.10:59–13:06 · Guest disagreement 0/10 Kedro Demo: Configuration and Data Catalog Guest details configuration settings and the data catalog features in Kedro. The segment is entirely guest-led with no host participation.13:06–16:13 · Guest disagreement 0/10 Kedro Demo: Pipeline Visualization with Kedro-Viz Guest demonstrates pipeline execution and Kedro-Viz UI, with brief interjections from the host acknowledging the transition back to questions. Host engagement remains purely receptive.16:13–19:08 · Guest disagreement 0/10 Presentation Wrap-up & AWS SageMaker Comparison Host opens Q&A by asking about SageMaker integration, product roadmap, and team composition. The dynamic is collaborative and inquisitive without tension.19:08–24:21 · Guest disagreement 1/10 Q&A: QuantumBlack Business Model and Audience Questions Host moderates audience questions regarding data catalogs, ETL tool comparisons, and internal QuantumBlack projects. Host demonstrates subject knowledge by contextualizing Great Expectations for the audience.0:08–3:26 · Matt pushing back 0/10 Background on QuantumBlack and McKinsey Monologue presentation segment setting up QuantumBlack's background and explaining the gap between proof-of-concept scripts and scalable production machine learning code. The host does not speak, resulting in 0s for host metrics.3:26–7:00 · Matt pushing back 0/10 Introduction to the Kedro Open-Source Framework Guest continues monologue introducing Kedro's architecture, software engineering concepts, and its distinction from orchestrators like Airflow or Dagster. The host remains silent throughout the segment.7:00–10:59 · Matt pushing back 0/10 Kedro Demo: Project Template Setup Guest presents a live terminal demonstration of Kedro project templates derived from cookiecutter data science. The monologue format leaves host metrics at 0.10:59–13:06 · Matt pushing back 0/10 Kedro Demo: Configuration and Data Catalog Guest details configuration settings and the data catalog features in Kedro. The segment is entirely guest-led with no host participation.13:06–16:13 · Matt pushing back 0/10 Kedro Demo: Pipeline Visualization with Kedro-Viz Guest demonstrates pipeline execution and Kedro-Viz UI, with brief interjections from the host acknowledging the transition back to questions. Host engagement remains purely receptive.16:13–19:08 · Matt pushing back 1/10 Presentation Wrap-up & AWS SageMaker Comparison Host opens Q&A by asking about SageMaker integration, product roadmap, and team composition. The dynamic is collaborative and inquisitive without tension.19:08–24:21 · Matt pushing back 1/10 Q&A: QuantumBlack Business Model and Audience Questions Host moderates audience questions regarding data catalogs, ETL tool comparisons, and internal QuantumBlack projects. Host demonstrates subject knowledge by contextualizing Great Expectations for the audience.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 8.4% · guest 91.6%15:00 · Matt 8.4% · guest 91.6%18:00 · Matt 29.5% · guest 70.5%18:00 · Matt 29.5% · guest 70.5%21:00 · Matt 27.3% · guest 72.7%21:00 · Matt 27.3% · guest 72.7%24:00 · Matt 94.2% · guest 5.8%24:00 · Matt 94.2% · guest 5.8%
Sharpest disagreement ▶ 21:33 Reframing Kedro vs traditional ETL tools

The guest gently rejects the audience framing that Kedro is merely an ETL tool like Talend, clarifying that it serves as pipeline scaffolding for data science.

Hardest push from Matt ▶ 19:01 Probing QuantumBlack and McKinsey relationship

Matt Turck asks the guest to explain how QuantumBlack's technical focus integrates with McKinsey's traditional strategic consulting model.

Biggest teaching moment ▶ 6:03 Distinguishing Kedro from workflow orchestrators

The guest clearly delineates Kedro's role as project scaffolding versus execution orchestrators like Airflow, Prefect, and Dagster.

Matt holds his own ▶ 23:17 Host defining Great Expectations

Matt Turck steps in to provide helpful background context to listeners regarding Great Expectations as an open-source data quality project.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Background on QuantumBlack and McKinsey 0500 Monologue presentation segment setting up QuantumBlack's background and explaining the gap between proof-of-concept scripts and scalable production machine learning code. The host does not speak, resulting in 0s for host metrics.
Introduction to the Kedro Open-Source Framework 0600 Guest continues monologue introducing Kedro's architecture, software engineering concepts, and its distinction from orchestrators like Airflow or Dagster. The host remains silent throughout the segment.
Kedro Demo: Project Template Setup 0500 Guest presents a live terminal demonstration of Kedro project templates derived from cookiecutter data science. The monologue format leaves host metrics at 0.
Kedro Demo: Configuration and Data Catalog 0500 Guest details configuration settings and the data catalog features in Kedro. The segment is entirely guest-led with no host participation.
Kedro Demo: Pipeline Visualization with Kedro-Viz 0500 Guest demonstrates pipeline execution and Kedro-Viz UI, with brief interjections from the host acknowledging the transition back to questions. Host engagement remains purely receptive.
Presentation Wrap-up & AWS SageMaker Comparison 2501 Host opens Q&A by asking about SageMaker integration, product roadmap, and team composition. The dynamic is collaborative and inquisitive without tension.
Q&A: QuantumBlack Business Model and Audience Questions 4411 Host moderates audience questions regarding data catalogs, ETL tool comparisons, and internal QuantumBlack projects. Host demonstrates subject knowledge by contextualizing Great Expectations for the audience.

Statements from this episode (5)

Insight
Dada argues building data pipelines is harder than machine learning itself
“And this is where we believe that machine learning is not really the hard part, but building and maintaining the data pipeline is”
Kedro Product Manager Feb 17, 2021 ▶ 3:18
Disclosure
Kedro leaves pipeline scheduling and failure monitoring to Airflow and Dagster
“How, what time will this pipeline run? How will I know if it failed? We'll leave those tools to Dagster, Airflow, Prefect, Luigi, and many others to actually handle for you, because they do that really well.”
Kedro Product Manager Feb 17, 2021 ▶ 6:25
Insight
Dada asserts notebooks are challenging for reproducible, production-ready data pipelines
“In terms of actually building your full pipeline and notebook that's where things get a little bit complicated because notebooks are a little bit challenging for reproducible and what we call production-ready code.”
Kedro Product Manager Feb 17, 2021 ▶ 10:27
Insight
Dada says data science experiment parameters must be separated from codebases
“So your experiment parameters for data science workflow should also be outside of your code base because we really believe it helps you make code that is generalizable and reusable by removing things that loading and load the code for loading and saving your d…”
Kedro Product Manager Feb 17, 2021 ▶ 12:25
Assertion Not checkable as stated
Dada argues AWS SageMaker is a deployment target, not a code-structuring tool
“We use, we think of Kedra or SageMaker as a deployment target for Kedra. So still that whole process of like creating like, well document, well structured, Modular data science code, but still having the freedom to deploy it on SageMaker is how we see those ro…”
Kedro Product Manager Feb 17, 2021 ▶ 16:27
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.