Feb 1, 2021 · 27m · mad

Data Observability and Pipelines: OpenLineage and Marquez

Julien Le Dem · 24m spoken Matt Turck · 60s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Julien Le Dem, Founder and CTO of Datakin, presents on data pipeline observability, introducing the open-source OpenLineage standard and Marquez reference implementation. He explains how standardized runtime metadata solves integration complexity, stabilizes data platform operations, and enhances data governance across modern organizations.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.4% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 3.6 Guest disagreement 0.2 Matt pushing back 0.4
05100:0010:0020:001:31–5:56 · Matt as informed peer 0/10 The Need for Metadata and Data Hierarchy This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero.5:56–12:42 · Matt as informed peer 0/10 Introduction to OpenLineage Standard and Purpose Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue.12:42–17:56 · Matt as informed peer 0/10 OpenLineage Architecture, Core Model, and Facets Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero.17:56–20:45 · Matt as informed peer 0/10 Marquez Reference Implementation and Datakin Commercial Layer Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation.20:45–27:43 · Matt as informed peer 4/10 Audience Q&A on Hive, Security, and Adoption Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide.1:31–5:56 · Guest teaching 4/10 The Need for Metadata and Data Hierarchy This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero.5:56–12:42 · Guest teaching 4/10 Introduction to OpenLineage Standard and Purpose Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue.12:42–17:56 · Guest teaching 3/10 OpenLineage Architecture, Core Model, and Facets Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero.17:56–20:45 · Guest teaching 3/10 Marquez Reference Implementation and Datakin Commercial Layer Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation.20:45–27:43 · Guest teaching 4/10 Audience Q&A on Hive, Security, and Adoption Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide.1:31–5:56 · Guest disagreement 0/10 The Need for Metadata and Data Hierarchy This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero.5:56–12:42 · Guest disagreement 0/10 Introduction to OpenLineage Standard and Purpose Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue.12:42–17:56 · Guest disagreement 0/10 OpenLineage Architecture, Core Model, and Facets Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero.17:56–20:45 · Guest disagreement 0/10 Marquez Reference Implementation and Datakin Commercial Layer Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation.20:45–27:43 · Guest disagreement 1/10 Audience Q&A on Hive, Security, and Adoption Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide.1:31–5:56 · Matt pushing back 0/10 The Need for Metadata and Data Hierarchy This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero.5:56–12:42 · Matt pushing back 0/10 Introduction to OpenLineage Standard and Purpose Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue.12:42–17:56 · Matt pushing back 0/10 OpenLineage Architecture, Core Model, and Facets Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero.17:56–20:45 · Matt pushing back 0/10 Marquez Reference Implementation and Datakin Commercial Layer Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation.20:45–27:43 · Matt pushing back 2/10 Audience Q&A on Hive, Security, and Adoption Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 10.4% · guest 89.6%18:00 · Matt 10.4% · guest 89.6%21:00 · Matt 11.2% · guest 88.8%21:00 · Matt 11.2% · guest 88.8%24:00 · Matt 15.2% · guest 84.8%24:00 · Matt 15.2% · guest 84.8%27:00 · Matt 19.7% · guest 80.3%27:00 · Matt 19.7% · guest 80.3%
Sharpest disagreement ▶ 24:43 Julien counters expectation of early proprietary adoption

Julien politely rejects the premise that OpenLineage needs immediate backing from commercial giants like Snowflake or BigQuery, citing Parquet's trajectory where open-source adoption preceded vendor support.

Hardest push from Matt ▶ 24:13 Matt raises audience question on lack of major database support

Matt pushes the guest on market momentum by posing Tony Bear's question regarding whether any major household database or ETL vendors have actually committed to the standard.

Biggest teaching moment ▶ 21:14 Julien explains runtime job lineage vs static metastores

Julien educates the audience on why traditional tools like Hive Metastore capture static schema states rather than the dynamic runtime execution lineage targeted by OpenLineage.

Matt holds his own ▶ 20:41 Matt moderates Q&A with precise ecosystem knowledge

Matt demonstrates his domain expertise across the data ecosystem by selecting targeted questions on Hive Metastore, prompting Julien to playfully reference Matt's famous Data Landscape chart.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Need for Metadata and Data Hierarchy 0400 This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero.
Introduction to OpenLineage Standard and Purpose 0400 Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue.
OpenLineage Architecture, Core Model, and Facets 0300 Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero.
Marquez Reference Implementation and Datakin Commercial Layer 0300 Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation.
Audience Q&A on Hive, Security, and Adoption 4412 Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide.

Statements from this episode (5)

Disclosure
Marquez started at WeWork as an OpenLineage reference implementation
“Marquez which is an open source project we started while I was at WeWork and which is also a reference implementation of the open lineage standard.”
Julien Le Dem Feb 1, 2021 ▶ 1:19
Insight
The best time to collect data pipeline metadata is during runtime
“The best time to collect metadata is when the job is running and we can inspect it to figure out what the inputs were, what the outputs were, how long it took you know, what was the version of the code and things like that.”
Julien Le Dem Feb 1, 2021 ▶ 9:22
Insight
Collecting massive amounts of metadata creates a signal-to-noise challenge
“I think the curse of collecting all that metadata that becomes really hard to figure out, to find the right information in the mass of data, and therefore that's what we strive to solve, right?”
Julien Le Dem Feb 1, 2021 ▶ 19:41
Assertion Not checkable as stated
90% of contacted open source projects want to contribute to OpenLineage
“90% of the project we reach out to see a lot of value in this and want to contribute”
Julien Le Dem Feb 1, 2021 ▶ 26:19
Prediction Not checkable as stated
OpenLineage will take a couple of years to achieve widespread adoption
“It will take a couple of years to get that but all the open source project are kind of the early adopters of the idea.”
Julien Le Dem Feb 1, 2021 ▶ 27:17
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.