Feb 1, 2021 · 27m · mad
Data Observability and Pipelines: OpenLineage and Marquez
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Julien Le Dem, Founder and CTO of Datakin, presents on data pipeline observability, introducing the open-source OpenLineage standard and Marquez reference implementation. He explains how standardized runtime metadata solves integration complexity, stabilizes data platform operations, and enhances data governance across modern organizations.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Julien politely rejects the premise that OpenLineage needs immediate backing from commercial giants like Snowflake or BigQuery, citing Parquet's trajectory where open-source adoption preceded vendor support.
Hardest push from Matt ▶ 24:13 Matt raises audience question on lack of major database supportMatt pushes the guest on market momentum by posing Tony Bear's question regarding whether any major household database or ETL vendors have actually committed to the standard.
Biggest teaching moment ▶ 21:14 Julien explains runtime job lineage vs static metastoresJulien educates the audience on why traditional tools like Hive Metastore capture static schema states rather than the dynamic runtime execution lineage targeted by OpenLineage.
Matt holds his own ▶ 20:41 Matt moderates Q&A with precise ecosystem knowledgeMatt demonstrates his domain expertise across the data ecosystem by selecting targeted questions on Hive Metastore, prompting Julien to playfully reference Matt's famous Data Landscape chart.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| The Need for Metadata and Data Hierarchy | 0 | 4 | 0 | 0 | This is a solo monologue presentation by Julien explaining data pipeline observability and the Maslow hierarchy of data needs. The host is not present during this section, requiring all host scores to be zero. | |
| Introduction to OpenLineage Standard and Purpose | 0 | 4 | 0 | 0 | Julien presents the core vision for OpenLineage using the digital camera EXIF metadata analogy to contrast standardized push integration against brittle custom scraping. The host remains silent throughout the monologue. | |
| OpenLineage Architecture, Core Model, and Facets | 0 | 3 | 0 | 0 | Julien details the OpenLineage core model (run, job, dataset) and the extensible facet architecture. As a monologue presentation segment without host interaction, host scores are strictly zero. | |
| Marquez Reference Implementation and Datakin Commercial Layer | 0 | 3 | 0 | 0 | Julien concludes his slide deck covering Marquez as the reference backend and Datakin as the commercial SaaS layer. The segment is a pure monologue without host participation. | |
| Audience Q&A on Hive, Security, and Adoption | 4 | 4 | 1 | 2 | Matt moderates audience questions covering Hive Metastore, security models, and enterprise database support. The exchange is highly collaborative and courteous, with Julien making a warm reference to Matt's data landscape slide. |