Julien Le Dem

Principal Engineer, Datadog · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

engineerfounderexecutive@J_ ↗LinkedIn ↗julien.ledem.net ↗

Julien Le Dem is a data infrastructure leader best known as a co-creator of open-source projects including Apache Parquet, Apache Arrow, and OpenLineage. He co-founded data observability startup Datakin and previously led data processing tools at Twitter.

5statements → 2claims → 0claims resolved → 3.6/5average certainty → 1.8/5average debate potential → 4said about them ↓

2 not checkable as stated how the 2 claims stand · each chip opens the sources

1 prediction · 1 assertion · 2 insights · 1 disclosure · every statement was checked. The prediction and assertion are the 2 claims: statements the public record can support or contradict. 0 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Julien argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

How they sound: speaking style how? →

218 words/min while actually speaking · 29.6 um and uh per 1k words

No argument clarity score for Julien Le Dem: only 1 usable question→answer exchange on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 4,331 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Julien Le Dem said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
The best time to collect data pipeline metadata is during runtime
“The best time to collect metadata is when the job is running and we can inspect it to figure out what the inputs were, what the outputs were, how long it took you know, what was the version of the code and things like that.”
Julien Le Dem Feb 1, 2021 ▶ 9:22 Data Observability and Pipelines: OpenLineage and Marquez
Insight
Collecting massive amounts of metadata creates a signal-to-noise challenge
“I think the curse of collecting all that metadata that becomes really hard to figure out, to find the right information in the mass of data, and therefore that's what we strive to solve, right?”
Julien Le Dem Feb 1, 2021 ▶ 19:41 Data Observability and Pipelines: OpenLineage and Marquez
Assertion Not checkable as stated
90% of contacted open source projects want to contribute to OpenLineage
“90% of the project we reach out to see a lot of value in this and want to contribute”
Julien Le Dem Feb 1, 2021 ▶ 26:19 Data Observability and Pipelines: OpenLineage and Marquez
Prediction Not checkable as stated
OpenLineage will take a couple of years to achieve widespread adoption
“It will take a couple of years to get that but all the open source project are kind of the early adopters of the idea.”
Julien Le Dem Feb 1, 2021 ▶ 27:17 Data Observability and Pipelines: OpenLineage and Marquez
Disclosure
Marquez started at WeWork as an OpenLineage reference implementation
“Marquez which is an open source project we started while I was at WeWork and which is also a reference implementation of the open lineage standard.”
Julien Le Dem Feb 1, 2021 ▶ 1:19 Data Observability and Pipelines: OpenLineage and Marquez

The other half of the tape: Julien Le Dem's own voice is left out of every number here. Other people bring the name up 4 times in 3 episodes on the MAD Podcast. every mention, with the transcript →

Who brings them up most Wes McKinney 2Pete DeJoy 1Matt Turck 1

Every mention by year

tap a year for its mentions
00213220212022episodesmentions
01220212022episodes it came up in
000.811.5220212022episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
Data Observability and Pipelines: OpenLineage and Marquez Feb 1, 2021 24m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.