Jul 13, 2017 · 23m · mad
AI, Big Data, and Data Governance // Stan Christiaens, Collibra (FirstMark's Data Driven)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC presentation and fireside chat, Collibra Founder & CTO Stan Christiaens explores the common pitfalls of artificial intelligence and demonstrates how robust data governance enables organizations to find, trust, and extract real value from their data assets.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Christiaens expresses sharp frustration with data scientists complaining about lack of access, declaring that the excuse pisses him off and telling them to hack into internal databases if necessary.
Hardest push from Matt ▶ 18:59 Challenging data lake centralisationTurck presses Christiaens on whether centralized data lakes are superior to distributed repositories, challenging the traditional architecture assumption.
Biggest teaching moment ▶ 17:49 Catalog vs governance explanationChristiaens clarifies a common industry misconception by explaining that standalone data catalogs become useless phone books without overarching governance frameworks.
Matt holds his own ▶ 18:59 Framing agile governance choicesTurck demonstrates strong domain familiarity by articulating the technical trade-offs between centralized data lakes and distributed repository models requiring agile software.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Frustrations in Unlocking Value from Data | 0 | 2 | 1 | 0 | This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero. | |
| What Was Missed: The Explosion of AI & Machine Learning | 0 | 2 | 1 | 0 | Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment. | |
| Pitfall 2: Falling into the Snake Oil Salesman's Trap | 0 | 3 | 2 | 0 | Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment. | |
| Pitfall 3: Betting Too Much on Algorithms | 0 | 3 | 3 | 0 | Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue. | |
| Fireside Chat: Data Governance vs. Data Cataloging | 4 | 5 | 2 | 2 | Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books. | |
| Audience Q&A on Data Ownership and Privacy | 1 | 4 | 2 | 0 | An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management. |