Jul 13, 2017 · 23m · mad

AI, Big Data, and Data Governance // Stan Christiaens, Collibra (FirstMark's Data Driven)

Stan Christiaens · 19m spoken Matt Turck · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC presentation and fireside chat, Collibra Founder & CTO Stan Christiaens explores the common pitfalls of artificial intelligence and demonstrates how robust data governance enables organizations to find, trust, and extract real value from their data assets.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.9% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 3.2 Guest disagreement 1.8 Matt pushing back 0.3
05100:0010:0020:000:15–4:50 · Matt as informed peer 0/10 Frustrations in Unlocking Value from Data This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero.4:50–9:24 · Matt as informed peer 0/10 What Was Missed: The Explosion of AI & Machine Learning Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment.9:24–11:52 · Matt as informed peer 0/10 Pitfall 2: Falling into the Snake Oil Salesman's Trap Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment.11:52–16:08 · Matt as informed peer 0/10 Pitfall 3: Betting Too Much on Algorithms Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue.16:08–20:39 · Matt as informed peer 4/10 Fireside Chat: Data Governance vs. Data Cataloging Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books.20:39–23:13 · Matt as informed peer 1/10 Audience Q&A on Data Ownership and Privacy An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management.0:15–4:50 · Guest teaching 2/10 Frustrations in Unlocking Value from Data This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero.4:50–9:24 · Guest teaching 2/10 What Was Missed: The Explosion of AI & Machine Learning Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment.9:24–11:52 · Guest teaching 3/10 Pitfall 2: Falling into the Snake Oil Salesman's Trap Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment.11:52–16:08 · Guest teaching 3/10 Pitfall 3: Betting Too Much on Algorithms Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue.16:08–20:39 · Guest teaching 5/10 Fireside Chat: Data Governance vs. Data Cataloging Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books.20:39–23:13 · Guest teaching 4/10 Audience Q&A on Data Ownership and Privacy An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management.0:15–4:50 · Guest disagreement 1/10 Frustrations in Unlocking Value from Data This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero.4:50–9:24 · Guest disagreement 1/10 What Was Missed: The Explosion of AI & Machine Learning Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment.9:24–11:52 · Guest disagreement 2/10 Pitfall 2: Falling into the Snake Oil Salesman's Trap Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment.11:52–16:08 · Guest disagreement 3/10 Pitfall 3: Betting Too Much on Algorithms Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue.16:08–20:39 · Guest disagreement 2/10 Fireside Chat: Data Governance vs. Data Cataloging Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books.20:39–23:13 · Guest disagreement 2/10 Audience Q&A on Data Ownership and Privacy An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management.0:15–4:50 · Matt pushing back 0/10 Frustrations in Unlocking Value from Data This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero.4:50–9:24 · Matt pushing back 0/10 What Was Missed: The Explosion of AI & Machine Learning Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment.9:24–11:52 · Matt pushing back 0/10 Pitfall 2: Falling into the Snake Oil Salesman's Trap Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment.11:52–16:08 · Matt pushing back 0/10 Pitfall 3: Betting Too Much on Algorithms Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue.16:08–20:39 · Matt pushing back 2/10 Fireside Chat: Data Governance vs. Data Cataloging Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books.20:39–23:13 · Matt pushing back 0/10 Audience Q&A on Data Ownership and Privacy An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 16.4% · guest 83.6%15:00 · Matt 16.4% · guest 83.6%18:00 · Matt 17.1% · guest 82.9%18:00 · Matt 17.1% · guest 82.9%21:00 · Matt 6.5% · guest 93.5%21:00 · Matt 6.5% · guest 93.5%
Sharpest disagreement ▶ 13:00 Rant on data scientist excuses

Christiaens expresses sharp frustration with data scientists complaining about lack of access, declaring that the excuse pisses him off and telling them to hack into internal databases if necessary.

Hardest push from Matt ▶ 18:59 Challenging data lake centralisation

Turck presses Christiaens on whether centralized data lakes are superior to distributed repositories, challenging the traditional architecture assumption.

Biggest teaching moment ▶ 17:49 Catalog vs governance explanation

Christiaens clarifies a common industry misconception by explaining that standalone data catalogs become useless phone books without overarching governance frameworks.

Matt holds his own ▶ 18:59 Framing agile governance choices

Turck demonstrates strong domain familiarity by articulating the technical trade-offs between centralized data lakes and distributed repository models requiring agile software.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Frustrations in Unlocking Value from Data 0210 This initial segment is a monologue presentation by guest Stan Christiaens outlining corporate frustrations with data value and predictions around the Chief Data Officer role. As a solo presentation, the host does not participate, keeping host-side metrics at zero.
What Was Missed: The Explosion of AI & Machine Learning 0210 Christiaens continues his presentation, explaining why Collibra previously missed the AI/ML surge due to past AI winter skepticism and NLP busts. The host remains silent throughout the presentation segment.
Pitfall 2: Falling into the Snake Oil Salesman's Trap 0320 Christiaens critiques AI hype, pointing out failures like IBM Watson's cancer initiative with MD Anderson and Google Flu Trends overfitting. The host is inactive during this presentation segment.
Pitfall 3: Betting Too Much on Algorithms 0330 Christiaens argues that data rather than algorithms provides true competitive differentiation, venting forcefully about data scientists making excuses about missing data. The host remains in listener mode with no dialogue.
Fireside Chat: Data Governance vs. Data Cataloging 4522 Host Matt Turck engages in a Q&A session, asking informed questions about data cataloging versus data governance and centralized data lakes versus distributed repos. Christiaens educates the audience on why catalogs without governance act like unmaintained phone books.
Audience Q&A on Data Ownership and Privacy 1420 An audience member asks about data ownership and privacy rights. Christiaens explains the contrast between European regulatory approaches and Silicon Valley practices while the host facilitates time management.

Statements from this episode (14)

Insight
Data scientists' biggest problem is finding data
“One of their first frustrations is, I can't find the data, right? That's like their biggest problem.”
Stan Christiaens Jul 13, 2017 ▶ 1:13
Insight
Data scientists quit when companies fail to operationalize their work
“So these talented people that you then hire are becoming demotivated and will actually go somewhere else just because their work is not actually adding any value to the business.”
Stan Christiaens Jul 13, 2017 ▶ 1:47
Prediction Not checkable as stated
Enterprise data needs a dedicated system of record like ERP or CRM
“Data will require a system of records, right, just like the chief financial officer as a CRM an ERP, or the VP of sales as a CRM system, data will have the same kind of need.”
Stan Christiaens Jul 13, 2017 ▶ 2:57
Prediction Not checkable as stated
Data users will revolt against Chief Data Officers who restrict access
“Data citizens who are all people who use data to do their work. In our company, product managers are data citizens, for example. Sales ops is data citizens. They will rise up against the data dictator. That's a chief data officer who tries too much to control …”
Stan Christiaens Jul 13, 2017 ▶ 3:42
Opinion
Big tech uses AI and machine learning as its next feature war
“And three, and this is a belief of me that could be wrong, is that actually the big tech firms, Amazon, Facebook, Google, Microsoft, and all the others, they're using AI and machine learning as their next feature war.”
Stan Christiaens Jul 13, 2017 ▶ 6:18
Prediction Not checkable as stated
AI will generate business value first in highly specialized applications
“So AI will have business value first in specialized applications.”
Stan Christiaens Jul 13, 2017 ▶ 8:26
Assertion Supported
IBM Watson's $60M MD Anderson oncology project was a complete failure
“Now several years later, and sixty million dollars down the drain, that initiative failed. It actually failed.”
Stan Christiaens Jul 13, 2017 ▶ 10:00
Assertion Supported
Google Flu Trends suffered a 140% prediction error in 2013
“So in 2013, they had a mismatch of a 140% prediction versus the actual situation in the world.”
Stan Christiaens Jul 13, 2017 ▶ 11:16
Insight
Competitive advantage in AI comes from proprietary data, not algorithms
“Our belief is that AI and machine learning, you will not win this war by having the better algorithm. The differentiator, the value proposition, is not in the algorithm. It's actually in the data.”
Stan Christiaens Jul 13, 2017 ▶ 11:58
Insight
Data treated as a strategic asset requires its own business process
“And data today, if you treat it as a strategic asset, should have its own business process.”
Stan Christiaens Jul 13, 2017 ▶ 15:51
Insight
Data catalogs without governance inevitably fail and fall into disuse
“If you have a Data catalog, which is really like a listing of all the data sets and attributes, dictionaries that are out there, without governance around it, it's just like a phone book. It doesn't do anything, it's not controlled, it doesn't work, it's gonna…”
Stan Christiaens Jul 13, 2017 ▶ 18:09
Assertion Not checkable as stated
Twitter and LinkedIn's open-source data catalogs go unused without governance
“Reason why I say that is because I've talked to all the high-tech companies like Twitter, LinkedIn, et cetera, and they all have their catalog projects. Many of them have actually open sourced catalog initiatives, But then you find that nobody's actually using…”
Stan Christiaens Jul 13, 2017 ▶ 18:26
Prediction Not checkable as stated
Big tech data lock-in risks market monopolization without strict regulation
“The data is already in the hands of these technology companies, so that's definitely going to continue, and I'm hoping that the regulations will impose enough sanctions so that these big, ah, you know, internet giants are actually allowing more control of an i…”
Stan Christiaens Jul 13, 2017 ▶ 21:58
Insight
Corporate data ownership must be snuck in to avoid executive reluctance
“Nobody wants to take responsibility. No, no, I don't want to be the data owner. So there, what you have to do, is you have to sort of sneak ownership under the door. You have to sort of say to the business executive, or the process or application owner, yeah, …”
Stan Christiaens Jul 13, 2017 ▶ 22:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.