Jul 13, 2017 · 28m · mad

Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven)

Sean Kandel · 23m spoken Holmer Gislason · 51s spoken Matt Turck · 39s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this presentation at Data Driven NYC, Sean Kandel, CTO and co-founder of Trifacta, demonstrates how organizations can eliminate data preparation bottlenecks by transitioning from dependency, batch, and disconnected loops to self-service, interactive, and networked analytics workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.3% of the talking time here. How this is scored →

Matt as informed peer 1.0 Guest teaching 2.2 Guest disagreement 0.2 Matt pushing back 0.8
05100:0010:0020:000:29–4:34 · Matt as informed peer 0/10 The Evolution of Data Infrastructure and the Data Prep Bottleneck Sean Kandel delivers a uninterrupted presentation detailing how data infrastructure advancements have shifted the primary analytics bottleneck to data preparation. The host does not participate in this monologue segment.4:34–8:37 · Matt as informed peer 0/10 Overcoming Dependency Loops Through Self-Service Data Preparation Sean presents the concept of dependency loops and explains how self-service predictive tools replace slow reliance on technical colleagues. The host remains silent throughout the slide presentation.8:37–12:14 · Matt as informed peer 0/10 Replacing Slow Batch Processing with Real-Time Interactive Feedback Sean contrasts slow batch processing with real-time interactive visual feedback to reduce the cost of change in data cleaning. Host engagement is zero during this monologue.12:14–19:05 · Matt as informed peer 0/10 Connecting Organizational Silos Through Shared Metadata and Networked Loops Sean outlines how disconnected organizational silos create duplicated work and how automated metadata catalogs build networked feedback loops. The host does not intervene.19:05–28:54 · Matt as informed peer 5/10 Audience Q&A on Data Engineering Roles, Missing Data, and Tool Interoperability Host Matt Turck opens Q&A by pressing Sean on whether non-technical analysts can truly replace deep data engineers or if complex production pipelines will always require specialists. The segment remains polite, constructive, and collaborative throughout.0:29–4:34 · Guest teaching 2/10 The Evolution of Data Infrastructure and the Data Prep Bottleneck Sean Kandel delivers a uninterrupted presentation detailing how data infrastructure advancements have shifted the primary analytics bottleneck to data preparation. The host does not participate in this monologue segment.4:34–8:37 · Guest teaching 2/10 Overcoming Dependency Loops Through Self-Service Data Preparation Sean presents the concept of dependency loops and explains how self-service predictive tools replace slow reliance on technical colleagues. The host remains silent throughout the slide presentation.8:37–12:14 · Guest teaching 2/10 Replacing Slow Batch Processing with Real-Time Interactive Feedback Sean contrasts slow batch processing with real-time interactive visual feedback to reduce the cost of change in data cleaning. Host engagement is zero during this monologue.12:14–19:05 · Guest teaching 2/10 Connecting Organizational Silos Through Shared Metadata and Networked Loops Sean outlines how disconnected organizational silos create duplicated work and how automated metadata catalogs build networked feedback loops. The host does not intervene.19:05–28:54 · Guest teaching 3/10 Audience Q&A on Data Engineering Roles, Missing Data, and Tool Interoperability Host Matt Turck opens Q&A by pressing Sean on whether non-technical analysts can truly replace deep data engineers or if complex production pipelines will always require specialists. The segment remains polite, constructive, and collaborative throughout.0:29–4:34 · Guest disagreement 0/10 The Evolution of Data Infrastructure and the Data Prep Bottleneck Sean Kandel delivers a uninterrupted presentation detailing how data infrastructure advancements have shifted the primary analytics bottleneck to data preparation. The host does not participate in this monologue segment.4:34–8:37 · Guest disagreement 0/10 Overcoming Dependency Loops Through Self-Service Data Preparation Sean presents the concept of dependency loops and explains how self-service predictive tools replace slow reliance on technical colleagues. The host remains silent throughout the slide presentation.8:37–12:14 · Guest disagreement 0/10 Replacing Slow Batch Processing with Real-Time Interactive Feedback Sean contrasts slow batch processing with real-time interactive visual feedback to reduce the cost of change in data cleaning. Host engagement is zero during this monologue.12:14–19:05 · Guest disagreement 0/10 Connecting Organizational Silos Through Shared Metadata and Networked Loops Sean outlines how disconnected organizational silos create duplicated work and how automated metadata catalogs build networked feedback loops. The host does not intervene.19:05–28:54 · Guest disagreement 1/10 Audience Q&A on Data Engineering Roles, Missing Data, and Tool Interoperability Host Matt Turck opens Q&A by pressing Sean on whether non-technical analysts can truly replace deep data engineers or if complex production pipelines will always require specialists. The segment remains polite, constructive, and collaborative throughout.0:29–4:34 · Matt pushing back 0/10 The Evolution of Data Infrastructure and the Data Prep Bottleneck Sean Kandel delivers a uninterrupted presentation detailing how data infrastructure advancements have shifted the primary analytics bottleneck to data preparation. The host does not participate in this monologue segment.4:34–8:37 · Matt pushing back 0/10 Overcoming Dependency Loops Through Self-Service Data Preparation Sean presents the concept of dependency loops and explains how self-service predictive tools replace slow reliance on technical colleagues. The host remains silent throughout the slide presentation.8:37–12:14 · Matt pushing back 0/10 Replacing Slow Batch Processing with Real-Time Interactive Feedback Sean contrasts slow batch processing with real-time interactive visual feedback to reduce the cost of change in data cleaning. Host engagement is zero during this monologue.12:14–19:05 · Matt pushing back 0/10 Connecting Organizational Silos Through Shared Metadata and Networked Loops Sean outlines how disconnected organizational silos create duplicated work and how automated metadata catalogs build networked feedback loops. The host does not intervene.19:05–28:54 · Matt pushing back 4/10 Audience Q&A on Data Engineering Roles, Missing Data, and Tool Interoperability Host Matt Turck opens Q&A by pressing Sean on whether non-technical analysts can truly replace deep data engineers or if complex production pipelines will always require specialists. The segment remains polite, constructive, and collaborative throughout.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 16.1% · guest 83.9%18:00 · Matt 16.1% · guest 83.9%21:00 · Matt 1.5% · guest 98.5%21:00 · Matt 1.5% · guest 98.5%24:00 · Matt 5.2% · guest 94.8%24:00 · Matt 5.2% · guest 94.8%27:00 · Matt 0.6% · guest 99.4%27:00 · Matt 0.6% · guest 99.4%
Sharpest disagreement ▶ 26:27 Audience member challenges self-service premise

Holmer Gislason offers a mild contrarian perspective, arguing that analysts prefer having data prep done by others and pointing out that tools remain isolated islands.

Hardest push from Matt ▶ 20:03 Matt Turck reframes question on technical roles

Matt Turck refuses to accept Sean's broad claim of full analyst self-service, pressing further on whether deep technical data scientists will always be needed for advanced prep.

Biggest teaching moment ▶ 20:19 Sean clarifies exploratory versus production pipelines

Sean educates the host on the distinction between initial exploratory data preparation and production hardening, which still requires specialized data engineering.

Matt holds his own ▶ 20:03 Matt Turck demonstrates domain knowledge on role boundaries

Matt Turck demonstrates industry awareness by distinguishing between business analysts and deep technical data scientists, refining his inquiry to challenge oversimplified claims.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Evolution of Data Infrastructure and the Data Prep Bottleneck 0200 Sean Kandel delivers a uninterrupted presentation detailing how data infrastructure advancements have shifted the primary analytics bottleneck to data preparation. The host does not participate in this monologue segment.
Overcoming Dependency Loops Through Self-Service Data Preparation 0200 Sean presents the concept of dependency loops and explains how self-service predictive tools replace slow reliance on technical colleagues. The host remains silent throughout the slide presentation.
Replacing Slow Batch Processing with Real-Time Interactive Feedback 0200 Sean contrasts slow batch processing with real-time interactive visual feedback to reduce the cost of change in data cleaning. Host engagement is zero during this monologue.
Connecting Organizational Silos Through Shared Metadata and Networked Loops 0200 Sean outlines how disconnected organizational silos create duplicated work and how automated metadata catalogs build networked feedback loops. The host does not intervene.
Audience Q&A on Data Engineering Roles, Missing Data, and Tool Interoperability 5314 Host Matt Turck opens Q&A by pressing Sean on whether non-technical analysts can truly replace deep data engineers or if complex production pipelines will always require specialists. The segment remains polite, constructive, and collaborative throughout.

Statements from this episode (9)

Assertion Not checkable as stated
Sean Kandel: Machine learning and statistical tools are becoming commoditized
“Machine learning and even statistical tools are becoming kind of commoditized.”
Sean Kandel Jul 13, 2017 ▶ 1:20
Opinion
Sean Kandel: Data preparation is the primary bottleneck in analytics
“We kind of think this is the new bottleneck in analytics. Broadly speaking now is, you know, what people refer to as the data preparation space.”
Sean Kandel Jul 13, 2017 ▶ 1:59
Assertion Not checkable as stated
Kandel: Typical organizations have few employees skilled in data preparation
“In a typical organization there's usually only kind of a small subset of people who have the skills necessary to prepare data.”
Sean Kandel Jul 13, 2017 ▶ 5:04
Assertion Not checkable as stated
Kandel: Hand-coding tools remain the most common data preparation method
“So probably still today the most common is using kind of hand coding tools so programming languages, Python, Spark SAS and so on.”
Sean Kandel Jul 13, 2017 ▶ 6:41
Insight
Kandel: Analytics tools lag behind software development in collaboration
“I'd say generally if you look at kind of software development process, there's been a lot more in terms of development there and aiding kind of collaboration across users and teams. So I think generally speaking with analytics, there's still a long way to go.”
Sean Kandel Jul 13, 2017 ▶ 16:12
Insight
Kandel: Automated metadata capture is essential to scale analytics efficiency
“I think one way to kind of really speed up this loop and make it much more efficient is instead of kind of relying on others to kind of manually capture metadata and manually leverage it, having systems that automatically capture it and automatically leverage …”
Sean Kandel Jul 13, 2017 ▶ 18:11
Prediction Not checkable as stated
Trifacta CTO: No data prep algorithm will ever be 100% correct
“We don't feel that the you know, algorithms or predictions that we provide are ever going to be a hundred percent correct, or any algorithm will be.”
Sean Kandel Jul 13, 2017 ▶ 25:39
Insight
Qlik VP Holmer Gislason: The best data prep is done by someone else
“Because the best data preparation is data preparation done by someone else.”
Holmer Gislason Jul 13, 2017 ▶ 26:52
Insight
Sean Kandel: Locking in internal metadata makes software interoperability nearly impossible
“So I think the important piece is not really locking in your internal metadata. Otherwise, it becomes incredibly hard to interoperate.”
Sean Kandel Jul 13, 2017 ▶ 27:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.