Oct 17, 2018 · 44m · mad

Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC)

Mike Tuchen · 32m spoken Matt Turck · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC, Talend CEO Mike Tuchen discusses the evolution of modern enterprise data architectures, open-source commercialization, and operational strategies for scaling a data integration company to a $1.9 billion public valuation. He details how solving the 'first mile' data preparation challenge enables real-time analytics, robust governance, and multi-cloud flexibility across global enterprises.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 13.1% of the talking time here. How this is scored →

Matt as informed peer 2.9 Guest teaching 4.3 Guest disagreement 0.6 Matt pushing back 0.9
05100:0015:0030:000:46–5:03 · Matt as informed peer 3/10 Defining the First Mile Problem in Data Management Matt sets up the discussion by framing industry terms like ETL, data integration, and data fabric. Mike explains the concept of the first mile problem in data management and how software stack constraints have shifted away from the data priesthood.5:03–7:26 · Matt as informed peer 3/10 Data Integration in Action: 360-Degree Customer View Matt prompts for a concrete, non-technical example, referencing an illustration he heard Mike use previously. Mike outlines the 360-degree customer view problem across inconsistent enterprise data stores.7:26–17:37 · Matt as informed peer 6/10 Architecting for Continuous Technological Change Matt demonstrates high domain familiarity by citing Talend's diagram and contrasting legacy engines like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift. When Mike gives a broad pitch on designing for change, Matt pushes back directly asking how technically Talend executes this abstraction.17:37–21:02 · Matt as informed peer 5/10 Open Source Strategy and Freemium Conversion Matt asks if open source is the primary differentiator against legacy competitors like Informatica and IBM and how it scales post-IPO. Mike slightly counters by reframing open source into a broader philosophy of extensibility and openness.21:02–25:50 · Matt as informed peer 3/10 Data Catalogs, Lineage, and Modern Governance Models Matt asks about data governance and lineage as part of the first mile. Mike delivers a detailed breakdown of cataloging, machine learning semantic tags, lineage tracking, and blending top-down and bottom-up governance.25:50–31:40 · Matt as informed peer 5/10 Executive Leadership and Scaling Talend to IPO Matt shifts to executive leadership, asking how to scale a company from $50M ARR to IPO. When Mike offers general leadership principles, Matt pushes for concrete details regarding whether existing executive leadership needed to be replaced.31:40–36:45 · Matt as informed peer 1/10 Q&A: Data Anonymization and GDPR Compliance Audience members ask about GDPR compliance and cloud vendor strategies. Mike provides an overview of data anonymization practices and analyzes multi-cloud dynamics between AWS, Azure, and Google.36:45–39:25 · Matt as informed peer 0/10 Q&A: Debugging Abstractions and ETL Automation Audience members ask about debugging generated code abstractions and bridging organizational data silos. Mike discusses automated error recovery goals and services ecosystem partnerships.39:25–43:51 · Matt as informed peer 0/10 Q&A: Balancing Self-Service Analytics with Governance An audience member asks about maintaining a single version of truth alongside self-service analytics. Mike explains how modern architectures invert the traditional governance model from upfront rigid design to post-hoc incremental governance.0:46–5:03 · Guest teaching 4/10 Defining the First Mile Problem in Data Management Matt sets up the discussion by framing industry terms like ETL, data integration, and data fabric. Mike explains the concept of the first mile problem in data management and how software stack constraints have shifted away from the data priesthood.5:03–7:26 · Guest teaching 3/10 Data Integration in Action: 360-Degree Customer View Matt prompts for a concrete, non-technical example, referencing an illustration he heard Mike use previously. Mike outlines the 360-degree customer view problem across inconsistent enterprise data stores.7:26–17:37 · Guest teaching 5/10 Architecting for Continuous Technological Change Matt demonstrates high domain familiarity by citing Talend's diagram and contrasting legacy engines like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift. When Mike gives a broad pitch on designing for change, Matt pushes back directly asking how technically Talend executes this abstraction.17:37–21:02 · Guest teaching 4/10 Open Source Strategy and Freemium Conversion Matt asks if open source is the primary differentiator against legacy competitors like Informatica and IBM and how it scales post-IPO. Mike slightly counters by reframing open source into a broader philosophy of extensibility and openness.21:02–25:50 · Guest teaching 6/10 Data Catalogs, Lineage, and Modern Governance Models Matt asks about data governance and lineage as part of the first mile. Mike delivers a detailed breakdown of cataloging, machine learning semantic tags, lineage tracking, and blending top-down and bottom-up governance.25:50–31:40 · Guest teaching 4/10 Executive Leadership and Scaling Talend to IPO Matt shifts to executive leadership, asking how to scale a company from $50M ARR to IPO. When Mike offers general leadership principles, Matt pushes for concrete details regarding whether existing executive leadership needed to be replaced.31:40–36:45 · Guest teaching 4/10 Q&A: Data Anonymization and GDPR Compliance Audience members ask about GDPR compliance and cloud vendor strategies. Mike provides an overview of data anonymization practices and analyzes multi-cloud dynamics between AWS, Azure, and Google.36:45–39:25 · Guest teaching 4/10 Q&A: Debugging Abstractions and ETL Automation Audience members ask about debugging generated code abstractions and bridging organizational data silos. Mike discusses automated error recovery goals and services ecosystem partnerships.39:25–43:51 · Guest teaching 5/10 Q&A: Balancing Self-Service Analytics with Governance An audience member asks about maintaining a single version of truth alongside self-service analytics. Mike explains how modern architectures invert the traditional governance model from upfront rigid design to post-hoc incremental governance.0:46–5:03 · Guest disagreement 1/10 Defining the First Mile Problem in Data Management Matt sets up the discussion by framing industry terms like ETL, data integration, and data fabric. Mike explains the concept of the first mile problem in data management and how software stack constraints have shifted away from the data priesthood.5:03–7:26 · Guest disagreement 0/10 Data Integration in Action: 360-Degree Customer View Matt prompts for a concrete, non-technical example, referencing an illustration he heard Mike use previously. Mike outlines the 360-degree customer view problem across inconsistent enterprise data stores.7:26–17:37 · Guest disagreement 1/10 Architecting for Continuous Technological Change Matt demonstrates high domain familiarity by citing Talend's diagram and contrasting legacy engines like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift. When Mike gives a broad pitch on designing for change, Matt pushes back directly asking how technically Talend executes this abstraction.17:37–21:02 · Guest disagreement 2/10 Open Source Strategy and Freemium Conversion Matt asks if open source is the primary differentiator against legacy competitors like Informatica and IBM and how it scales post-IPO. Mike slightly counters by reframing open source into a broader philosophy of extensibility and openness.21:02–25:50 · Guest disagreement 0/10 Data Catalogs, Lineage, and Modern Governance Models Matt asks about data governance and lineage as part of the first mile. Mike delivers a detailed breakdown of cataloging, machine learning semantic tags, lineage tracking, and blending top-down and bottom-up governance.25:50–31:40 · Guest disagreement 1/10 Executive Leadership and Scaling Talend to IPO Matt shifts to executive leadership, asking how to scale a company from $50M ARR to IPO. When Mike offers general leadership principles, Matt pushes for concrete details regarding whether existing executive leadership needed to be replaced.31:40–36:45 · Guest disagreement 0/10 Q&A: Data Anonymization and GDPR Compliance Audience members ask about GDPR compliance and cloud vendor strategies. Mike provides an overview of data anonymization practices and analyzes multi-cloud dynamics between AWS, Azure, and Google.36:45–39:25 · Guest disagreement 0/10 Q&A: Debugging Abstractions and ETL Automation Audience members ask about debugging generated code abstractions and bridging organizational data silos. Mike discusses automated error recovery goals and services ecosystem partnerships.39:25–43:51 · Guest disagreement 0/10 Q&A: Balancing Self-Service Analytics with Governance An audience member asks about maintaining a single version of truth alongside self-service analytics. Mike explains how modern architectures invert the traditional governance model from upfront rigid design to post-hoc incremental governance.0:46–5:03 · Matt pushing back 0/10 Defining the First Mile Problem in Data Management Matt sets up the discussion by framing industry terms like ETL, data integration, and data fabric. Mike explains the concept of the first mile problem in data management and how software stack constraints have shifted away from the data priesthood.5:03–7:26 · Matt pushing back 0/10 Data Integration in Action: 360-Degree Customer View Matt prompts for a concrete, non-technical example, referencing an illustration he heard Mike use previously. Mike outlines the 360-degree customer view problem across inconsistent enterprise data stores.7:26–17:37 · Matt pushing back 3/10 Architecting for Continuous Technological Change Matt demonstrates high domain familiarity by citing Talend's diagram and contrasting legacy engines like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift. When Mike gives a broad pitch on designing for change, Matt pushes back directly asking how technically Talend executes this abstraction.17:37–21:02 · Matt pushing back 1/10 Open Source Strategy and Freemium Conversion Matt asks if open source is the primary differentiator against legacy competitors like Informatica and IBM and how it scales post-IPO. Mike slightly counters by reframing open source into a broader philosophy of extensibility and openness.21:02–25:50 · Matt pushing back 0/10 Data Catalogs, Lineage, and Modern Governance Models Matt asks about data governance and lineage as part of the first mile. Mike delivers a detailed breakdown of cataloging, machine learning semantic tags, lineage tracking, and blending top-down and bottom-up governance.25:50–31:40 · Matt pushing back 4/10 Executive Leadership and Scaling Talend to IPO Matt shifts to executive leadership, asking how to scale a company from $50M ARR to IPO. When Mike offers general leadership principles, Matt pushes for concrete details regarding whether existing executive leadership needed to be replaced.31:40–36:45 · Matt pushing back 0/10 Q&A: Data Anonymization and GDPR Compliance Audience members ask about GDPR compliance and cloud vendor strategies. Mike provides an overview of data anonymization practices and analyzes multi-cloud dynamics between AWS, Azure, and Google.36:45–39:25 · Matt pushing back 0/10 Q&A: Debugging Abstractions and ETL Automation Audience members ask about debugging generated code abstractions and bridging organizational data silos. Mike discusses automated error recovery goals and services ecosystem partnerships.39:25–43:51 · Matt pushing back 0/10 Q&A: Balancing Self-Service Analytics with Governance An audience member asks about maintaining a single version of truth alongside self-service analytics. Mike explains how modern architectures invert the traditional governance model from upfront rigid design to post-hoc incremental governance.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 40.2% · guest 59.8%0:00 · Matt 40.2% · guest 59.8%3:00 · Matt 12.3% · guest 87.7%3:00 · Matt 12.3% · guest 87.7%6:00 · Matt 32.9% · guest 67.1%6:00 · Matt 32.9% · guest 67.1%9:00 · Matt 5.7% · guest 94.3%9:00 · Matt 5.7% · guest 94.3%12:00 · Matt 0.5% · guest 99.5%12:00 · Matt 0.5% · guest 99.5%15:00 · Matt 12.5% · guest 87.5%15:00 · Matt 12.5% · guest 87.5%18:00 · Matt 20.2% · guest 79.8%18:00 · Matt 20.2% · guest 79.8%21:00 · Matt 7.1% · guest 92.9%21:00 · Matt 7.1% · guest 92.9%24:00 · Matt 39.6% · guest 60.4%24:00 · Matt 39.6% · guest 60.4%27:00 · Matt 9.2% · guest 90.8%27:00 · Matt 9.2% · guest 90.8%30:00 · Matt 8.5% · guest 91.5%30:00 · Matt 8.5% · guest 91.5%33:00 · Matt 0% · guest 100%33:00 · Matt 0% · guest 100%36:00 · Matt 0% · guest 100%36:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%42:00 · Matt 6.7% · guest 93.3%42:00 · Matt 6.7% · guest 93.3%
Sharpest disagreement ▶ 18:08 Reframing open source positioning

Mike pushes back on Matt's premise that open source is their main differentiator, clarifying that openness and API extensibility matter more than just open core code.

Hardest push from Matt ▶ 29:44 Demanding concrete executive scaling tactics

Matt refuses to accept high-level leadership platitudes and explicitly challenges Mike to explain if scaling required firing and replacing the existing leadership team.

Biggest teaching moment ▶ 41:30 Inversion of governance paradigms

Mike educates the audience on the structural shift from old schema-on-write data warehouse governance to modern schema-on-read collaborative data lake governance.

Matt holds his own ▶ 7:26 Host demonstrates deep data stack knowledge

Matt demonstrates high technical fluency by walking through Talend's architectural stack diagram and contrasting legacy vendors like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining the First Mile Problem in Data Management 3410 Matt sets up the discussion by framing industry terms like ETL, data integration, and data fabric. Mike explains the concept of the first mile problem in data management and how software stack constraints have shifted away from the data priesthood.
Data Integration in Action: 360-Degree Customer View 3300 Matt prompts for a concrete, non-technical example, referencing an illustration he heard Mike use previously. Mike outlines the 360-degree customer view problem across inconsistent enterprise data stores.
Architecting for Continuous Technological Change 6513 Matt demonstrates high domain familiarity by citing Talend's diagram and contrasting legacy engines like Teradata and Oracle with modern tools like Spark, Snowflake, and Redshift. When Mike gives a broad pitch on designing for change, Matt pushes back directly asking how technically Talend executes this abstraction.
Open Source Strategy and Freemium Conversion 5421 Matt asks if open source is the primary differentiator against legacy competitors like Informatica and IBM and how it scales post-IPO. Mike slightly counters by reframing open source into a broader philosophy of extensibility and openness.
Data Catalogs, Lineage, and Modern Governance Models 3600 Matt asks about data governance and lineage as part of the first mile. Mike delivers a detailed breakdown of cataloging, machine learning semantic tags, lineage tracking, and blending top-down and bottom-up governance.
Executive Leadership and Scaling Talend to IPO 5414 Matt shifts to executive leadership, asking how to scale a company from $50M ARR to IPO. When Mike offers general leadership principles, Matt pushes for concrete details regarding whether existing executive leadership needed to be replaced.
Q&A: Data Anonymization and GDPR Compliance 1400 Audience members ask about GDPR compliance and cloud vendor strategies. Mike provides an overview of data anonymization practices and analyzes multi-cloud dynamics between AWS, Azure, and Google.
Q&A: Debugging Abstractions and ETL Automation 0400 Audience members ask about debugging generated code abstractions and bridging organizational data silos. Mike discusses automated error recovery goals and services ecosystem partnerships.
Q&A: Balancing Self-Service Analytics with Governance 0500 An audience member asks about maintaining a single version of truth alongside self-service analytics. Mike explains how modern architectures invert the traditional governance model from upfront rigid design to post-hoc incremental governance.

Statements from this episode (10)

Assertion Supported
Gartner: Companies spend 60% of their data effort on the first mile
“According to folks like Gartner companies spend about 60% of their time and effort solving that first mile, so that's what we do.”
Mike Tuchen Oct 17, 2018 ▶ 1:46
Disclosure
Talend partnered with Google to transition Dataflow into Apache Beam
“We talked to Google, and they had open sourced a part of that, and said, would you be willing to make that an Apache project, and, you know, co-sponsor it with us, and so we can now jointly push this thing forward? And they said yes, and we did. And so that be…”
Mike Tuchen Oct 17, 2018 ▶ 14:44
Assertion Not checkable as stated
Early Apache Spark was so unstable that clusters crashed within two hours
“At the time we made that bet, Spark was barely working. I mean, seriously. It was a really promising technology, very complete programming model. We loved the fact it was in memory, but we couldn't find a single customer that could keep their cluster running f…”
Mike Tuchen Oct 17, 2018 ▶ 16:31
Assertion Not checkable as stated
Over half of Talend's sales pipeline originates from free tier users
“What we find is that still today well over half of our sales opportunities come from people that are trying it out for free, right?”
Mike Tuchen Oct 17, 2018 ▶ 19:37
Insight
Tuchen: Open source matters less in the cloud era than on-premise
“I firmly believe that open sourceness is actually less important in a cloud world than it was in the premise world.”
Mike Tuchen Oct 17, 2018 ▶ 20:51
Assertion Supported
IDC: Data analysts waste up to three hours daily duplicating work
“Data analysts and data engineers tend to waste around one to three hours per day on duplicating work that someone else has already done.”
Mike Tuchen Oct 17, 2018 ▶ 21:54
Disclosure
Tuchen: Talend will surpass $200 million in revenue in 2018
“And then when I joined Talend it was already a fifty million dollar company, about 350 people. We'll do a little over two hundred million this year. We've got a little over a thousand people.”
Mike Tuchen Oct 17, 2018 ▶ 27:31
Opinion
Tuchen: Microsoft is by far the easiest cloud vendor to partner with
“They're by far the easiest cloud vendor to partner with right now.”
Mike Tuchen Oct 17, 2018 ▶ 35:15
Assertion Partly supported
Tuchen: Google BigQuery is the only fully serverless major cloud data engine
“Google BigQuery is a very impressive data engine. It's probably the only one of the cloud majors that's delivering in a fully serverless kind of way.”
Mike Tuchen Oct 17, 2018 ▶ 35:32
Disclosure
Talend relies on 10 to 20 partner services people per internal employee
“For every one talent services person, there'll probably be 10 or 20 or more partner services people.”
Mike Tuchen Oct 17, 2018 ▶ 38:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.