Jun 12, 2019 · 20m · mad

Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)

FJ Yang · 16m spoken Matt Turck · 36s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At FirstMark's Data Driven NYC, Imply co-founder FJ Yang presents Apache Druid, explaining the paradigm shift toward real-time 'data rivers' and demonstrating how Druid's open-core architecture powers sub-second interactive analytics for continuous streaming data.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 2.8 Guest disagreement 0.0 Matt pushing back 0.3
05100:0010:0020:000:14–4:54 · Matt as informed peer 0/10 FJ Yang's Background and Imply Overview FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction.4:54–10:26 · Matt as informed peer 0/10 Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment.10:26–12:51 · Matt as informed peer 0/10 What is Apache Druid and How It Stores Segments FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment.12:51–17:41 · Matt as informed peer 3/10 Druid Microservice Architecture and Wikipedia Edits Live Demo Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model.0:14–4:54 · Guest teaching 2/10 FJ Yang's Background and Imply Overview FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction.4:54–10:26 · Guest teaching 3/10 Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment.10:26–12:51 · Guest teaching 3/10 What is Apache Druid and How It Stores Segments FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment.12:51–17:41 · Guest teaching 3/10 Druid Microservice Architecture and Wikipedia Edits Live Demo Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model.0:14–4:54 · Guest disagreement 0/10 FJ Yang's Background and Imply Overview FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction.4:54–10:26 · Guest disagreement 0/10 Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment.10:26–12:51 · Guest disagreement 0/10 What is Apache Druid and How It Stores Segments FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment.12:51–17:41 · Guest disagreement 0/10 Druid Microservice Architecture and Wikipedia Edits Live Demo Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model.0:14–4:54 · Matt pushing back 0/10 FJ Yang's Background and Imply Overview FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction.4:54–10:26 · Matt pushing back 0/10 Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment.10:26–12:51 · Matt pushing back 0/10 What is Apache Druid and How It Stores Segments FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment.12:51–17:41 · Matt pushing back 1/10 Druid Microservice Architecture and Wikipedia Edits Live Demo Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 15.6% · guest 84.4%15:00 · Matt 15.6% · guest 84.4%18:00 · Matt 13.5% · guest 86.5%18:00 · Matt 13.5% · guest 86.5%
Sharpest disagreement ▶ 19:32 Clarifying stream replay functionality

FJ gently reframes a question about data loss by noting that people make mistakes all the time and explaining how stream replay solves it.

Hardest push from Matt ▶ 16:27 Matt shifts focus to monetization

Matt intervenes as the presentation concludes to steer the discussion away from technical features and onto commercial open-source sales motion.

Biggest teaching moment ▶ 17:02 Engine versus car open-source business model

FJ educates the audience on open-source commercialization by contrasting Druid as an open engine with Imply as a complete turnkey car.

Matt holds his own ▶ 16:56 Host recognizes board member Martin Casado

Matt displays network familiarity by identifying Andreessen Horowitz partner Martin Casado as an Imply board member.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
FJ Yang's Background and Imply Overview 0200 FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction.
Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers 0300 FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment.
What is Apache Druid and How It Stores Segments 0300 FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment.
Druid Microservice Architecture and Wikipedia Edits Live Demo 3301 Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model.

Statements from this episode (7)

Assertion Not checkable as stated
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 3:29
Prediction Not checkable as stated
FJ Yang: Data architecture is trending toward real-time streaming over batch files
“Where I believe the world is starting to trend to is toward a new world where data is not just batched in a static file, but data is constantly in motion, a constant flow of information.”
FJ Yang Jun 12, 2019 ▶ 8:25
Assertion Not checkable as stated
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
FJ Yang Jun 12, 2019 ▶ 10:28
Assertion Not checkable as stated
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
FJ Yang Jun 12, 2019 ▶ 12:10
Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49
Assertion Supported
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 15:49
Opinion
Yang: Imply's main competitors are legacy data warehouses unsuited for streaming
“The largest competition we probably see are more traditional data warehouses. We see a lot of people trying to shoehorn in various Use cases into legacy data warehouses where, you know, data houses, one, they're not built for, like, kind of live streaming inge…”
FJ Yang Jun 12, 2019 ▶ 20:21
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.