Jun 12, 2019 · 20m · mad
Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At FirstMark's Data Driven NYC, Imply co-founder FJ Yang presents Apache Druid, explaining the paradigm shift toward real-time 'data rivers' and demonstrating how Druid's open-core architecture powers sub-second interactive analytics for continuous streaming data.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
FJ gently reframes a question about data loss by noting that people make mistakes all the time and explaining how stream replay solves it.
Hardest push from Matt ▶ 16:27 Matt shifts focus to monetizationMatt intervenes as the presentation concludes to steer the discussion away from technical features and onto commercial open-source sales motion.
Biggest teaching moment ▶ 17:02 Engine versus car open-source business modelFJ educates the audience on open-source commercialization by contrasting Druid as an open engine with Imply as a complete turnkey car.
Matt holds his own ▶ 16:56 Host recognizes board member Martin CasadoMatt displays network familiarity by identifying Andreessen Horowitz partner Martin Casado as an Imply board member.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| FJ Yang's Background and Imply Overview | 0 | 2 | 0 | 0 | FJ Yang presents his background at Metamarkets and how Apache Druid was created to handle high-volume event data. The segment is a solo presentation with no host interaction. | |
| Evolution of Data Infrastructure: Warehouses, Lakes, and Rivers | 0 | 3 | 0 | 0 | FJ Yang delivers a solo lecture detailing the transition from traditional data warehouses to data lakes and coining the concept of 'data rivers'. The host is completely absent from this segment. | |
| What is Apache Druid and How It Stores Segments | 0 | 3 | 0 | 0 | FJ Yang explains how Druid combines features of data warehouses, time series databases, and search systems into hyper-optimized segments. The host does not speak or participate in this segment. | |
| Druid Microservice Architecture and Wikipedia Edits Live Demo | 3 | 3 | 0 | 1 | Matt Turck enters at 16:27 to conclude the talk demo and redirect the conversation toward VC funding and open-source commercialization. FJ Yang politely explains Imply's product-versus-engine business model. |