Apache Druid

product on 2 shows · 6 statements across 2 episodes · said 56 times in 7 episodes since 2019

the MAD Podcast 53 the a16z Podcast 3

Mentions by year, every show

tap a year for its mentions
00202404201920202021202220232024episodesmentions
024201920202021202220232024episodes it came up in
00102204201920202021202220232024episodesmentions per episode

the MAD Podcast 53the a16z Podcast 3

2024 3 mentions in 1 episode
2021 13 mentions in 4 episodes 3 per episode
2019 40 mentions in 2 episodes 20 per episode

every mention on every show, scene by scene, with the transcript →

6 statements about Apache Druid, every show

MAD Assertion Not checkable as stated
Pinterest reduced maintenance costs, latency, and infrastructure expenses by migrating to Druid
“And so we decided to migrate that to Druid. And so the maintenance cost has gone way down. The latency has gone way down. The actual cost of running the infrastructure way down. So really, really happy.”
Dave Burgess Apr 5, 2021 ▶ 27:25 Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
MAD Assertion Not checkable as stated
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 3:29 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
MAD Assertion Not checkable as stated
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
FJ Yang Jun 12, 2019 ▶ 10:28 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
MAD Assertion Not checkable as stated
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
FJ Yang Jun 12, 2019 ▶ 12:10 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
MAD Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
MAD Assertion Supported
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 15:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.