Apache Druid

6 statements across 2 episodes · 3 bullish · 0 bearish · 2 people on the record · first statement Jun 12, 2019 by FJ Yang · said 53 times in 6 episodes since 2019 · across every show →

Mentions by year

brought up most by FJ Yang (37), Matt Turck (5), Dave Burgess (4), Jack Hanlon (3), Aaron Katz (3), DeVaris Brown (1)

tap a year for its mentions
00202404201920202021202220232024episodesmentions
024201920202021202220232024episodes it came up in
00202404201920202021202220232024episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Apache Druid, oldest first

Jun 12, 2019 positive
Assertion Not checkable as stated
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
FJ Yang Jun 12, 2019 ▶ 12:10 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Jun 12, 2019 positive
Assertion Not checkable as stated
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
FJ Yang Jun 12, 2019 ▶ 10:28 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Jun 12, 2019
Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Jun 12, 2019
Assertion Not checkable as stated
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 3:29 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Jun 12, 2019
Assertion Supported
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 15:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Apr 5, 2021 positive
Assertion Not checkable as stated
Pinterest reduced maintenance costs, latency, and infrastructure expenses by migrating to Druid
“And so we decided to migrate that to Druid. And so the maintenance cost has gone way down. The latency has gone way down. The actual cost of running the infrastructure way down. So really, really happy.”
Dave Burgess Apr 5, 2021 ▶ 27:25 Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.