Prediction certainty 3/5 debate potential 2/5

FJ Yang: Data architecture is trending toward real-time streaming over batch files

FJ Yang · Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC) · Jun 12, 2019 · at 8:25

FJ Yang, co-creator of Apache Druid and CEO of Imply, speaks at FirstMark's Data Driven NYC about the evolution of data infrastructure from warehouses to real-time 'data rivers'.

0:00 / 0:10exact quote · 10.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Where I believe the world is starting to trend to is toward a new world where data is not just batched in a static file, but data is constantly in motion, a constant flow of information.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from FJ Yang

Opinion
Yang: Imply's main competitors are legacy data warehouses unsuited for streaming
“The largest competition we probably see are more traditional data warehouses. We see a lot of people trying to shoehorn in various Use cases into legacy data warehouses where, you know, data houses, one, they're not built for, like, kind of live streaming inge…”
FJ Yang Jun 12, 2019 ▶ 20:21 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
FJ Yang Jun 12, 2019 ▶ 10:28 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
FJ Yang Jun 12, 2019 ▶ 12:10 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 3:29 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Supported
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 15:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.