Everything FJ Yang said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Yang: Imply's main competitors are legacy data warehouses unsuited for streaming
“The largest competition we probably see are more traditional data warehouses. We see a lot of people trying to shoehorn in various Use cases into legacy data warehouses where, you know, data houses, one, they're not built for, like, kind of live streaming inge…”
FJ Yang: Data architecture is trending toward real-time streaming over batch files
“Where I believe the world is starting to trend to is toward a new world where data is not just batched in a static file, but data is constantly in motion, a constant flow of information.”
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”