FJ Yang

Co-Founder and CEO, Imply · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveengineer@fangjin ↗imply.io ↗

Fangjin "FJ" Yang is one of the original creators and authors of Apache Druid, an open-source distributed real-time analytics database. In 2015, he co-founded Imply to commercialize Apache Druid and serves as its CEO.

7statements → 6claims → 2claims resolved → 3.57/5average certainty → 1.71/5average debate potential →

2 supported 0 partly supported 0 contradicted 4 not checkable as stated how the 6 claims stand · each chip opens the sources

1 prediction · 5 assertions · 1 opinion · every statement was checked. The prediction and assertions are the 6 claims: statements the public record can support or contradict. 2 are resolved, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how FJ argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)

How they sound: speaking style how? →

255 words/min while actually speaking · 42.3 um and uh per 1k words

No argument clarity score for FJ Yang: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 3,475 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything FJ Yang said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Yang: Imply's main competitors are legacy data warehouses unsuited for streaming
“The largest competition we probably see are more traditional data warehouses. We see a lot of people trying to shoehorn in various Use cases into legacy data warehouses where, you know, data houses, one, they're not built for, like, kind of live streaming inge…”
FJ Yang Jun 12, 2019 ▶ 20:21 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Prediction Not checkable as stated
FJ Yang: Data architecture is trending toward real-time streaming over batch files
“Where I believe the world is starting to trend to is toward a new world where data is not just batched in a static file, but data is constantly in motion, a constant flow of information.”
FJ Yang Jun 12, 2019 ▶ 8:25 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid merges data warehouses, time series, and search systems
“Druid is a combination of a data warehouse merged with a time series database merged with a search system.”
FJ Yang Jun 12, 2019 ▶ 10:28 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid can condense raw data 100x through roll-up aggregation
“If you have raw data that's, you know, a hundred gigabytes, Druid can sometimes condense it through roll-up to about a gigabyte in size, so a hundred X reduction.”
FJ Yang Jun 12, 2019 ▶ 12:10 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Not checkable as stated
Yang: Apache Druid was created to handle hundreds of billions of daily events
“It was created because the volume of data we were dealing with was reaching millions of events per second, and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 3:29 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Supported
Yang: Wikimedia Foundation uses Apache Druid for internal analytics
“And this is actually how the Wikimedia Foundation itself does a lot of internal analytics on who's editing what on Wikipedia.”
FJ Yang Jun 12, 2019 ▶ 13:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)
Assertion Supported
Yang: Production Druid clusters process tens of millions of events per second
“Companies today in production have you know, deployed Druid clusters to handle tens of millions of events per second and hundreds of billions of events per day.”
FJ Yang Jun 12, 2019 ▶ 15:49 Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC)

Appearances (1)

EpisodeDateSpeaking time
Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven N Jun 12, 2019 16m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.