Mar 15, 2021 · 29m · mad

Fireside Chat: Jack Hanlon (VP Data, Reddit) with Matt Turck (Partner, FirstMark)

Jack Hanlon · 22m spoken Matt Turck · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this fireside chat hosted by FirstMark's Matt Turck, Jack Hanlon, VP of Data at Reddit, shares insights on scaling Reddit's data organization, engineering robust hybrid infrastructure, and advancing ethical AI and privacy practices.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 14.9% of the talking time here. How this is scored →

Matt as informed peer 3.1 Guest teaching 4.6 Guest disagreement 0.4 Matt pushing back 0.3
05100:0010:0020:000:11–2:53 · Matt as informed peer 1/10 Jack Hanlon's Career Path to Reddit Host opens with a standard, open-ended question about the guest's career progression to Reddit. Guest shares a non-linear path from music major through sales, ad tech startups, and Jet.com to advising Reddit.2:53–5:45 · Matt as informed peer 2/10 Overview of Reddit's Four Data Organizations Host asks organizational questions about how data is structured at Reddit. Guest breaks down the four core data groups and the hub-and-spoke matrix model.5:45–9:00 · Matt as informed peer 3/10 Roadmap Prioritization and Scaling Data Engineering Host probes into prioritization and interrupts to clarify headcount numbers. Guest reveals the surprising fact that Reddit had only one data engineer managing billions of daily events when he joined.9:00–14:35 · Matt as informed peer 5/10 Reddit's Data Engineering Stack and Quality Tools Host asks targeted questions about Reddit's tech stack across BI, notebooks, and data lineage. Guest acknowledges the host's market expertise regarding venture-backed data startup pressures.14:35–17:24 · Matt as informed peer 4/10 Great Expectations Unit Testing and Streaming Pipelines Host relays technical audience questions regarding Great Expectations and ETL pipelines. Guest playfully notes the audience picked his weakest area before explaining data unit testing principles.17:24–22:58 · Matt as informed peer 4/10 Machine Learning Applications, Personalization, and Anonymity Host prompts guest on specific machine learning applications and infrastructure choices like Imply and Druid. Guest explains how Reddit's unique conversational corpus was leveraged by major AI labs while preserving user anonymity.22:58–27:18 · Matt as informed peer 4/10 Infrastructure Migration from AWS Redshift to GCP Host relays audience questions on cloud migrations and AI bias. Guest gives a candid evaluation rejecting Redshift's capabilities and offers an insightful lecture on framework strategies for preventing machine learning bias.27:18–29:10 · Matt as informed peer 2/10 Future Data Trends, Ethical Design, and Conclusion Host transitions to quick closing questions on industry trends and recommendations. Guest compares data product engineering ethics to civil engineering standards for bridge safety.0:11–2:53 · Guest teaching 2/10 Jack Hanlon's Career Path to Reddit Host opens with a standard, open-ended question about the guest's career progression to Reddit. Guest shares a non-linear path from music major through sales, ad tech startups, and Jet.com to advising Reddit.2:53–5:45 · Guest teaching 4/10 Overview of Reddit's Four Data Organizations Host asks organizational questions about how data is structured at Reddit. Guest breaks down the four core data groups and the hub-and-spoke matrix model.5:45–9:00 · Guest teaching 5/10 Roadmap Prioritization and Scaling Data Engineering Host probes into prioritization and interrupts to clarify headcount numbers. Guest reveals the surprising fact that Reddit had only one data engineer managing billions of daily events when he joined.9:00–14:35 · Guest teaching 5/10 Reddit's Data Engineering Stack and Quality Tools Host asks targeted questions about Reddit's tech stack across BI, notebooks, and data lineage. Guest acknowledges the host's market expertise regarding venture-backed data startup pressures.14:35–17:24 · Guest teaching 4/10 Great Expectations Unit Testing and Streaming Pipelines Host relays technical audience questions regarding Great Expectations and ETL pipelines. Guest playfully notes the audience picked his weakest area before explaining data unit testing principles.17:24–22:58 · Guest teaching 6/10 Machine Learning Applications, Personalization, and Anonymity Host prompts guest on specific machine learning applications and infrastructure choices like Imply and Druid. Guest explains how Reddit's unique conversational corpus was leveraged by major AI labs while preserving user anonymity.22:58–27:18 · Guest teaching 7/10 Infrastructure Migration from AWS Redshift to GCP Host relays audience questions on cloud migrations and AI bias. Guest gives a candid evaluation rejecting Redshift's capabilities and offers an insightful lecture on framework strategies for preventing machine learning bias.27:18–29:10 · Guest teaching 4/10 Future Data Trends, Ethical Design, and Conclusion Host transitions to quick closing questions on industry trends and recommendations. Guest compares data product engineering ethics to civil engineering standards for bridge safety.0:11–2:53 · Guest disagreement 0/10 Jack Hanlon's Career Path to Reddit Host opens with a standard, open-ended question about the guest's career progression to Reddit. Guest shares a non-linear path from music major through sales, ad tech startups, and Jet.com to advising Reddit.2:53–5:45 · Guest disagreement 0/10 Overview of Reddit's Four Data Organizations Host asks organizational questions about how data is structured at Reddit. Guest breaks down the four core data groups and the hub-and-spoke matrix model.5:45–9:00 · Guest disagreement 0/10 Roadmap Prioritization and Scaling Data Engineering Host probes into prioritization and interrupts to clarify headcount numbers. Guest reveals the surprising fact that Reddit had only one data engineer managing billions of daily events when he joined.9:00–14:35 · Guest disagreement 0/10 Reddit's Data Engineering Stack and Quality Tools Host asks targeted questions about Reddit's tech stack across BI, notebooks, and data lineage. Guest acknowledges the host's market expertise regarding venture-backed data startup pressures.14:35–17:24 · Guest disagreement 1/10 Great Expectations Unit Testing and Streaming Pipelines Host relays technical audience questions regarding Great Expectations and ETL pipelines. Guest playfully notes the audience picked his weakest area before explaining data unit testing principles.17:24–22:58 · Guest disagreement 0/10 Machine Learning Applications, Personalization, and Anonymity Host prompts guest on specific machine learning applications and infrastructure choices like Imply and Druid. Guest explains how Reddit's unique conversational corpus was leveraged by major AI labs while preserving user anonymity.22:58–27:18 · Guest disagreement 2/10 Infrastructure Migration from AWS Redshift to GCP Host relays audience questions on cloud migrations and AI bias. Guest gives a candid evaluation rejecting Redshift's capabilities and offers an insightful lecture on framework strategies for preventing machine learning bias.27:18–29:10 · Guest disagreement 0/10 Future Data Trends, Ethical Design, and Conclusion Host transitions to quick closing questions on industry trends and recommendations. Guest compares data product engineering ethics to civil engineering standards for bridge safety.0:11–2:53 · Matt pushing back 0/10 Jack Hanlon's Career Path to Reddit Host opens with a standard, open-ended question about the guest's career progression to Reddit. Guest shares a non-linear path from music major through sales, ad tech startups, and Jet.com to advising Reddit.2:53–5:45 · Matt pushing back 0/10 Overview of Reddit's Four Data Organizations Host asks organizational questions about how data is structured at Reddit. Guest breaks down the four core data groups and the hub-and-spoke matrix model.5:45–9:00 · Matt pushing back 1/10 Roadmap Prioritization and Scaling Data Engineering Host probes into prioritization and interrupts to clarify headcount numbers. Guest reveals the surprising fact that Reddit had only one data engineer managing billions of daily events when he joined.9:00–14:35 · Matt pushing back 1/10 Reddit's Data Engineering Stack and Quality Tools Host asks targeted questions about Reddit's tech stack across BI, notebooks, and data lineage. Guest acknowledges the host's market expertise regarding venture-backed data startup pressures.14:35–17:24 · Matt pushing back 0/10 Great Expectations Unit Testing and Streaming Pipelines Host relays technical audience questions regarding Great Expectations and ETL pipelines. Guest playfully notes the audience picked his weakest area before explaining data unit testing principles.17:24–22:58 · Matt pushing back 0/10 Machine Learning Applications, Personalization, and Anonymity Host prompts guest on specific machine learning applications and infrastructure choices like Imply and Druid. Guest explains how Reddit's unique conversational corpus was leveraged by major AI labs while preserving user anonymity.22:58–27:18 · Matt pushing back 0/10 Infrastructure Migration from AWS Redshift to GCP Host relays audience questions on cloud migrations and AI bias. Guest gives a candid evaluation rejecting Redshift's capabilities and offers an insightful lecture on framework strategies for preventing machine learning bias.27:18–29:10 · Matt pushing back 0/10 Future Data Trends, Ethical Design, and Conclusion Host transitions to quick closing questions on industry trends and recommendations. Guest compares data product engineering ethics to civil engineering standards for bridge safety.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 10.4% · guest 89.6%0:00 · Matt 10.4% · guest 89.6%3:00 · Matt 13.7% · guest 86.3%3:00 · Matt 13.7% · guest 86.3%6:00 · Matt 12.1% · guest 87.9%6:00 · Matt 12.1% · guest 87.9%9:00 · Matt 15.6% · guest 84.4%9:00 · Matt 15.6% · guest 84.4%12:00 · Matt 18.7% · guest 81.3%12:00 · Matt 18.7% · guest 81.3%15:00 · Matt 15.1% · guest 84.9%15:00 · Matt 15.1% · guest 84.9%18:00 · Matt 2.3% · guest 97.7%18:00 · Matt 2.3% · guest 97.7%21:00 · Matt 19.4% · guest 80.6%21:00 · Matt 19.4% · guest 80.6%24:00 · Matt 14.5% · guest 85.5%24:00 · Matt 14.5% · guest 85.5%27:00 · Matt 32% · guest 68%27:00 · Matt 32% · guest 68%
Sharpest disagreement ▶ 23:01 Blunt dismissal of AWS Redshift architecture

Guest forcefully rejects AWS Redshift as inadequate for Reddit's needs, plainly stating that its core architecture has failed to keep up with modern data demands.

Hardest push from Matt ▶ 7:16 Host presses guest on company headcount context

Host interrupts to explicitly question and clarify the scale of Reddit's total company headcount relative to its single data engineer at the time.

Biggest teaching moment ▶ 24:35 Masterclass on detecting and mitigating algorithmic data bias

Guest provides an in-depth breakdown on systemic data collection bias, referencing Carolyn Criado Perez's book Invisible Women and outlining practical auditing frameworks.

Matt holds his own ▶ 13:35 Host's domain expertise acknowledged by guest

Host demonstrates deep knowledge of data lineage and market trends, prompting the guest to explicitly defer to Matt's superior market knowledge on venture startup behavior.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Jack Hanlon's Career Path to Reddit 1200 Host opens with a standard, open-ended question about the guest's career progression to Reddit. Guest shares a non-linear path from music major through sales, ad tech startups, and Jet.com to advising Reddit.
Overview of Reddit's Four Data Organizations 2400 Host asks organizational questions about how data is structured at Reddit. Guest breaks down the four core data groups and the hub-and-spoke matrix model.
Roadmap Prioritization and Scaling Data Engineering 3501 Host probes into prioritization and interrupts to clarify headcount numbers. Guest reveals the surprising fact that Reddit had only one data engineer managing billions of daily events when he joined.
Reddit's Data Engineering Stack and Quality Tools 5501 Host asks targeted questions about Reddit's tech stack across BI, notebooks, and data lineage. Guest acknowledges the host's market expertise regarding venture-backed data startup pressures.
Great Expectations Unit Testing and Streaming Pipelines 4410 Host relays technical audience questions regarding Great Expectations and ETL pipelines. Guest playfully notes the audience picked his weakest area before explaining data unit testing principles.
Machine Learning Applications, Personalization, and Anonymity 4600 Host prompts guest on specific machine learning applications and infrastructure choices like Imply and Druid. Guest explains how Reddit's unique conversational corpus was leveraged by major AI labs while preserving user anonymity.
Infrastructure Migration from AWS Redshift to GCP 4720 Host relays audience questions on cloud migrations and AI bias. Guest gives a candid evaluation rejecting Redshift's capabilities and offers an insightful lecture on framework strategies for preventing machine learning bias.
Future Data Trends, Ethical Design, and Conclusion 2400 Host transitions to quick closing questions on industry trends and recommendations. Guest compares data product engineering ethics to civil engineering standards for bridge safety.

Statements from this episode (12)

Disclosure
Hanlon: Reddit uses a hub-and-spoke model for data science
“Data science is hub and spoke. So so effectively that is all of those folks report into the director of data science, who's, who's one of my team members and, but they actually, when we're in a, when we're in an office setting, they sit with the downstream tea…”
Jack Hanlon Mar 15, 2021 ▶ 4:38
Assertion Not checkable as stated
Hanlon: Reddit grew from around 600 to 800 employees in 14 months
“When I got to Reddit overall, I'd say Reddit was maybe 600 people 14 months ago. Right now, oh, maybe 800.”
Jack Hanlon Mar 15, 2021 ▶ 7:34
Prediction Not checkable as stated
Hanlon: Reddit's data team will probably reach ~120 people by end of 2021
“By the end of the year, the data organization I just described will be probably the third largest organization at Reddit with a 120 or so people.”
Jack Hanlon Mar 15, 2021 ▶ 7:43
Assertion Not checkable as stated
Hanlon: Reddit ingested 55-60B daily events managed by one person in 2019
“I came in, we were doing probably about 55 or sixty billion events a day into a data warehouse that one person was managing.”
Jack Hanlon Mar 15, 2021 ▶ 8:44
Disclosure
Reddit runs production on AWS and analytics/ML on GCP
“From the top view, I'd say we're AWS prod with originally a lot of Postgres For a variety of prod systems, we actually transit almost all that stuff over to GCP, and we use TensorFlow and BigQuery for our analytics and our model concerns”
Jack Hanlon Mar 15, 2021 ▶ 9:36
Insight
Data startups push horizontal platforms, but enterprise buyers want single features
“The companies I see, it feels like a number of the startups I see in the space get pressure to expand to be more horizontal solutions. And then frankly, like we want them for one thing, but not for the three other things that they want to do.”
Jack Hanlon Mar 15, 2021 ▶ 13:45
Disclosure
Reddit isolates malformed data events in an automated penalty box
“You can do stuff for event validation to say, ah, events that don't match certain things get dropped and sort of get put in the penalty box, which we've done and said, okay, developers, you can't actually join these to anything. You can't actually use them if …”
Jack Hanlon Mar 15, 2021 ▶ 15:32
Disclosure
Reddit is leaning heavily into becoming a Kafka shop for streaming
“So, you know, I think we're ending up leaning heavily into being a Kafka shop, building more of the stuff into base plate, into more of these core services”
Jack Hanlon Mar 15, 2021 ▶ 17:00
Assertion Partly supported
Google, Facebook, Microsoft, and OpenAI used Reddit data for AI models
“Google, Facebook, Microsoft, and OpenAI all used Reddit's data to train their conversational AI models.”
Jack Hanlon Mar 15, 2021 ▶ 17:52
Assertion Not checkable as stated
Jack Hanlon: 2020 was the first year every Reddit surface implemented personalization
“Last year was the first year that every Reddit surface saw personalization.”
Jack Hanlon Mar 15, 2021 ▶ 19:00
Disclosure
Reddit's 2021 home feed will include recommendations beyond user subscriptions
“You'll see this year that the home feed and people's home feed will be breaking the subscription wall and having recommendations in there for people.”
Jack Hanlon Mar 15, 2021 ▶ 19:36
Opinion
Hanlon: AWS Redshift's core architecture has fallen behind competitors
“Redshift feels way behind, whereas a number of years ago, it was fantastic. The core architecture there has not kept up.”
Jack Hanlon Mar 15, 2021 ▶ 23:10
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.