Jan 16, 2014 · 27m · mad
Scott Sorensen, CTO of Ancestry.com // Data Driven 19 // October 2013 (Hosted by FirstMark Capital)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Scott Sorensen, CTO of Ancestry.com, presents how the company manages petabytes of genealogical data and complex genomic algorithms through big data infrastructure like Hadoop and HBase. He highlights the integration of machine learning, DNA matching, and organizational alignment to deliver personalized historical discoveries at massive scale.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.3% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Scott dismisses conventional management approaches, explicitly stating that holding weekly alignment meetings is a waste of time compared to physical co-location.
Hardest push from Matt ▶ 26:00 Host controlling Q&A time limitHost Matt Turck steps in to control the event schedule, firmly capping audience Q&A by declaring 'One last one'.
Biggest teaching moment ▶ 7:20 Predictive vs. actionable machine learning modelsScott educates the audience on machine learning pitfalls, illustrating how a highly predictive 1,000-feature churn model failed to yield actionable business decisions.
Matt holds his own ▶ 0:01 Host introduction and presentation handoffHost Matt Turck opens the event and sets up Scott Sorensen's talk on Ancestry.com's big data stack.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Ancestry.com Company Overview and Mission | 0 | 2 | 0 | 0 | Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment. | |
| Data Scale, User Content, and Behavioral Algorithms | 0 | 3 | 0 | 0 | Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation. | |
| Graph Mathematics and Structure of Family Trees | 0 | 4 | 0 | 0 | Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue. | |
| AncestryDNA Product Science and Network Effects | 0 | 5 | 0 | 0 | Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment. | |
| Scaling DNA Processing Pipeline with HBase and Hadoop | 0 | 6 | 0 | 0 | Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue. | |
| Organizational Lessons Learned and Audience Q&A | 0 | 5 | 0 | 0 | Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs. |