Jan 16, 2014 · 27m · mad

Scott Sorensen, CTO of Ancestry.com // Data Driven 19 // October 2013 (Hosted by FirstMark Capital)

Scott Sorensen · 23m spoken Dave Machala · 14s spoken Margaret Oest · 11s spoken Serge Gekker · 9s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Scott Sorensen, CTO of Ancestry.com, presents how the company manages petabytes of genealogical data and complex genomic algorithms through big data infrastructure like Hadoop and HBase. He highlights the integration of machine learning, DNA matching, and organizational alignment to deliver personalized historical discoveries at massive scale.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.3% of the talking time here. How this is scored →

Matt as informed peer 0.0 Guest teaching 4.2 Guest disagreement 0.0 Matt pushing back 0.0
05100:0010:0020:000:12–3:32 · Matt as informed peer 0/10 Ancestry.com Company Overview and Mission Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment.3:32–5:49 · Matt as informed peer 0/10 Data Scale, User Content, and Behavioral Algorithms Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation.5:49–9:53 · Matt as informed peer 0/10 Graph Mathematics and Structure of Family Trees Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue.9:53–12:47 · Matt as informed peer 0/10 AncestryDNA Product Science and Network Effects Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment.12:47–20:52 · Matt as informed peer 0/10 Scaling DNA Processing Pipeline with HBase and Hadoop Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue.20:52–27:11 · Matt as informed peer 0/10 Organizational Lessons Learned and Audience Q&A Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs.0:12–3:32 · Guest teaching 2/10 Ancestry.com Company Overview and Mission Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment.3:32–5:49 · Guest teaching 3/10 Data Scale, User Content, and Behavioral Algorithms Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation.5:49–9:53 · Guest teaching 4/10 Graph Mathematics and Structure of Family Trees Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue.9:53–12:47 · Guest teaching 5/10 AncestryDNA Product Science and Network Effects Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment.12:47–20:52 · Guest teaching 6/10 Scaling DNA Processing Pipeline with HBase and Hadoop Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue.20:52–27:11 · Guest teaching 5/10 Organizational Lessons Learned and Audience Q&A Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs.0:12–3:32 · Guest disagreement 0/10 Ancestry.com Company Overview and Mission Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment.3:32–5:49 · Guest disagreement 0/10 Data Scale, User Content, and Behavioral Algorithms Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation.5:49–9:53 · Guest disagreement 0/10 Graph Mathematics and Structure of Family Trees Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue.9:53–12:47 · Guest disagreement 0/10 AncestryDNA Product Science and Network Effects Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment.12:47–20:52 · Guest disagreement 0/10 Scaling DNA Processing Pipeline with HBase and Hadoop Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue.20:52–27:11 · Guest disagreement 0/10 Organizational Lessons Learned and Audience Q&A Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs.0:12–3:32 · Matt pushing back 0/10 Ancestry.com Company Overview and Mission Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment.3:32–5:49 · Matt pushing back 0/10 Data Scale, User Content, and Behavioral Algorithms Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation.5:49–9:53 · Matt pushing back 0/10 Graph Mathematics and Structure of Family Trees Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue.9:53–12:47 · Matt pushing back 0/10 AncestryDNA Product Science and Network Effects Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment.12:47–20:52 · Matt pushing back 0/10 Scaling DNA Processing Pipeline with HBase and Hadoop Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue.20:52–27:11 · Matt pushing back 0/10 Organizational Lessons Learned and Audience Q&A Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 6.2% · guest 93.8%0:00 · Matt 6.2% · guest 93.8%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 4.9% · guest 95.1%21:00 · Matt 4.9% · guest 95.1%24:00 · Matt 0.3% · guest 99.7%24:00 · Matt 0.3% · guest 99.7%27:00 · Matt 28.8% · guest 71.2%27:00 · Matt 28.8% · guest 71.2%
Sharpest disagreement ▶ 26:25 Rejecting weekly meetings as ineffective

Scott dismisses conventional management approaches, explicitly stating that holding weekly alignment meetings is a waste of time compared to physical co-location.

Hardest push from Matt ▶ 26:00 Host controlling Q&A time limit

Host Matt Turck steps in to control the event schedule, firmly capping audience Q&A by declaring 'One last one'.

Biggest teaching moment ▶ 7:20 Predictive vs. actionable machine learning models

Scott educates the audience on machine learning pitfalls, illustrating how a highly predictive 1,000-feature churn model failed to yield actionable business decisions.

Matt holds his own ▶ 0:01 Host introduction and presentation handoff

Host Matt Turck opens the event and sets up Scott Sorensen's talk on Ancestry.com's big data stack.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Ancestry.com Company Overview and Mission 0200 Scott delivers an uninterrupted talk introducing Ancestry.com's user scale, revenue, and a personal family story illustrating how historical records generate leaves and hints. The host does not speak during this monologue segment.
Data Scale, User Content, and Behavioral Algorithms 0300 Scott outlines Ancestry's 12 billion records and 10 petabytes of data, explaining how user attachments provide training data for search algorithms similar to Amazon recommendation engines. The host remains silent throughout the presentation.
Graph Mathematics and Structure of Family Trees 0400 Scott details graph theory applications in family trees and Ancestry's adoption of Hadoop, MapReduce, and R. He shares insights into why a 1,000-feature churn model was predictive but non-actionable. The host provides no input during the monologue.
AncestryDNA Product Science and Network Effects 0500 Scott explains the genetics science behind AncestryDNA, describing how comparing 700,000 SNPs creates an N-squared computational challenge offset by quadratic cousin network effects. The host is not active in this segment.
Scaling DNA Processing Pipeline with HBase and Hadoop 0600 Scott breaks down how Ancestry parallelized academic genomic algorithms onto HBase and Hadoop using Battlestar Galactica character matrices. The host is absent from this technical deep-dive monologue.
Organizational Lessons Learned and Audience Q&A 0500 Scott outlines operational learnings regarding cross-functional team structures and responds to audience questions about false positives and team alignment. Host Matt Turck acts solely as a session moderator managing question handoffs.

Statements from this episode (13)

Assertion Supported
Ancestry has 2.7 million paying subscribers and leads online family history
“So, ah, at Ancestry.com, we're the largest online resource for family history, and, ah, we've got 2.7 million paying subscribers.”
Scott Sorensen Jan 16, 2014 ▶ 0:42
Assertion Partly supported
Ancestry generated $480 million in 2012 revenue across eight countries
“We had, we did, ah, four hundred and eighty million dollars in revenue last year, and we've got offices in eight countries.”
Scott Sorensen Jan 16, 2014 ▶ 0:52
Assertion Not checkable as stated
Ancestry holds over 12 billion records and 10 petabytes of data
“We have over twelve billion records and images, 10 petabytes of data, and we've worked with government agencies to collect this institutional content.”
Scott Sorensen Jan 16, 2014 ▶ 3:37
Assertion Not checkable as stated
Ancestry users built over 50 million family trees containing 5 billion profiles
“We have our customers who have built over fifty million family trees with more than five billion people in these family trees and many, many of these trees have stories and photos uploaded to these trees.”
Scott Sorensen Jan 16, 2014 ▶ 4:02
Disclosure
Ancestry uses user family tree attachments to train information retrieval algorithms
“Every time that a user associates some content with somebody in their family tree, they're making a judgment that we can use for training data in order to train our information retrieval algorithms.”
Scott Sorensen Jan 16, 2014 ▶ 4:37
Insight
Adding social features shifted Ancestry family trees from deep to broad
“Initially we didn't have a lot of social features on our site, and our trees were very deep, and they were, ah, very balanced. But as we added social features to our site, it made it easier for our customers to collaborate with cousins. Ah, then they started b…”
Scott Sorensen Jan 16, 2014 ▶ 6:19
Insight
Predictive machine learning models are useless if they cannot inform interventions
“One of the key learnings was that just because something's predictive doesn't mean that it's actionable.”
Scott Sorensen Jan 16, 2014 ▶ 8:25
Assertion Not checkable as stated
Entity extraction elevated Ancestry's worst-performing data collections into its top 10
“It took what was some of our worst performing collections out of our 30,000, ah, collections, and put some of those collections into our top 10 best, ah, best collections.”
Scott Sorensen Jan 16, 2014 ▶ 9:35
Assertion Not checkable as stated
Ancestry's DNA testing service achieves 90% accuracy when matching fourth cousins
“For fourth cousins, we're 90% accurate”
Scott Sorensen Jan 16, 2014 ▶ 11:16
Assertion Not checkable as stated
Ancestry's database finds over 40 fourth-cousin matches per user on average
“With over 200,000 samples, we can find on average over 40 fourth cousin only matches and more third and second cousins for every other person on our, in our database.”
Scott Sorensen Jan 16, 2014 ▶ 12:57
Insight
Ancestry embeds data scientists into software teams to ensure production readiness
“What we've found, at least at Ancestry, is that we have to embed our data science Scientists into the actual engineering teams and make them talk every day. It improves the feature engineering, improves the domain knowledge, improves the likelihood that engine…”
Scott Sorensen Jan 16, 2014 ▶ 20:23
Insight
Technical architecture mirrors organizational structure, requiring coordinated team and system decoupling
“It's uncanny how technical architecture follows an organization, and so if you want to decouple systems, you have to decouple the organization. If you want tightly, ah, integrated, if you want something tightly integrated, you have to tightly integrate the org…”
Scott Sorensen Jan 16, 2014 ▶ 21:52
Insight
Granting academics access to large datasets creates a data science recruiting pipeline
“If you have a large data set and can make it available for academic research, then you could create yourself a pipeline of potentially good data science candidates because they know how to deal with large amounts of data, but they also know your data pretty we…”
Scott Sorensen Jan 16, 2014 ▶ 22:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.