Nov 20, 2014 · 22m · mad

Michael Rubenstein and Catherine Williams, App Nexus // Data Driven #31 // Nov 2014

Catherine Williams · 13m spoken Michael Rubenstein · 4m spoken Carter Schoenwald · 43s spoken Matt Turck · 23s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this presentation at Data Driven NYC, AppNexus executives Michael Rubenstein and Catherine Williams explore the scale, structural evolution, and real-world algorithmic applications of data science within the programmatic advertising ecosystem. They detail how AppNexus transitioned from makeshift analytics into an enterprise platform processing tens of billions of daily transactions through advanced machine learning and streaming technologies.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2% of the talking time here. How this is scored →

Matt as informed peer 0.2 Guest teaching 1.4 Guest disagreement 0.4 Matt pushing back 0.2
05100:0010:0020:004:18–7:47 · Matt as informed peer 0/10 Catherine Williams' Background, Dark Ages, and Renaissance of Data Catherine Williams delivers a solo presentation on her background in pure mathematics and her transition to AppNexus quants during a data renaissance. The host does not participate in this monologue segment.7:47–10:48 · Matt as informed peer 0/10 Case Study: Pacing Algorithm for Even Budget Delivery Catherine presents a case study explaining her first project at AppNexus, creating a pacing algorithm lever across 1500 disconnected bidder instances. The host is inactive during this presentation segment.10:48–14:21 · Matt as informed peer 0/10 Case Study: Programmatic Adult Content Detection and Word Splitting Catherine illustrates how adult content detection was improved through probabilistic domain word splitting using the Google trillion-word corpus. The host does not speak in this monologue segment.14:21–16:54 · Matt as informed peer 0/10 The Modern Era of Data Science at AppNexus Catherine outlines the modern infrastructure of data science at AppNexus, including Hadoop clusters and real-time streaming tools. The host is silent throughout this segment.16:54–22:57 · Matt as informed peer 1/10 Looking Forward: Visual Metaphor for Predictive Systems Matt Turck opens the floor for audience Q&A, where Catherine and Michael answer technical questions on auction theory and click-through rates. Catherine gently reframes audience questions about truth-telling mechanisms and long-term incentives.4:18–7:47 · Guest teaching 1/10 Catherine Williams' Background, Dark Ages, and Renaissance of Data Catherine Williams delivers a solo presentation on her background in pure mathematics and her transition to AppNexus quants during a data renaissance. The host does not participate in this monologue segment.7:47–10:48 · Guest teaching 1/10 Case Study: Pacing Algorithm for Even Budget Delivery Catherine presents a case study explaining her first project at AppNexus, creating a pacing algorithm lever across 1500 disconnected bidder instances. The host is inactive during this presentation segment.10:48–14:21 · Guest teaching 1/10 Case Study: Programmatic Adult Content Detection and Word Splitting Catherine illustrates how adult content detection was improved through probabilistic domain word splitting using the Google trillion-word corpus. The host does not speak in this monologue segment.14:21–16:54 · Guest teaching 1/10 The Modern Era of Data Science at AppNexus Catherine outlines the modern infrastructure of data science at AppNexus, including Hadoop clusters and real-time streaming tools. The host is silent throughout this segment.16:54–22:57 · Guest teaching 3/10 Looking Forward: Visual Metaphor for Predictive Systems Matt Turck opens the floor for audience Q&A, where Catherine and Michael answer technical questions on auction theory and click-through rates. Catherine gently reframes audience questions about truth-telling mechanisms and long-term incentives.4:18–7:47 · Guest disagreement 0/10 Catherine Williams' Background, Dark Ages, and Renaissance of Data Catherine Williams delivers a solo presentation on her background in pure mathematics and her transition to AppNexus quants during a data renaissance. The host does not participate in this monologue segment.7:47–10:48 · Guest disagreement 0/10 Case Study: Pacing Algorithm for Even Budget Delivery Catherine presents a case study explaining her first project at AppNexus, creating a pacing algorithm lever across 1500 disconnected bidder instances. The host is inactive during this presentation segment.10:48–14:21 · Guest disagreement 0/10 Case Study: Programmatic Adult Content Detection and Word Splitting Catherine illustrates how adult content detection was improved through probabilistic domain word splitting using the Google trillion-word corpus. The host does not speak in this monologue segment.14:21–16:54 · Guest disagreement 0/10 The Modern Era of Data Science at AppNexus Catherine outlines the modern infrastructure of data science at AppNexus, including Hadoop clusters and real-time streaming tools. The host is silent throughout this segment.16:54–22:57 · Guest disagreement 2/10 Looking Forward: Visual Metaphor for Predictive Systems Matt Turck opens the floor for audience Q&A, where Catherine and Michael answer technical questions on auction theory and click-through rates. Catherine gently reframes audience questions about truth-telling mechanisms and long-term incentives.4:18–7:47 · Matt pushing back 0/10 Catherine Williams' Background, Dark Ages, and Renaissance of Data Catherine Williams delivers a solo presentation on her background in pure mathematics and her transition to AppNexus quants during a data renaissance. The host does not participate in this monologue segment.7:47–10:48 · Matt pushing back 0/10 Case Study: Pacing Algorithm for Even Budget Delivery Catherine presents a case study explaining her first project at AppNexus, creating a pacing algorithm lever across 1500 disconnected bidder instances. The host is inactive during this presentation segment.10:48–14:21 · Matt pushing back 0/10 Case Study: Programmatic Adult Content Detection and Word Splitting Catherine illustrates how adult content detection was improved through probabilistic domain word splitting using the Google trillion-word corpus. The host does not speak in this monologue segment.14:21–16:54 · Matt pushing back 0/10 The Modern Era of Data Science at AppNexus Catherine outlines the modern infrastructure of data science at AppNexus, including Hadoop clusters and real-time streaming tools. The host is silent throughout this segment.16:54–22:57 · Matt pushing back 1/10 Looking Forward: Visual Metaphor for Predictive Systems Matt Turck opens the floor for audience Q&A, where Catherine and Michael answer technical questions on auction theory and click-through rates. Catherine gently reframes audience questions about truth-telling mechanisms and long-term incentives.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 2.5% · guest 97.5%15:00 · Matt 2.5% · guest 97.5%18:00 · Matt 4% · guest 96%18:00 · Matt 4% · guest 96%21:00 · Matt 16% · guest 84%21:00 · Matt 16% · guest 84%
Sharpest disagreement ▶ 18:36 Rejecting Truth-Telling Premise

Catherine challenges audience member Carter's assumption regarding truth-telling in repeated second-price auctions, noting she does not consider it an ideal end-state.

Hardest push from Matt ▶ 21:58 Rephrasing Audience Question

Matt Turck steps in to rephrase and clarify Ariel's audience question about click-through rate improvements after Catherine asks for clarification.

Biggest teaching moment ▶ 12:48 Domain Splitting via Google Corpus

Catherine details how her team used mathematical word-splitting probabilities based on the Google trillion-word corpus to eliminate false positive porn flags.

Matt holds his own ▶ 17:55 Host Moderating Audience Q&A

Matt Turck steps back in smoothly to transition from the presentation phase directly into audience Q&A while keeping time.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Catherine Williams' Background, Dark Ages, and Renaissance of Data 0100 Catherine Williams delivers a solo presentation on her background in pure mathematics and her transition to AppNexus quants during a data renaissance. The host does not participate in this monologue segment.
Case Study: Pacing Algorithm for Even Budget Delivery 0100 Catherine presents a case study explaining her first project at AppNexus, creating a pacing algorithm lever across 1500 disconnected bidder instances. The host is inactive during this presentation segment.
Case Study: Programmatic Adult Content Detection and Word Splitting 0100 Catherine illustrates how adult content detection was improved through probabilistic domain word splitting using the Google trillion-word corpus. The host does not speak in this monologue segment.
The Modern Era of Data Science at AppNexus 0100 Catherine outlines the modern infrastructure of data science at AppNexus, including Hadoop clusters and real-time streaming tools. The host is silent throughout this segment.
Looking Forward: Visual Metaphor for Predictive Systems 1321 Matt Turck opens the floor for audience Q&A, where Catherine and Michael answer technical questions on auction theory and click-through rates. Catherine gently reframes audience questions about truth-telling mechanisms and long-term incentives.

Statements from this episode (9)

Assertion Not checkable as stated
AppNexus processes over 100 terabytes of data daily
“To a company today that's processing over a hundred terabytes of data daily and has got massive internet scale.”
Michael Rubenstein Nov 20, 2014 ▶ 0:27
Assertion Supported
AppNexus acquired WPP's tech assets and powers its digital media buying
“We acquired their technology assets, they made an investment in our company, and today our technology powers the digital buying of the largest ad agency on earth.”
Michael Rubenstein Nov 20, 2014 ▶ 2:38
Disclosure
AppNexus has raised around $250 million in venture capital
“We haven't quite raised three hundred million, but we've raised almost that much, around two hundred fifty million dollars, so a lot of venture capital”
Michael Rubenstein Nov 20, 2014 ▶ 3:43
Assertion Not checkable as stated
AppNexus operated about 1,500 global bidding server instances
“These bidding instances, there's actually about 1500 of them spread in data centers around the world”
Catherine Williams Nov 20, 2014 ▶ 8:23
Disclosure
AppNexus previously relied on human auditors to detect adult content
“We had a team of human auditors who would literally just go look at websites and determine whether or not they were porn.”
Catherine Williams Nov 20, 2014 ▶ 11:02
Disclosure
AppNexus leveraged Google's trillion-word corpus for domain word-splitting algorithms
“The data here is the Google trillion word corpus, right, where they scrape the entire internet, look for the frequencies of all the different strings, including misspellings, right, and that basically gives you a probability of any given string being, like, a …”
Catherine Williams Nov 20, 2014 ▶ 13:52
Assertion Not checkable as stated
AppNexus processes 30 billion daily impressions on a 16-node Hadoop cluster
“We have a 16 node Hadoop cluster currently, and I have on my proposed budget for 2015, a 200 node Hadoop cluster so that we can really get our hands on all that raw data of the thirty billion impressions we're transacting daily.”
Catherine Williams Nov 20, 2014 ▶ 14:48
Assertion Not checkable as stated
AppNexus has no first-party data
“We don't have any first party data at all.”
Catherine Williams Nov 20, 2014 ▶ 17:13
Insight
Product teams lacking data science, engineering, or product management are unstable
“The data science product engineering trifecta was just a very, very powerful combination. And when you were lacking one, it didn't, it was unstable.”
Catherine Williams Nov 20, 2014 ▶ 20:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.