Dec 5, 2013 · 23m · mad

John Foreman, Mailchimp // Data Driven NYC 19 // October 2013

John Foreman · 20m spoken Matt Turck · 56s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC, MailChimp's Chief Data Scientist John Foreman presents a practical framework for applying data science in business, illustrating how data work should be structured across internal operational tools, external product capabilities, and user-centered design. He emphasizes prioritizing actionable user utility and simple design integration over overly complex analytics and vanity visualizations.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.4% of the talking time here. How this is scored →

Matt as informed peer 0.1 Guest teaching 3.9 Guest disagreement 1.4 Matt pushing back 0.0
05100:0010:0020:001:05–3:20 · Matt as informed peer 0/10 John Foreman's Background and MailChimp Scale Overview John Foreman presents MailChimp's growth metrics, company history, and design-led engineering culture in a solo talk. Matt Turck does not speak during this segment, keeping host-side scores at zero.3:20–5:47 · Matt as informed peer 0/10 Actionable Data Science vs. Unhelpful Data Visualizations Foreman critiques vanity data visualizations like city hipster heatmaps and complex social network graphs, arguing they fail to improve user experience. Host participation is absent during this presentation segment.5:47–9:04 · Matt as informed peer 0/10 Internal Insights: Support Scheduling and the Kobayashi Maru Strategy Foreman explains applying linear programming to schedule support chat staff and advocates adopting a Kobayashi Maru mindset when data models hit operational limits. The segment is a monologue, so host scores remain zero.9:04–11:31 · Matt as informed peer 0/10 External Insights: Analyzing Email Engagement Patterns Foreman shares data on how Gmail tabbed inbox updates and federal government shutdowns altered email open rates across MailChimp. No host intervention occurs during the presentation.11:31–15:27 · Matt as informed peer 0/10 Internal Capabilities: Machine Learning for Compliance and Anti-Spam Foreman details MailChimp's automated compliance architecture, highlighting why they chose Postgres, Redis, and Random Forest models in R to eliminate spam pre-send. The host is inactive throughout.15:27–18:16 · Matt as informed peer 0/10 External Capabilities: Send Time Optimization and Demographic Predictions Foreman explains how cross-subscription graphs and email address tokens enable send time optimization and demographic predictions. Host scores are zero due to the solo presentation format.18:16–23:09 · Matt as informed peer 0/10 User-Centric Data Product Design: Discovered Segments Case Study Foreman demonstrates how Discovered Segments replaces complex visualizations with simple UI buttons to achieve significant open-rate and revenue uplifts. As a monologue segment, host scores remain zero.23:09–23:44 · Matt as informed peer 1/10 Book Announcement and Q&A Transition Foreman plugs his upcoming spreadsheet data science book, and Matt Turck re-enters briefly to compliment his author slide photo and thank him. Host expertise is 1 for the brief concluding exchange.1:05–3:20 · Guest teaching 3/10 John Foreman's Background and MailChimp Scale Overview John Foreman presents MailChimp's growth metrics, company history, and design-led engineering culture in a solo talk. Matt Turck does not speak during this segment, keeping host-side scores at zero.3:20–5:47 · Guest teaching 4/10 Actionable Data Science vs. Unhelpful Data Visualizations Foreman critiques vanity data visualizations like city hipster heatmaps and complex social network graphs, arguing they fail to improve user experience. Host participation is absent during this presentation segment.5:47–9:04 · Guest teaching 5/10 Internal Insights: Support Scheduling and the Kobayashi Maru Strategy Foreman explains applying linear programming to schedule support chat staff and advocates adopting a Kobayashi Maru mindset when data models hit operational limits. The segment is a monologue, so host scores remain zero.9:04–11:31 · Guest teaching 4/10 External Insights: Analyzing Email Engagement Patterns Foreman shares data on how Gmail tabbed inbox updates and federal government shutdowns altered email open rates across MailChimp. No host intervention occurs during the presentation.11:31–15:27 · Guest teaching 5/10 Internal Capabilities: Machine Learning for Compliance and Anti-Spam Foreman details MailChimp's automated compliance architecture, highlighting why they chose Postgres, Redis, and Random Forest models in R to eliminate spam pre-send. The host is inactive throughout.15:27–18:16 · Guest teaching 4/10 External Capabilities: Send Time Optimization and Demographic Predictions Foreman explains how cross-subscription graphs and email address tokens enable send time optimization and demographic predictions. Host scores are zero due to the solo presentation format.18:16–23:09 · Guest teaching 5/10 User-Centric Data Product Design: Discovered Segments Case Study Foreman demonstrates how Discovered Segments replaces complex visualizations with simple UI buttons to achieve significant open-rate and revenue uplifts. As a monologue segment, host scores remain zero.23:09–23:44 · Guest teaching 1/10 Book Announcement and Q&A Transition Foreman plugs his upcoming spreadsheet data science book, and Matt Turck re-enters briefly to compliment his author slide photo and thank him. Host expertise is 1 for the brief concluding exchange.1:05–3:20 · Guest disagreement 1/10 John Foreman's Background and MailChimp Scale Overview John Foreman presents MailChimp's growth metrics, company history, and design-led engineering culture in a solo talk. Matt Turck does not speak during this segment, keeping host-side scores at zero.3:20–5:47 · Guest disagreement 2/10 Actionable Data Science vs. Unhelpful Data Visualizations Foreman critiques vanity data visualizations like city hipster heatmaps and complex social network graphs, arguing they fail to improve user experience. Host participation is absent during this presentation segment.5:47–9:04 · Guest disagreement 2/10 Internal Insights: Support Scheduling and the Kobayashi Maru Strategy Foreman explains applying linear programming to schedule support chat staff and advocates adopting a Kobayashi Maru mindset when data models hit operational limits. The segment is a monologue, so host scores remain zero.9:04–11:31 · Guest disagreement 1/10 External Insights: Analyzing Email Engagement Patterns Foreman shares data on how Gmail tabbed inbox updates and federal government shutdowns altered email open rates across MailChimp. No host intervention occurs during the presentation.11:31–15:27 · Guest disagreement 2/10 Internal Capabilities: Machine Learning for Compliance and Anti-Spam Foreman details MailChimp's automated compliance architecture, highlighting why they chose Postgres, Redis, and Random Forest models in R to eliminate spam pre-send. The host is inactive throughout.15:27–18:16 · Guest disagreement 1/10 External Capabilities: Send Time Optimization and Demographic Predictions Foreman explains how cross-subscription graphs and email address tokens enable send time optimization and demographic predictions. Host scores are zero due to the solo presentation format.18:16–23:09 · Guest disagreement 2/10 User-Centric Data Product Design: Discovered Segments Case Study Foreman demonstrates how Discovered Segments replaces complex visualizations with simple UI buttons to achieve significant open-rate and revenue uplifts. As a monologue segment, host scores remain zero.23:09–23:44 · Guest disagreement 0/10 Book Announcement and Q&A Transition Foreman plugs his upcoming spreadsheet data science book, and Matt Turck re-enters briefly to compliment his author slide photo and thank him. Host expertise is 1 for the brief concluding exchange.1:05–3:20 · Matt pushing back 0/10 John Foreman's Background and MailChimp Scale Overview John Foreman presents MailChimp's growth metrics, company history, and design-led engineering culture in a solo talk. Matt Turck does not speak during this segment, keeping host-side scores at zero.3:20–5:47 · Matt pushing back 0/10 Actionable Data Science vs. Unhelpful Data Visualizations Foreman critiques vanity data visualizations like city hipster heatmaps and complex social network graphs, arguing they fail to improve user experience. Host participation is absent during this presentation segment.5:47–9:04 · Matt pushing back 0/10 Internal Insights: Support Scheduling and the Kobayashi Maru Strategy Foreman explains applying linear programming to schedule support chat staff and advocates adopting a Kobayashi Maru mindset when data models hit operational limits. The segment is a monologue, so host scores remain zero.9:04–11:31 · Matt pushing back 0/10 External Insights: Analyzing Email Engagement Patterns Foreman shares data on how Gmail tabbed inbox updates and federal government shutdowns altered email open rates across MailChimp. No host intervention occurs during the presentation.11:31–15:27 · Matt pushing back 0/10 Internal Capabilities: Machine Learning for Compliance and Anti-Spam Foreman details MailChimp's automated compliance architecture, highlighting why they chose Postgres, Redis, and Random Forest models in R to eliminate spam pre-send. The host is inactive throughout.15:27–18:16 · Matt pushing back 0/10 External Capabilities: Send Time Optimization and Demographic Predictions Foreman explains how cross-subscription graphs and email address tokens enable send time optimization and demographic predictions. Host scores are zero due to the solo presentation format.18:16–23:09 · Matt pushing back 0/10 User-Centric Data Product Design: Discovered Segments Case Study Foreman demonstrates how Discovered Segments replaces complex visualizations with simple UI buttons to achieve significant open-rate and revenue uplifts. As a monologue segment, host scores remain zero.23:09–23:44 · Matt pushing back 0/10 Book Announcement and Q&A Transition Foreman plugs his upcoming spreadsheet data science book, and Matt Turck re-enters briefly to compliment his author slide photo and thank him. Host expertise is 1 for the brief concluding exchange.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 33.6% · guest 66.4%0:00 · Matt 33.6% · guest 66.4%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 2% · guest 98%21:00 · Matt 2% · guest 98%
Sharpest disagreement ▶ 3:45 Dismissal of vanity data visualizations

Foreman openly mocks features on other platforms that visualize social graphs or search keywords, calling them unhelpful attempts to flex data muscle rather than solve real user problems.

Hardest push from Matt ▶ 23:37 Host wrap-up acknowledgment

Because this episode is a solo presentation, the host offers zero pushback or counter-framing, re-entering only at the end to compliment the speaker's slide photo and transition to Q&A.

Biggest teaching moment ▶ 6:55 Explaining linear programming optimization for support scheduling

Foreman educates the audience on defining support scheduling decision spaces as polytopes and formulating linear program files to automate complex shift planning.

Matt holds his own ▶ 23:37 Host brief session conclusion

In a talk format where the host stays off-mic, the host displays minimal active expertise, speaking briefly at the end to thank the speaker and transition to the next segment.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
John Foreman's Background and MailChimp Scale Overview 0310 John Foreman presents MailChimp's growth metrics, company history, and design-led engineering culture in a solo talk. Matt Turck does not speak during this segment, keeping host-side scores at zero.
Actionable Data Science vs. Unhelpful Data Visualizations 0420 Foreman critiques vanity data visualizations like city hipster heatmaps and complex social network graphs, arguing they fail to improve user experience. Host participation is absent during this presentation segment.
Internal Insights: Support Scheduling and the Kobayashi Maru Strategy 0520 Foreman explains applying linear programming to schedule support chat staff and advocates adopting a Kobayashi Maru mindset when data models hit operational limits. The segment is a monologue, so host scores remain zero.
External Insights: Analyzing Email Engagement Patterns 0410 Foreman shares data on how Gmail tabbed inbox updates and federal government shutdowns altered email open rates across MailChimp. No host intervention occurs during the presentation.
Internal Capabilities: Machine Learning for Compliance and Anti-Spam 0520 Foreman details MailChimp's automated compliance architecture, highlighting why they chose Postgres, Redis, and Random Forest models in R to eliminate spam pre-send. The host is inactive throughout.
External Capabilities: Send Time Optimization and Demographic Predictions 0410 Foreman explains how cross-subscription graphs and email address tokens enable send time optimization and demographic predictions. Host scores are zero due to the solo presentation format.
User-Centric Data Product Design: Discovered Segments Case Study 0520 Foreman demonstrates how Discovered Segments replaces complex visualizations with simple UI buttons to achieve significant open-rate and revenue uplifts. As a monologue segment, host scores remain zero.
Book Announcement and Q&A Transition 1100 Foreman plugs his upcoming spreadsheet data science book, and Matt Turck re-enters briefly to compliment his author slide photo and thank him. Host expertise is 1 for the brief concluding exchange.

Statements from this episode (13)

Assertion Supported
Mailchimp sends 6-7 billion monthly emails across 4 million users
“The last time I checked, four million users, somewhere between six and seven billion emails sent every month.”
John Foreman Dec 5, 2013 ▶ 1:50
Prediction Didn’t hold up
Mailchimp expects to hit 10 billion monthly emails by late 2013
“That'll probably be ten billion by the end of the year.”
John Foreman Dec 5, 2013 ▶ 1:56
Opinion
Most tech data features are built to impress investors, not help users
“When you see examples from other companies out in the wild, generally they're to improve, they're sort of to impress investors or the media.”
John Foreman Dec 5, 2013 ▶ 3:55
Disclosure
Mailchimp's data science team spends 80 percent of time building tools
“And right now we spend about 20% of our time doing insight, which is just one-off reporting or one-off sort of consulting engagements, and we spend about 80% of our time building tools or capabilities”
John Foreman Dec 5, 2013 ▶ 5:20
Assertion Not checkable as stated
Mailchimp had 200 total employees in 2013, with 100 in support
“We have 200 employees. About a hundred of those are support personnel.”
John Foreman Dec 5, 2013 ▶ 5:51
Insight
Data scientists should consider changing the business itself when models fail
“If a model's not working out, you should always ask yourself, can we just change the business? You know, as opposed to somehow optimally solving this problem, can we just stop doing this thing I'm trying to solve?”
John Foreman Dec 5, 2013 ▶ 8:03
Assertion Not checkable as stated
Gmail's tabbed inbox launch dropped Mailchimp open rates by 10 percent
“This is three weeks before and after tabs were introduced, and we can see about a raw one percent difference, so maybe about a 10% decrease in engagement. We've got three weeks. Blue is weekday. Yellow is, ah, weekend, and these are open rates. So you can see …”
John Foreman Dec 5, 2013 ▶ 9:53
Assertion Not checkable as stated
EPA email engagement fell to 10 percent during the government shutdown
“Down there at the bottom, EPA was, like, garbage, right? 10% of the engagement after the shutdown as before”
John Foreman Dec 5, 2013 ▶ 10:48
Assertion Not checkable as stated
One sixth of Mailchimp compliance tickets were solely for user vetting
“A sixth of our compliance tickets had nothing to do with being bad. We just wanted to vet you.”
John Foreman Dec 5, 2013 ▶ 13:10
Disclosure
Mailchimp used Redis to maintain dossiers on billions of email addresses
“For our fast storage, we used Redis. We've got one dossier per email address of all the email addresses we've ever seen, so billions of unique addresses and a bunch of events about everything they've engaged with in the past.”
John Foreman Dec 5, 2013 ▶ 15:06
Disclosure
Mailchimp infers recipient age and gender via cross-network subscription graphs
“We can actually determine a lot of demographic data through this, ah, age, gender, things like that.”
John Foreman Dec 5, 2013 ▶ 18:11
Assertion Not checkable as stated
Mailchimp's Discovered Segments boosted a developer newsletter engagement by 50 percent
“We took a group of devs that we knew about, threw it into discovered segments, grabbed some more, and then sent it out and actually got 50% better engagement on that newsletter than typical.”
John Foreman Dec 5, 2013 ▶ 21:08
Insight
Data science deserves no special treatment over simple UI design changes
“Data science products should receive no special treatment. If a designer Can do a better job solving a problem by changing the color of a form than I can do with some AI model or something, then that design product should win.”
John Foreman Dec 5, 2013 ▶ 22:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.