Jan 16, 2015 · 27m · mad

Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital)

Chris Wiggins · 21m spoken Matt Turck · 3m spoken Elliot Noma · 43s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this presentation at Data Driven NYC, Chris Wiggins, Head of Data Science at The New York Times, discusses how computational statistics, predictive modeling, and data engineering are applied to support digital journalism, reader engagement, and organizational transformation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.4% of the talking time here. How this is scored →

Matt as informed peer 2.7 Guest teaching 6.0 Guest disagreement 1.7 Matt pushing back 1.1
05100:0010:0020:001:40–5:42 · Matt as informed peer 4/10 Horizontal Data Science Skills vs. Vertical Domain Expertise Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting.5:42–8:01 · Matt as informed peer 3/10 Organizational Culture and the Changing Newspaper Business Model Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data.8:01–12:56 · Matt as informed peer 4/10 Predictive Modeling and Machine Learning at The Times Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated.12:56–17:36 · Matt as informed peer 3/10 Newsroom Analytics and Digital Journalism Innovation Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution.17:36–20:00 · Matt as informed peer 3/10 Data Engineering Architecture and Technology Stack Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist.20:00–23:54 · Matt as informed peer 1/10 Q&A: Cultural Failures and Respecting Journalistic Craft An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning.23:54–27:50 · Matt as informed peer 1/10 Q&A: Cultivating Data Literacy Across Society An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities.1:40–5:42 · Guest teaching 7/10 Horizontal Data Science Skills vs. Vertical Domain Expertise Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting.5:42–8:01 · Guest teaching 6/10 Organizational Culture and the Changing Newspaper Business Model Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data.8:01–12:56 · Guest teaching 6/10 Predictive Modeling and Machine Learning at The Times Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated.12:56–17:36 · Guest teaching 5/10 Newsroom Analytics and Digital Journalism Innovation Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution.17:36–20:00 · Guest teaching 5/10 Data Engineering Architecture and Technology Stack Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist.20:00–23:54 · Guest teaching 6/10 Q&A: Cultural Failures and Respecting Journalistic Craft An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning.23:54–27:50 · Guest teaching 7/10 Q&A: Cultivating Data Literacy Across Society An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities.1:40–5:42 · Guest disagreement 1/10 Horizontal Data Science Skills vs. Vertical Domain Expertise Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting.5:42–8:01 · Guest disagreement 2/10 Organizational Culture and the Changing Newspaper Business Model Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data.8:01–12:56 · Guest disagreement 3/10 Predictive Modeling and Machine Learning at The Times Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated.12:56–17:36 · Guest disagreement 2/10 Newsroom Analytics and Digital Journalism Innovation Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution.17:36–20:00 · Guest disagreement 1/10 Data Engineering Architecture and Technology Stack Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist.20:00–23:54 · Guest disagreement 2/10 Q&A: Cultural Failures and Respecting Journalistic Craft An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning.23:54–27:50 · Guest disagreement 1/10 Q&A: Cultivating Data Literacy Across Society An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities.1:40–5:42 · Matt pushing back 1/10 Horizontal Data Science Skills vs. Vertical Domain Expertise Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting.5:42–8:01 · Matt pushing back 1/10 Organizational Culture and the Changing Newspaper Business Model Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data.8:01–12:56 · Matt pushing back 3/10 Predictive Modeling and Machine Learning at The Times Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated.12:56–17:36 · Matt pushing back 2/10 Newsroom Analytics and Digital Journalism Innovation Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution.17:36–20:00 · Matt pushing back 1/10 Data Engineering Architecture and Technology Stack Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist.20:00–23:54 · Matt pushing back 0/10 Q&A: Cultural Failures and Respecting Journalistic Craft An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning.23:54–27:50 · Matt pushing back 0/10 Q&A: Cultivating Data Literacy Across Society An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 33.5% · guest 66.5%0:00 · Matt 33.5% · guest 66.5%3:00 · Matt 9.4% · guest 90.6%3:00 · Matt 9.4% · guest 90.6%6:00 · Matt 18.1% · guest 81.9%6:00 · Matt 18.1% · guest 81.9%9:00 · Matt 10.1% · guest 89.9%9:00 · Matt 10.1% · guest 89.9%12:00 · Matt 23.7% · guest 76.3%12:00 · Matt 23.7% · guest 76.3%15:00 · Matt 14.3% · guest 85.7%15:00 · Matt 14.3% · guest 85.7%18:00 · Matt 3.9% · guest 96.1%18:00 · Matt 3.9% · guest 96.1%21:00 · Matt 0.9% · guest 99.1%21:00 · Matt 0.9% · guest 99.1%24:00 · Matt 0% · guest 100%24:00 · Matt 0% · guest 100%27:00 · Matt 3.6% · guest 96.4%27:00 · Matt 3.6% · guest 96.4%
Sharpest disagreement ▶ 11:13 Challenging the pure data-driven premise

Chris explicitly rejects Matt's premise that data science is simply about letting data tell the story, detailing how human feature engineering inherently introduces hypotheses.

Hardest push from Matt ▶ 11:04 Host probing on hypothesis vs data-driven modeling

Matt directly challenges Chris to clarify whether their predictive modeling approach lets data tell the story rather than verifying pre-existing hypotheses.

Biggest teaching moment ▶ 2:40 Historical origins of applied statistical algorithms

Chris educates the room on how foundational machine learning tools like CART and Random Forests arose from mathematical statisticians doing hands-on consulting with messy real-world data.

Matt holds his own ▶ 1:40 Framing horizontal technical skills versus domain verticality

Matt demonstrates sharp conceptual understanding by framing Chris's career transition in terms of horizontal data science capabilities translating into vertical domain fields.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Horizontal Data Science Skills vs. Vertical Domain Expertise 4711 Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting.
Organizational Culture and the Changing Newspaper Business Model 3621 Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data.
Predictive Modeling and Machine Learning at The Times 4633 Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated.
Newsroom Analytics and Digital Journalism Innovation 3522 Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution.
Data Engineering Architecture and Technology Stack 3511 Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist.
Q&A: Cultural Failures and Respecting Journalistic Craft 1620 An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning.
Q&A: Cultivating Data Literacy Across Society 1710 An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities.

Statements from this episode (8)

Insight
Data science differs from ML through interdisciplinary domain collaboration
“The thing that makes data science different from machine learning is not just getting epsilon better predictive accuracy on learning, you know, cat's faces from pictures. It's this thing where you interact with somebody from a different discipline, and then so…”
Chris Wiggins Jan 16, 2015 ▶ 4:07
Insight
Every publishing company is now a startup searching for a business model
“I like to use Steve Blank's definition of a startup, that a startup is a temporary organization in search of a scalable and repeatable business model. And in that sense, every publisher is now a startup, because the business model of publishing just completely…”
Chris Wiggins Jan 16, 2015 ▶ 6:43
Assertion Partly supported
US print advertising spend fell about 50% from 2008 to 2012
“Print advertising spend in the United States lost about 50% of its value in like four years, 2008 to 2012.”
Chris Wiggins Jan 16, 2015 ▶ 7:06
Insight
Tech companies and digital media now operate as church, state, and engineering
“I like to think about the New York Times or any technology company now as church, state, and engineering”
Chris Wiggins Jan 16, 2015 ▶ 9:00
Insight
Wiggins: Supervised models beat clustering because errors are clear
“Working on, on supervised learning or predictive models to be reassuring because I know if I'm wrong. Whereas, you know, models where I'm clustering, I sort of never know at the end of the day, should I have clustered things a different way?”
Chris Wiggins Jan 16, 2015 ▶ 10:38
Disclosure
NYT data science prefers buying or using open source over building custom
“In terms of build or buy, if we, if there's something out there we can buy, we'll buy it. And if there's an open source alternative, then we'll definitely use that because most of the group know open source.”
Chris Wiggins Jan 16, 2015 ▶ 19:42
Opinion
The New Republic and First Look Media failures were total people failures
“Those are both, like, total people failures, right”
Chris Wiggins Jan 16, 2015 ▶ 20:57
Insight
Data literacy requires critical, rhetorical, and functional skills equally
“Critical literacy, rhetorical literacy, and functional literacy, I think, are all equally important parts of having a data literate society.”
Chris Wiggins Jan 16, 2015 ▶ 27:34
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.