Jan 16, 2015 · 27m · mad
Chris Wiggins, NY Times // Data Science at The New York Times (Hosted by FirstMark Capital)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this presentation at Data Driven NYC, Chris Wiggins, Head of Data Science at The New York Times, discusses how computational statistics, predictive modeling, and data engineering are applied to support digital journalism, reader engagement, and organizational transformation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Chris explicitly rejects Matt's premise that data science is simply about letting data tell the story, detailing how human feature engineering inherently introduces hypotheses.
Hardest push from Matt ▶ 11:04 Host probing on hypothesis vs data-driven modelingMatt directly challenges Chris to clarify whether their predictive modeling approach lets data tell the story rather than verifying pre-existing hypotheses.
Biggest teaching moment ▶ 2:40 Historical origins of applied statistical algorithmsChris educates the room on how foundational machine learning tools like CART and Random Forests arose from mathematical statisticians doing hands-on consulting with messy real-world data.
Matt holds his own ▶ 1:40 Framing horizontal technical skills versus domain verticalityMatt demonstrates sharp conceptual understanding by framing Chris's career transition in terms of horizontal data science capabilities translating into vertical domain fields.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Horizontal Data Science Skills vs. Vertical Domain Expertise | 4 | 7 | 1 | 1 | Matt opens with an insightful framing comparing horizontal technical skills to vertical domain expertise. Chris enthusiastically agrees and educates the audience on the history of applied statistics, highlighting how luminaries like Leo Breiman and John Tukey developed core ML algorithms like CART and Random Forests through hands-on domain consulting. | |
| Organizational Culture and the Changing Newspaper Business Model | 3 | 6 | 2 | 1 | Matt asks about the cultural readiness of New York Times staff to adopt data. Chris gently reframes the issue from purely cultural willingness to economic necessity, citing Steve Blank's startup definition and detailing how the collapse of print ad revenues forced traditional publishers to rethink reader data. | |
| Predictive Modeling and Machine Learning at The Times | 4 | 6 | 3 | 3 | Matt prompts Chris on predictive analytics and asks whether they simply let data speak without prior hypotheses. Chris politely pushes back on this simplistic premise, explaining how human bias in feature engineering and deep learning examples demonstrate that pure data-driven discovery is rarely unmediated. | |
| Newsroom Analytics and Digital Journalism Innovation | 3 | 5 | 2 | 2 | Matt asks whether digital-first models like BuzzFeed are viewed as heresy in the newsroom. Chris responds by explaining that the Times newsroom is deeply self-critical and forward-thinking, citing the internal Innovation Report leaked to BuzzFeed as evidence of their openness to digital evolution. | |
| Data Engineering Architecture and Technology Stack | 3 | 5 | 1 | 1 | Matt inquires about data engineering infrastructure and tooling. Chris provides a detailed breakdown of their stack, including Python, scikit-learn, MapReduce on AWS EMR, S3, and HDFS, while explaining how open-source and vendor tools co-exist. | |
| Q&A: Cultural Failures and Respecting Journalistic Craft | 1 | 6 | 2 | 0 | An audience member asks about high-profile media blowups like The New Republic. Chris delivers a thoughtful lesson on respecting newsroom culture and journalistic craft, arguing that data scientists must be good listeners rather than trying to unilaterally replace guild values with machine learning. | |
| Q&A: Cultivating Data Literacy Across Society | 1 | 7 | 1 | 0 | An audience member asks how society can become data-driven without universal coding skills. Chris delivers a masterclass outlining Mark Hansen's concept of multi-literacies, breaking data literacy down into functional, rhetorical, and critical capacities. |