Dec 5, 2013 · 56m · mad
Panel Discussion // Data Driven #16 // May 2013
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Hosted by Matt Turck at Data Driven NYC, this panel discussion featuring Cathy O'Neil, Max Shron, Chris Wiggins, Claudia Perlich, and Drew Conway explores the evolving field of data science, examining practitioner career paths, educational foundations, real-world applications across industry and government, and practical technical challenges.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 8% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Max Shron interrupting and explicitly stating 'Strongly disagree' to Cathy O'Neil's assertion that feature selection is always paramount represents the transcript's most direct panelist clash.
Hardest push from Matt ▶ 28:42 Matt Turck reframes audience question around founder hiring choicesMatt Turck interrupts the general discussion to pivot the query specifically toward whether startup founders should prioritize vertical domain experience over horizontal technical skills.
Biggest teaching moment ▶ 18:30 Chris Wiggins corrects host on data science historyChris Wiggins educates the host by challenging the notion that data science is a brand-new term, citing Bill Cleveland's 2001 proposal and roots extending back to Tukey and Deming in 1940 and 1962.
Matt holds his own ▶ 34:55 Matt Turck focuses the discussion on hiring practicalities for entrepreneursMatt Turck demonstrates domain awareness of his startup audience by clearly framing the core hiring trade-offs facing startup CEOs in data science.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility | 2 | 3 | 4 | 1 | Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time. | |
| Panelist Background: Max Shron on Applied Projects and Portfolio Building | 2 | 2 | 2 | 1 | Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback. | |
| Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo | 2 | 5 | 4 | 1 | Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing. | |
| Panelist Background: Claudia Perlich on Academic Research and Digital Advertising | 0 | 1 | 1 | 0 | Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics. | |
| Audience Q&A: Formulating Business Needs versus Loose Data Questions | 1 | 2 | 1 | 1 | Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational. | |
| Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities | 1 | 3 | 4 | 0 | Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator. | |
| Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact | 3 | 3 | 5 | 2 | After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration. | |
| Audience Q&A: Data Science in Academia and Degree Programs | 0 | 3 | 1 | 0 | An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment. | |
| Audience Q&A: Civic Data Science and Municipal Government Leadership | 0 | 2 | 0 | 0 | Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation. | |
| Audience Q&A: Information Geometry and Social Science Disciplines | 2 | 4 | 2 | 1 | An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker. | |
| Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work | 0 | 3 | 3 | 0 | An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely. | |
| Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity | 1 | 4 | 6 | 0 | An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability. |