Dec 5, 2013 · 56m · mad

Panel Discussion // Data Driven #16 // May 2013

Chris Wiggins · 11m spoken Max Shron · 10m spoken Cathy O'Neil · 7m spoken Drew Conway · 6m spoken Claudia Perlich · 6m spoken Matt Turck · 3m spoken AB Mendez · 51s spoken Carlos Medina · 48s spoken Tom Olds · 43s spoken Jeroen Janssens · 38s spoken Dawn Barber · 32s spoken Victor Olex · 17s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Hosted by Matt Turck at Data Driven NYC, this panel discussion featuring Cathy O'Neil, Max Shron, Chris Wiggins, Claudia Perlich, and Drew Conway explores the evolving field of data science, examining practitioner career paths, educational foundations, real-world applications across industry and government, and practical technical challenges.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 8% of the talking time here. How this is scored →

Matt as informed peer 1.2 Guest teaching 2.9 Guest disagreement 2.8 Matt pushing back 0.6
05100:0015:0030:0045:001:09–7:11 · Matt as informed peer 2/10 Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time.7:11–13:11 · Matt as informed peer 2/10 Panelist Background: Max Shron on Applied Projects and Portfolio Building Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback.13:11–20:14 · Matt as informed peer 2/10 Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing.20:14–23:29 · Matt as informed peer 0/10 Panelist Background: Claudia Perlich on Academic Research and Digital Advertising Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics.23:29–25:31 · Matt as informed peer 1/10 Audience Q&A: Formulating Business Needs versus Loose Data Questions Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational.25:31–28:29 · Matt as informed peer 1/10 Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator.28:29–38:31 · Matt as informed peer 3/10 Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration.38:31–42:06 · Matt as informed peer 0/10 Audience Q&A: Data Science in Academia and Degree Programs An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment.42:06–44:28 · Matt as informed peer 0/10 Audience Q&A: Civic Data Science and Municipal Government Leadership Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation.44:28–48:13 · Matt as informed peer 2/10 Audience Q&A: Information Geometry and Social Science Disciplines An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker.48:13–51:31 · Matt as informed peer 0/10 Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely.51:31–56:11 · Matt as informed peer 1/10 Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability.1:09–7:11 · Guest teaching 3/10 Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time.7:11–13:11 · Guest teaching 2/10 Panelist Background: Max Shron on Applied Projects and Portfolio Building Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback.13:11–20:14 · Guest teaching 5/10 Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing.20:14–23:29 · Guest teaching 1/10 Panelist Background: Claudia Perlich on Academic Research and Digital Advertising Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics.23:29–25:31 · Guest teaching 2/10 Audience Q&A: Formulating Business Needs versus Loose Data Questions Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational.25:31–28:29 · Guest teaching 3/10 Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator.28:29–38:31 · Guest teaching 3/10 Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration.38:31–42:06 · Guest teaching 3/10 Audience Q&A: Data Science in Academia and Degree Programs An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment.42:06–44:28 · Guest teaching 2/10 Audience Q&A: Civic Data Science and Municipal Government Leadership Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation.44:28–48:13 · Guest teaching 4/10 Audience Q&A: Information Geometry and Social Science Disciplines An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker.48:13–51:31 · Guest teaching 3/10 Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely.51:31–56:11 · Guest teaching 4/10 Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability.1:09–7:11 · Guest disagreement 4/10 Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time.7:11–13:11 · Guest disagreement 2/10 Panelist Background: Max Shron on Applied Projects and Portfolio Building Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback.13:11–20:14 · Guest disagreement 4/10 Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing.20:14–23:29 · Guest disagreement 1/10 Panelist Background: Claudia Perlich on Academic Research and Digital Advertising Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics.23:29–25:31 · Guest disagreement 1/10 Audience Q&A: Formulating Business Needs versus Loose Data Questions Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational.25:31–28:29 · Guest disagreement 4/10 Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator.28:29–38:31 · Guest disagreement 5/10 Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration.38:31–42:06 · Guest disagreement 1/10 Audience Q&A: Data Science in Academia and Degree Programs An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment.42:06–44:28 · Guest disagreement 0/10 Audience Q&A: Civic Data Science and Municipal Government Leadership Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation.44:28–48:13 · Guest disagreement 2/10 Audience Q&A: Information Geometry and Social Science Disciplines An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker.48:13–51:31 · Guest disagreement 3/10 Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely.51:31–56:11 · Guest disagreement 6/10 Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability.1:09–7:11 · Matt pushing back 1/10 Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time.7:11–13:11 · Matt pushing back 1/10 Panelist Background: Max Shron on Applied Projects and Portfolio Building Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback.13:11–20:14 · Matt pushing back 1/10 Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing.20:14–23:29 · Matt pushing back 0/10 Panelist Background: Claudia Perlich on Academic Research and Digital Advertising Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics.23:29–25:31 · Matt pushing back 1/10 Audience Q&A: Formulating Business Needs versus Loose Data Questions Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational.25:31–28:29 · Matt pushing back 0/10 Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator.28:29–38:31 · Matt pushing back 2/10 Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration.38:31–42:06 · Matt pushing back 0/10 Audience Q&A: Data Science in Academia and Degree Programs An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment.42:06–44:28 · Matt pushing back 0/10 Audience Q&A: Civic Data Science and Municipal Government Leadership Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation.44:28–48:13 · Matt pushing back 1/10 Audience Q&A: Information Geometry and Social Science Disciplines An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker.48:13–51:31 · Matt pushing back 0/10 Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely.51:31–56:11 · Matt pushing back 0/10 Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 37.4% · guest 62.6%0:00 · Matt 37.4% · guest 62.6%3:00 · Matt 1.1% · guest 98.9%3:00 · Matt 1.1% · guest 98.9%6:00 · Matt 19.4% · guest 80.6%6:00 · Matt 19.4% · guest 80.6%9:00 · Matt 4.8% · guest 95.2%9:00 · Matt 4.8% · guest 95.2%12:00 · Matt 0.7% · guest 99.3%12:00 · Matt 0.7% · guest 99.3%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 5.9% · guest 94.1%18:00 · Matt 5.9% · guest 94.1%21:00 · Matt 15.9% · guest 84.1%21:00 · Matt 15.9% · guest 84.1%24:00 · Matt 2.3% · guest 97.7%24:00 · Matt 2.3% · guest 97.7%27:00 · Matt 10.8% · guest 89.2%27:00 · Matt 10.8% · guest 89.2%30:00 · Matt 0% · guest 100%30:00 · Matt 0% · guest 100%33:00 · Matt 14.4% · guest 85.6%33:00 · Matt 14.4% · guest 85.6%36:00 · Matt 5.9% · guest 94.1%36:00 · Matt 5.9% · guest 94.1%39:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%42:00 · Matt 14.5% · guest 85.5%42:00 · Matt 14.5% · guest 85.5%45:00 · Matt 3.2% · guest 96.8%45:00 · Matt 3.2% · guest 96.8%48:00 · Matt 5.1% · guest 94.9%48:00 · Matt 5.1% · guest 94.9%51:00 · Matt 0.4% · guest 99.6%51:00 · Matt 0.4% · guest 99.6%54:00 · Matt 10.1% · guest 89.9%54:00 · Matt 10.1% · guest 89.9%
Sharpest disagreement ▶ 52:57 Max Shron directly counters Cathy O'Neil on feature selection

Max Shron interrupting and explicitly stating 'Strongly disagree' to Cathy O'Neil's assertion that feature selection is always paramount represents the transcript's most direct panelist clash.

Hardest push from Matt ▶ 28:42 Matt Turck reframes audience question around founder hiring choices

Matt Turck interrupts the general discussion to pivot the query specifically toward whether startup founders should prioritize vertical domain experience over horizontal technical skills.

Biggest teaching moment ▶ 18:30 Chris Wiggins corrects host on data science history

Chris Wiggins educates the host by challenging the notion that data science is a brand-new term, citing Bill Cleveland's 2001 proposal and roots extending back to Tukey and Deming in 1940 and 1962.

Matt holds his own ▶ 34:55 Matt Turck focuses the discussion on hiring practicalities for entrepreneurs

Matt Turck demonstrates domain awareness of his startup audience by clearly framing the core hiring trade-offs facing startup CEOs in data science.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Panelist Background: Cathy O'Neil on Mathematics, Social Science, and Fallibility 2341 Matt Turck opens with a general query about the panelists' educational backgrounds and career entry into data science. Cathy O'Neil and Drew Conway engage in a lively discussion on whether mathematicians or social scientists handle being wrong and understanding human aspects better. Host involvement remains strictly limited to setting the initial prompt and keeping time.
Panelist Background: Max Shron on Applied Projects and Portfolio Building 2221 Max Shron outlines his non-traditional route into data science via applied projects, portfolio building, and his time at OkCupid. Matt Turck briefly interjects to clarify whether Max was self-taught or mentored. The exchange remains friendly and conversational with minimal pushback.
Panelist Background: Chris Wiggins on Theoretical Physics, Biology, and Data Science Histo 2541 Chris Wiggins details his physics background before gently correcting the host's earlier assertion that data science is a brand-new term. Wiggins traces the intellectual history of the discipline back to John Tukey, W. Edwards Deming, and Bill Cleveland's 2001 paper. Matt Turck listens without challenging Wiggins's historical framing.
Panelist Background: Claudia Perlich on Academic Research and Digital Advertising 0110 Claudia Perlich delivers a monologue detailing her computer science background, academic career, IBM Watson experience, and ad tech work. The host does not speak during her monologue, resulting in zero scores for host metrics.
Audience Q&A: Formulating Business Needs versus Loose Data Questions 1211 Matt Turck opens the floor to audience Q&A, and an attendee asks about balancing tight versus loose data questions. Max Shron responds by advising practitioners to focus on defining a core business need first. The dynamic is collaborative and informational.
Audience Q&A: Ad Tech Metrics, Botnet Fraud, and Measurement Realities 1340 Audience member AB Mendez questions Claudia Perlich about ad tech metrics and botnet fraud. Claudia offers a candid, critical critique of ad industry incentives, describing metrics as a mess and calling fraud oversight an ostrich policy. The host acts purely as a session moderator.
Audience Q&A & Discussion: Industry Applications, Domain Expertise, and Social Impact 3352 After an audience member asks about preferred industries, Matt Turck reframes the topic by asking whether startups need deep domain expertise or horizontal technical skill. Cathy O'Neil presents a passionate, adversarial view of commercial data science as predatory modeling against vulnerable people. Drew Conway and Chris Wiggins offer counterbalances emphasizing domain collaboration.
Audience Q&A: Data Science in Academia and Degree Programs 0310 An audience member asks about academic degree programs in data science. Chris Wiggins and Drew Conway detail the rise of master's and upcoming PhD programs. The host does not intervene or comment during this segment.
Audience Q&A: Civic Data Science and Municipal Government Leadership 0200 Dawn Barber asks how city governments can utilize data science. Drew Conway answers by highlighting Mike Flowers's leadership as Chief Analytics Officer of NYC. Host engagement is nonexistent beyond moderation.
Audience Q&A: Information Geometry and Social Science Disciplines 2421 An audience member asks technical questions regarding information geometry and social science reading lists. Chris Wiggins explains Amari's 1988 work on Fisher metrics and its application in natural gradients, while Drew Conway outlines social science entry points. Matt Turck manages the mic and prompts the speaker.
Audience Q&A: Pain Points, Tedium, and Debugging in Data Science Work 0330 An audience member asks for the panelists' single biggest pain points and tedious tasks. The panelists share varied grievances ranging from buggy academic software to explaining error bars to resistant executives and formatting plots. The host allows the panel to respond freely.
Audience Q&A: Feature Selection, Feature Creation, and Classifier Complexity 1460 An audience question about feature selection vs classifier complexity sparks direct debate across the panel. Cathy O'Neil argues feature selection is always paramount, which Max Shron explicitly disagrees with, supported by Claudia Perlich. The panelists clash constructively over prediction accuracy versus model interpretability.

Statements from this episode (18)

Assertion Supported
Turck: In 2013, data science rapidly became a highly sought-after job
“The term sort of didn't exist a few years ago. It appeared, and now it's one of the most, ah, sought-after jobs”
Matt Turck Dec 5, 2013 ▶ 0:39
Insight
Conway: Human elements of data science are harder to teach than math
“I think it's much harder to learn and be trained on the human aspect than it is to be trained on the math and the computer science.”
Drew Conway Dec 5, 2013 ▶ 4:30
Insight
Cathy O'Neil: Bad data science stems from using algorithms without understanding them
“I think a lot of bad data science happens because people are like, I don't really know how this algorithm works. I'm hoping when I press this button, something good comes out. And look, it converged, so it must be okay.”
Cathy O'Neil Dec 5, 2013 ▶ 6:13
Opinion
Shron: PhD hires in data science often act entitled about personal research
“I feel like the people who I've seen hired in as PhDs often have a sense of entitlement about their, how much they're going to have to be able to work on their own problems at the exclusion of the needs of the company.”
Max Shron Dec 5, 2013 ▶ 8:20
Insight
Shron: Practical projects yield more value in data science than deep specialization
“There is a lot more to be gained, especially in data science, for having done projects than having necessarily spent two or three years going in depth on one topic.”
Max Shron Dec 5, 2013 ▶ 8:32
Assertion Partly supported
Shron: OkCupid data showed beer drinkers twice as open to first-date sex
“We found one of the most interesting ones was that people who said they liked the taste of beer were twice as likely to be open to, ah, having sex on a first date.”
Max Shron Dec 5, 2013 ▶ 12:09
Opinion
Claudia Perlich: Cleaning data was more valuable than subsequent academic research
“That was a very valuable experience, much more so than the research of what happened afterwards.”
Claudia Perlich Dec 5, 2013 ▶ 21:13
Opinion
Claudia Perlich: Digital advertising is the golden age of experimentation
“What I love about that, that job is it's a playground. It's a lovely playground. I can try it all. I mean, big deal if I make a mistake. I show you the wrong ad, ball, you know. I can really, it's like the golden age of experimentation.”
Claudia Perlich Dec 5, 2013 ▶ 22:42
Insight
Shron: Data science projects should start with a need, not a question
“So I don't think it's necessarily that you start with a question. I think you should start with a need. You should first trigger out, what is the problem I'm actually trying to solve before I do anything else?”
Max Shron Dec 5, 2013 ▶ 24:26
Opinion
Perlich: Online advertising metrics and analytics are a huge mess
“I mean, honestly, the whole question of metrics and analytics is a huge mess”
Claudia Perlich Dec 5, 2013 ▶ 26:25
Assertion Not checkable as stated
Perlich: Industry bonuses encourage advertising professionals to ignore ad fraud
“I mean, there are so many people who are much better off looking the wrong way when it comes to fraud. Everybody's bonus just basically hinges on getting the wrong metrics a little bit up.”
Claudia Perlich Dec 5, 2013 ▶ 27:06
Insight
Wiggins: Data science hiring should prioritize listening skills over domain expertise
“So, I think what you're looking for is not a particularly somebody with a domain background, but somebody who's proven themselves to be a good listener.”
Chris Wiggins Dec 5, 2013 ▶ 35:45
Opinion
O'Neil: Data science is a war of moneyed interests against vulnerable people
“I think of data science as And the general modelization of everything in sight, including education, including getting a job insurance, health, it's a war. And we're losing. Like we are, this is a war of the people who have money and can go hire data scientist…”
Cathy O'Neil Dec 5, 2013 ▶ 37:10
Insight
Wiggins: Students incorrectly assume that published academic papers are inherently true
“My biggest pain point is, is trying to re-educate students who have read a bad paper, and because it was published, they think it's true.”
Chris Wiggins Dec 5, 2013 ▶ 50:18
Insight
O'Neil: The hardest data science role is conveying fundamental uncertainty
“The hardest role you have with data scientists, and this is when you get the label negative, is when you're saying, we actually don't know the answer to this, and you don't either, and no one knows the answer, and we can't pretend to know the answer.”
Cathy O'Neil Dec 5, 2013 ▶ 50:35
Insight
O'Neil: Data problems that bypass feature selection will not yield financial returns
“Feature selection is always gonna be very important. If you have something that's so easy to answer that you don't have to care about feature selection, you're not gonna make any money doing it.”
Cathy O'Neil Dec 5, 2013 ▶ 52:50
Disclosure
Perlich: Digital advertising machine learning models operate in ten million dimensions
“Today we are building models in, I build models in, you know, ten million dimensions with, on a good day, a 100,000 positives, on a bad day, 5000, and good, it's not good enough to survive on Wall Street, but hey, for advertising, it's just fine.”
Claudia Perlich Dec 5, 2013 ▶ 53:38
Insight
Conway: Extra model features carry statistical and cognitive explainability penalties
“Think very parsimoniously about what you bring into your model, because with each one of those, not only are you incurring a statistical penalty, you're incurring a sort of brain penalty and having to explain why the result is what it is when you get to whatev…”
Drew Conway Dec 5, 2013 ▶ 54:56
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.