Dec 5, 2013 · 58m · mad

Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012

Mike Driscoll · 16m spoken Matt Turck · 2m spoken Alex Jacqueline · 58s spoken Michael Selick · 53s spoken Carter Schoenwald · 50s spoken Garrett Vidal · 39s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC #4, panel moderator Matt Turck hosts industry leaders Mike Driscoll, Jeremy Howard, and Sean Gourley to discuss early-stage startup data infrastructure, talent acquisition strategies, enterprise hype, and the evolving landscape of competitive machine learning.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.6% of the talking time here. How this is scored →

Matt as informed peer 1.9 Guest teaching 5.6 Guest disagreement 2.4 Matt pushing back 1.4
05100:0015:0030:0045:005:04–11:21 · Matt as informed peer 4/10 Hiring Data Science Talent: Domain vs. Machine Learning Expertise Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds.11:21–14:49 · Matt as informed peer 1/10 Kaggle Competitions and Private Market Liquidity An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations.14:49–26:59 · Matt as informed peer 1/10 Consumer Algorithms, Data Conglomerates, and Targeting Limits Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems.26:59–32:18 · Matt as informed peer 1/10 Data Team Ratios, Infrastructure Contracts, and Coding Skills An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice.32:18–41:49 · Matt as informed peer 4/10 Industry Hype, Enterprise Sales, and the Hadoop Mandate An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop.41:49–45:46 · Matt as informed peer 1/10 The Evolution from Data Mining to Data Science An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization.45:46–48:32 · Matt as informed peer 3/10 Core Purpose of Data Science and Automation of Models Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks.48:32–54:01 · Matt as informed peer 1/10 Integrating Machine Learning Algorithms into Production Architecture An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP.54:01–58:11 · Matt as informed peer 1/10 Data Quality, Probabilistic Mindset, and Panel Conclusion In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets.5:04–11:21 · Guest teaching 5/10 Hiring Data Science Talent: Domain vs. Machine Learning Expertise Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds.11:21–14:49 · Guest teaching 6/10 Kaggle Competitions and Private Market Liquidity An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations.14:49–26:59 · Guest teaching 7/10 Consumer Algorithms, Data Conglomerates, and Targeting Limits Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems.26:59–32:18 · Guest teaching 5/10 Data Team Ratios, Infrastructure Contracts, and Coding Skills An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice.32:18–41:49 · Guest teaching 6/10 Industry Hype, Enterprise Sales, and the Hadoop Mandate An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop.41:49–45:46 · Guest teaching 5/10 The Evolution from Data Mining to Data Science An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization.45:46–48:32 · Guest teaching 5/10 Core Purpose of Data Science and Automation of Models Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks.48:32–54:01 · Guest teaching 6/10 Integrating Machine Learning Algorithms into Production Architecture An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP.54:01–58:11 · Guest teaching 5/10 Data Quality, Probabilistic Mindset, and Panel Conclusion In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets.5:04–11:21 · Guest disagreement 2/10 Hiring Data Science Talent: Domain vs. Machine Learning Expertise Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds.11:21–14:49 · Guest disagreement 1/10 Kaggle Competitions and Private Market Liquidity An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations.14:49–26:59 · Guest disagreement 5/10 Consumer Algorithms, Data Conglomerates, and Targeting Limits Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems.26:59–32:18 · Guest disagreement 3/10 Data Team Ratios, Infrastructure Contracts, and Coding Skills An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice.32:18–41:49 · Guest disagreement 4/10 Industry Hype, Enterprise Sales, and the Hadoop Mandate An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop.41:49–45:46 · Guest disagreement 2/10 The Evolution from Data Mining to Data Science An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization.45:46–48:32 · Guest disagreement 2/10 Core Purpose of Data Science and Automation of Models Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks.48:32–54:01 · Guest disagreement 2/10 Integrating Machine Learning Algorithms into Production Architecture An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP.54:01–58:11 · Guest disagreement 1/10 Data Quality, Probabilistic Mindset, and Panel Conclusion In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets.5:04–11:21 · Matt pushing back 3/10 Hiring Data Science Talent: Domain vs. Machine Learning Expertise Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds.11:21–14:49 · Matt pushing back 0/10 Kaggle Competitions and Private Market Liquidity An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations.14:49–26:59 · Matt pushing back 1/10 Consumer Algorithms, Data Conglomerates, and Targeting Limits Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems.26:59–32:18 · Matt pushing back 0/10 Data Team Ratios, Infrastructure Contracts, and Coding Skills An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice.32:18–41:49 · Matt pushing back 3/10 Industry Hype, Enterprise Sales, and the Hadoop Mandate An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop.41:49–45:46 · Matt pushing back 0/10 The Evolution from Data Mining to Data Science An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization.45:46–48:32 · Matt pushing back 6/10 Core Purpose of Data Science and Automation of Models Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks.48:32–54:01 · Matt pushing back 0/10 Integrating Machine Learning Algorithms into Production Architecture An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP.54:01–58:11 · Matt pushing back 0/10 Data Quality, Probabilistic Mindset, and Panel Conclusion In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 39.9% · guest 60.1%0:00 · Matt 39.9% · guest 60.1%3:00 · Matt 15.4% · guest 84.6%3:00 · Matt 15.4% · guest 84.6%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 5.6% · guest 94.4%9:00 · Matt 5.6% · guest 94.4%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%24:00 · Matt 1.7% · guest 98.3%24:00 · Matt 1.7% · guest 98.3%27:00 · Matt 1.3% · guest 98.7%27:00 · Matt 1.3% · guest 98.7%30:00 · Matt 0% · guest 100%30:00 · Matt 0% · guest 100%33:00 · Matt 0% · guest 100%33:00 · Matt 0% · guest 100%36:00 · Matt 6.5% · guest 93.5%36:00 · Matt 6.5% · guest 93.5%39:00 · Matt 5.8% · guest 94.2%39:00 · Matt 5.8% · guest 94.2%42:00 · Matt 0% · guest 100%42:00 · Matt 0% · guest 100%45:00 · Matt 6.4% · guest 93.6%45:00 · Matt 6.4% · guest 93.6%48:00 · Matt 0% · guest 100%48:00 · Matt 0% · guest 100%51:00 · Matt 5.3% · guest 94.7%51:00 · Matt 5.3% · guest 94.7%54:00 · Matt 0.8% · guest 99.2%54:00 · Matt 0.8% · guest 99.2%57:00 · Matt 4.4% · guest 95.6%57:00 · Matt 4.4% · guest 95.6%
Sharpest disagreement ▶ 35:10 Direct dismissal over chemoinformatics literature

Jeremy Howard dismisses Mike Driscoll's speculation on pharma R&D progress by directly stating 'You don't know' regarding recent machine learning literature in chemoinformatics.

Hardest push from Matt ▶ 46:00 Host shuts down audience promotional pitch

Host Matt Turck immediately cuts off an audience questioner who starts giving a marketing pitch for her company, instructing her to skip the intro and ask a question directly.

Biggest teaching moment ▶ 19:40 Correcting the misapplication of Norvig's data thesis

Jeremy Howard corrects Mike Driscoll on Peter Norvig's quote regarding data overriding algorithms, clarifying that Norvig specifically referred to human language translation rather than standard tabular machine learning problems.

Matt holds his own ▶ 5:06 Host demonstrates industry knowledge via Strata debate

Host Matt Turck demonstrates domain knowledge by citing a recent Strata conference debate to formulate a structured question on domain expertise versus pure machine learning skill.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Hiring Data Science Talent: Domain vs. Machine Learning Expertise 4523 Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds.
Kaggle Competitions and Private Market Liquidity 1610 An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations.
Consumer Algorithms, Data Conglomerates, and Targeting Limits 1751 Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems.
Data Team Ratios, Infrastructure Contracts, and Coding Skills 1530 An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice.
Industry Hype, Enterprise Sales, and the Hadoop Mandate 4643 An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop.
The Evolution from Data Mining to Data Science 1520 An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization.
Core Purpose of Data Science and Automation of Models 3526 Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks.
Integrating Machine Learning Algorithms into Production Architecture 1620 An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP.
Data Quality, Probabilistic Mindset, and Panel Conclusion 1510 In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets.

Statements from this episode (10)

Insight
Mike Driscoll: Solving real data problems beats taking data science courses
“I think that a lot of people who are interested in data science will, Decide that, you know, I gotta, I should go take a bunch of courses and read a bunch of books, and check a series of boxes, you know and I think that's actually the wrong approach. I think t…”
Mike Driscoll Dec 5, 2013 ▶ 1:56
Insight
Mike Driscoll: Hire ML talent for structured problems, domain experts otherwise
“If your data is already structured well enough, and you've got a set of criteria to define the success of your machine learning algorithm, then you should hire the strongest pure machine learner you can find. You should find someone who would, you know, who's …”
Mike Driscoll Dec 5, 2013 ▶ 6:19
Insight
Mike Driscoll: Reject data science candidates who fail basic command-line tasks
“If someone can't get an answer to that, cannot pull that column off and turn it into a human readable date string, In the order of minutes, you do not want that person as a data scientist.”
Mike Driscoll Dec 5, 2013 ▶ 9:41
Prediction Not checkable as stated
Mike Driscoll: Data-holding entities will win as algorithms move to data
“Analytics and algorithms are light. Data is very heavy. So the natural course of things will be that the algorithms will move to where the data lives. And the entities that have the data will ultimately win because of that strategic advantage”
Mike Driscoll Dec 5, 2013 ▶ 18:00
Assertion Not checkable as stated
Mike Driscoll: Telcos hold the world's most powerfully predictive social graphs
“If you want to know who has the most powerfully predictive social graphs on the planet, it's the telcos, because who you call and who you text is an extraordinarily strong indicator of who you actually are connected to in a social way.”
Mike Driscoll Dec 5, 2013 ▶ 24:06
Assertion Not checkable as stated
Mike Driscoll: Installing Hive on a Hadoop cluster increases usage by 10x
“Once you install Hive onto Hadoop cluster, usage of that cluster typically goes up by a factor of 10 because you lower the friction of getting data out of that system.”
Mike Driscoll Dec 5, 2013 ▶ 31:06
Assertion Partly supported
Mike Driscoll: Big pharma R&D cuts create data science recruiting opportunities
“Most big pharma actually have turned, reduced their investment in R&D, and many of, it's a great place to recruit data scientists is out of these pharma R&D labs.”
Mike Driscoll Dec 5, 2013 ▶ 35:20
Opinion
Mike Driscoll: Klout is a bad company due to poor algorithms
“Clout's maybe a crappy company because their algorithms are bad”
Mike Driscoll Dec 5, 2013 ▶ 36:05
Opinion
Mike Driscoll: Running algorithms via Apache Mahout on Hadoop is too slow
“I think the problem with Mahoot is that anything, it's, many of these things are, if you run in Hadoop, you're slow. You need to be able to run in an environment that's fast”
Mike Driscoll Dec 5, 2013 ▶ 53:02
Insight
Mike Driscoll: Combining human data labeling with machine learning outperforms pure automation
“People make this sort of distinction of either it's machine learning or it's people. Either you've got a team in India or, you know, Romania working through the data, trying to make sense of it or you've got some algorithm working on it, but I think there is a…”
Mike Driscoll Dec 5, 2013 ▶ 57:32
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.