Dec 5, 2013 · 58m · mad
Panel: Metamarkets, Kaggle and Quid // Data Driven NYC #4 // Mar 2012
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At Data Driven NYC #4, panel moderator Matt Turck hosts industry leaders Mike Driscoll, Jeremy Howard, and Sean Gourley to discuss early-stage startup data infrastructure, talent acquisition strategies, enterprise hype, and the evolving landscape of competitive machine learning.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.6% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Jeremy Howard dismisses Mike Driscoll's speculation on pharma R&D progress by directly stating 'You don't know' regarding recent machine learning literature in chemoinformatics.
Hardest push from Matt ▶ 46:00 Host shuts down audience promotional pitchHost Matt Turck immediately cuts off an audience questioner who starts giving a marketing pitch for her company, instructing her to skip the intro and ask a question directly.
Biggest teaching moment ▶ 19:40 Correcting the misapplication of Norvig's data thesisJeremy Howard corrects Mike Driscoll on Peter Norvig's quote regarding data overriding algorithms, clarifying that Norvig specifically referred to human language translation rather than standard tabular machine learning problems.
Matt holds his own ▶ 5:06 Host demonstrates industry knowledge via Strata debateHost Matt Turck demonstrates domain knowledge by citing a recent Strata conference debate to formulate a structured question on domain expertise versus pure machine learning skill.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Hiring Data Science Talent: Domain vs. Machine Learning Expertise | 4 | 5 | 2 | 3 | Host Matt Turck cites a recent Strata conference debate to frame a question about hiring machine learning generalists versus domain experts. The guests expand on the premise, using Kaggle data and practical coding tests to explain how pragmatism outweighs traditional academic backgrounds. | |
| Kaggle Competitions and Private Market Liquidity | 1 | 6 | 1 | 0 | An audience member asks why Kaggle prize money seems low given the business value generated. Jeremy Howard educates the room on unpublicized private competitions, market liquidity, and participant motivations. | |
| Consumer Algorithms, Data Conglomerates, and Targeting Limits | 1 | 7 | 5 | 1 | Jeremy Howard vigorously challenges Mike Driscoll's claim that more data always beats better algorithms, pointing out that Peter Norvig's famous quote applied specifically to complex natural language translation rather than standard tabular machine learning problems. | |
| Data Team Ratios, Infrastructure Contracts, and Coding Skills | 1 | 5 | 3 | 0 | An audience member asks about data team support ratios. Driscoll and Howard debate the technical boundaries for data scientists, discussing whether writing MapReduce in Java is necessary or if SQL and high-level scripting suffice. | |
| Industry Hype, Enterprise Sales, and the Hadoop Mandate | 4 | 6 | 4 | 3 | An audience question about biotech hype leads to direct friction when Howard bluntly tells Driscoll he does not know the current chemoinformatics literature. Host Matt Turck steers the conversation toward enterprise sales challenges and the top-down mandate for Hadoop. | |
| The Evolution from Data Mining to Data Science | 1 | 5 | 2 | 0 | An audience member asks what truly changed between the era of data mining and data science. The guests reframe the issue, emphasizing open-source tooling, lower barriers to entry, and community democratization. | |
| Core Purpose of Data Science and Automation of Models | 3 | 5 | 2 | 6 | Host Matt Turck forcefully intervenes when an audience questioner turns her introduction into an explicit company sales pitch. Jeremy Howard then draws an analogy between 1940s human computers and the inevitable automation of data science tasks. | |
| Integrating Machine Learning Algorithms into Production Architecture | 1 | 6 | 2 | 0 | An audience member asks how model design accounts for production backend architecture. The guests explain a two-step process where optimal statistical models are simplified or translated into production code like PHP. | |
| Data Quality, Probabilistic Mindset, and Panel Conclusion | 1 | 5 | 1 | 0 | In response to a query about messy enterprise data quality, the panelists explain that organizations must adopt a probabilistic mindset rather than expecting deterministic perfection from real-world datasets. |