Sep 14, 2015 · 22m · mad
Lukas Biewald, CrowdFlower // Enriching Your Data (Hosted by FirstMark Capital)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In a DataDrivenNYC presentation, CrowdFlower CEO Lukas Biewald demonstrates why increasing dataset size and data quality impacts AI performance more than algorithm tuning, while advocating for open data sharing through real-world case studies and company initiatives.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.7% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Lukas politely rejects an audience member's premise about using CrowdFlower for trading, explaining that asking a hundred crowd workers about stock movements provides no signal.
Hardest push from Matt ▶ 15:36 Host re-anchoring from presentation to company productMatt Turck steps in at the end of the keynote to pivot the topic away from open data theory toward CrowdFlower's commercial operations.
Biggest teaching moment ▶ 15:50 Preempting and clarifying Mechanical Turk differencesLukas anticipates the standard audience question regarding Amazon Mechanical Turk, detailing CrowdFlower's specialized focus on machine learning quality control.
Matt holds his own ▶ 15:36 Host citing data enrichment product capabilitiesMatt Turck demonstrates solid domain awareness by referencing specific features like data enrichment and categorization when introducing the Q&A section.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Presentation Opening: The Thesis for Open Data | 0 | 0 | 0 | 0 | This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero. | |
| The Impact of Cleaner Data on Machine Learning | 0 | 0 | 0 | 0 | Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment. | |
| Case Study: Apple Watch Sentiment and Open Media Data | 0 | 0 | 0 | 0 | Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present. | |
| Exploring CrowdFlower Open Datasets and Kaggle Competitions | 0 | 0 | 0 | 0 | Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation. | |
| Keynote Conclusion and Open Data Resources | 2 | 2 | 1 | 1 | Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction. |