Sep 14, 2015 · 22m · mad

Lukas Biewald, CrowdFlower // Enriching Your Data (Hosted by FirstMark Capital)

Lukas Biewald · 18m spoken Matt Turck · 20s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In a DataDrivenNYC presentation, CrowdFlower CEO Lukas Biewald demonstrates why increasing dataset size and data quality impacts AI performance more than algorithm tuning, while advocating for open data sharing through real-world case studies and company initiatives.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.7% of the talking time here. How this is scored →

Matt as informed peer 0.4 Guest teaching 0.4 Guest disagreement 0.2 Matt pushing back 0.2
05100:0010:0020:000:08–4:25 · Matt as informed peer 0/10 Presentation Opening: The Thesis for Open Data This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero.4:25–9:24 · Matt as informed peer 0/10 The Impact of Cleaner Data on Machine Learning Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment.9:24–12:50 · Matt as informed peer 0/10 Case Study: Apple Watch Sentiment and Open Media Data Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present.12:50–15:04 · Matt as informed peer 0/10 Exploring CrowdFlower Open Datasets and Kaggle Competitions Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation.15:04–22:59 · Matt as informed peer 2/10 Keynote Conclusion and Open Data Resources Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction.0:08–4:25 · Guest teaching 0/10 Presentation Opening: The Thesis for Open Data This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero.4:25–9:24 · Guest teaching 0/10 The Impact of Cleaner Data on Machine Learning Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment.9:24–12:50 · Guest teaching 0/10 Case Study: Apple Watch Sentiment and Open Media Data Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present.12:50–15:04 · Guest teaching 0/10 Exploring CrowdFlower Open Datasets and Kaggle Competitions Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation.15:04–22:59 · Guest teaching 2/10 Keynote Conclusion and Open Data Resources Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction.0:08–4:25 · Guest disagreement 0/10 Presentation Opening: The Thesis for Open Data This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero.4:25–9:24 · Guest disagreement 0/10 The Impact of Cleaner Data on Machine Learning Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment.9:24–12:50 · Guest disagreement 0/10 Case Study: Apple Watch Sentiment and Open Media Data Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present.12:50–15:04 · Guest disagreement 0/10 Exploring CrowdFlower Open Datasets and Kaggle Competitions Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation.15:04–22:59 · Guest disagreement 1/10 Keynote Conclusion and Open Data Resources Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction.0:08–4:25 · Matt pushing back 0/10 Presentation Opening: The Thesis for Open Data This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero.4:25–9:24 · Matt pushing back 0/10 The Impact of Cleaner Data on Machine Learning Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment.9:24–12:50 · Matt pushing back 0/10 Case Study: Apple Watch Sentiment and Open Media Data Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present.12:50–15:04 · Matt pushing back 0/10 Exploring CrowdFlower Open Datasets and Kaggle Competitions Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation.15:04–22:59 · Matt pushing back 1/10 Keynote Conclusion and Open Data Resources Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 11.9% · guest 88.1%15:00 · Matt 11.9% · guest 88.1%18:00 · Matt 0.2% · guest 99.8%18:00 · Matt 0.2% · guest 99.8%21:00 · Matt 2.5% · guest 97.5%21:00 · Matt 2.5% · guest 97.5%
Sharpest disagreement ▶ 21:42 Dismissing hedge fund market prediction premise

Lukas politely rejects an audience member's premise about using CrowdFlower for trading, explaining that asking a hundred crowd workers about stock movements provides no signal.

Hardest push from Matt ▶ 15:36 Host re-anchoring from presentation to company product

Matt Turck steps in at the end of the keynote to pivot the topic away from open data theory toward CrowdFlower's commercial operations.

Biggest teaching moment ▶ 15:50 Preempting and clarifying Mechanical Turk differences

Lukas anticipates the standard audience question regarding Amazon Mechanical Turk, detailing CrowdFlower's specialized focus on machine learning quality control.

Matt holds his own ▶ 15:36 Host citing data enrichment product capabilities

Matt Turck demonstrates solid domain awareness by referencing specific features like data enrichment and categorization when introducing the Q&A section.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Presentation Opening: The Thesis for Open Data 0000 This segment is a uninterrupted monologue presentation by guest Lukas Biewald on the value of data quantity over complex algorithms. Because the host does not speak, host expertise and pushback are scored zero.
The Impact of Cleaner Data on Machine Learning 0000 Lukas continues his presentation monologue, illustrating how data cleanup significantly reduces error rates and sharing early CrowdFlower dataset experiments. The host remains silent throughout the segment.
Case Study: Apple Watch Sentiment and Open Media Data 0000 Lukas presents case studies on Apple Watch sentiment and CrowdFlower's Data for Everyone initiative in a solo talk format. There is no host involvement or combativeness present.
Exploring CrowdFlower Open Datasets and Kaggle Competitions 0000 Lukas concludes his keynote presentation by reviewing open datasets and Kaggle search relevance competitions. The segment consists entirely of monologue presentation.
Keynote Conclusion and Open Data Resources 2211 Host Matt Turck opens Q&A with polite operational questions, while audience members ask about crowd quality and use cases. Lukas cordially explains the platform mechanics and gently clarifies that crowdsourcing is unsuitable for stock market prediction.

Statements from this episode (8)

Insight
Biewald: Only five out of 100 feature engineering attempts actually improve models
“Every task that you work on has kind of different different features work better or worse, and I spent, when I first was working as a data scientist, I spent all my time on this feature selection, and as you guys know, you try a hundred things and maybe five o…”
Lukas Biewald Sep 14, 2015 ▶ 2:45
Assertion Not checkable as stated
Increasing dataset accuracy from 90% to 95% repeatedly halves error rates
“So this is my same data set that I published online, but if you take it from 90% to 95%, you actually have the error rate, and then going up to a hundred percent, you have it again, right?”
Lukas Biewald Sep 14, 2015 ▶ 4:43
Disclosure
CrowdFlower spent just $100 to collect 10,000 initial color labels
“We collected that data. It maybe cost us a hundred bucks to get 10,000 labels.”
Lukas Biewald Sep 14, 2015 ▶ 6:55
Disclosure
CrowdFlower supplied raw Apple Watch survey data to journalists instead of PR
“Instead of sending out to journalists the our, like, the crowd flower analysis, we actually sent them the data and kind of let them draw their own conclusions.”
Lukas Biewald Sep 14, 2015 ▶ 9:57
Assertion Not checkable as stated
CrowdFlower claims to have the largest dataset determining if images are funny
“We have a gigantic, maybe the biggest data set available on, isn't image funny?”
Lukas Biewald Sep 14, 2015 ▶ 14:12
Disclosure
CrowdFlower pays all global crowd workers the same flat compensation rate
“We pay everyone the same regardless of what country you come in from.”
Lukas Biewald Sep 14, 2015 ▶ 19:13
Assertion Not checkable as stated
Thirty percent of CrowdFlower's crowdsourced workforce is based in the US
“It's also 30% US-based.”
Lukas Biewald Sep 14, 2015 ▶ 19:28
Disclosure
CrowdFlower charges a platform fee plus up to a 20% take rate
“We have a platform fee to use our software, and then we take a, up to a 20% cut of the work that, that flows through our platform.”
Lukas Biewald Sep 14, 2015 ▶ 22:41
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.