Jan 2, 2019 · 34m · a16z
a16z Podcast | It's Not What You Say, It's How You Say It -- When Language Meets Big Data
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the a16z Podcast, Textio CEO and co-founder Kieran Snyder discusses how machine learning and natural language processing analyze written text to predict candidate draw, uncover gender bias in workplace performance reviews, and optimize real-world business outcomes.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 21.9% of the talking time here. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Kieran firmly refutes Michael's suggestion that outside variables like seasonality or geography account for Kickstarter success, stating those factors made little statistical difference.
Hardest push from the host ▶ 30:53 Michael challenges optimization limits citing Demand MediaMichael challenges the core premise of language optimization, questioning whether algorithmic tuning inevitably results in bland, homogenized content like Demand Media.
Biggest teaching moment ▶ 27:36 Kieran highlights gendered evaluation language in performance reviewsKieran educates the hosts on empirical findings from performance reviews, revealing how identical traits like aggressiveness are praised in men but criticized as abrasiveness in women.
The host holds their own ▶ 16:12 Sonal contextualizes text analytics within 30-year NLP historySonal demonstrates deep domain knowledge by discussing the 30-year progression of natural language processing and contrasting early data collection limits with modern machine learning corpora.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Predicting Success: Applying Machine Learning to Kickstarter Text | 2 | 6 | 1 | 3 | Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions. | |
| Text Analytics for Job Listings and Shifting Vocabulary Trends | 3 | 5 | 1 | 2 | Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time. | |
| Corporate Jargon, Authorship Intent, and Emerging Industry Terms | 4 | 4 | 1 | 1 | Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings. | |
| Training Models with Outcome Signals and the Evolution of NLP | 5 | 5 | 0 | 1 | Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling. | |
| Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews | 4 | 6 | 0 | 1 | Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals. | |
| Language Optimization Feedback Loops and Podcast Conclusion | 5 | 4 | 1 | 6 | Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal. |