Jan 2, 2019 · 34m · a16z

a16z Podcast | It's Not What You Say, It's How You Say It -- When Language Meets Big Data

Kieran Snyder · 22m spoken Sonal Chokshi · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, Textio CEO and co-founder Kieran Snyder discusses how machine learning and natural language processing analyze written text to predict candidate draw, uncover gender bias in workplace performance reviews, and optimize real-world business outcomes.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 21.9% of the talking time here. How this is scored →

The host as informed peer 3.8 Guest teaching 5.0 Guest disagreement 0.7 The host pushing back 2.3
05100:0010:0020:0030:001:50–5:43 · The host as informed peer 2/10 Predicting Success: Applying Machine Learning to Kickstarter Text Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions.5:43–10:00 · The host as informed peer 3/10 Text Analytics for Job Listings and Shifting Vocabulary Trends Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time.10:00–13:57 · The host as informed peer 4/10 Corporate Jargon, Authorship Intent, and Emerging Industry Terms Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings.13:57–21:12 · The host as informed peer 5/10 Training Models with Outcome Signals and the Evolution of NLP Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling.21:12–30:53 · The host as informed peer 4/10 Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals.30:53–34:06 · The host as informed peer 5/10 Language Optimization Feedback Loops and Podcast Conclusion Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal.1:50–5:43 · Guest teaching 6/10 Predicting Success: Applying Machine Learning to Kickstarter Text Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions.5:43–10:00 · Guest teaching 5/10 Text Analytics for Job Listings and Shifting Vocabulary Trends Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time.10:00–13:57 · Guest teaching 4/10 Corporate Jargon, Authorship Intent, and Emerging Industry Terms Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings.13:57–21:12 · Guest teaching 5/10 Training Models with Outcome Signals and the Evolution of NLP Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling.21:12–30:53 · Guest teaching 6/10 Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals.30:53–34:06 · Guest teaching 4/10 Language Optimization Feedback Loops and Podcast Conclusion Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal.1:50–5:43 · Guest disagreement 1/10 Predicting Success: Applying Machine Learning to Kickstarter Text Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions.5:43–10:00 · Guest disagreement 1/10 Text Analytics for Job Listings and Shifting Vocabulary Trends Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time.10:00–13:57 · Guest disagreement 1/10 Corporate Jargon, Authorship Intent, and Emerging Industry Terms Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings.13:57–21:12 · Guest disagreement 0/10 Training Models with Outcome Signals and the Evolution of NLP Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling.21:12–30:53 · Guest disagreement 0/10 Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals.30:53–34:06 · Guest disagreement 1/10 Language Optimization Feedback Loops and Podcast Conclusion Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal.1:50–5:43 · The host pushing back 3/10 Predicting Success: Applying Machine Learning to Kickstarter Text Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions.5:43–10:00 · The host pushing back 2/10 Text Analytics for Job Listings and Shifting Vocabulary Trends Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time.10:00–13:57 · The host pushing back 1/10 Corporate Jargon, Authorship Intent, and Emerging Industry Terms Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings.13:57–21:12 · The host pushing back 1/10 Training Models with Outcome Signals and the Evolution of NLP Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling.21:12–30:53 · The host pushing back 1/10 Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals.30:53–34:06 · The host pushing back 6/10 Language Optimization Feedback Loops and Podcast Conclusion Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 57.9% · guest 42.1%0:00 · the host 57.9% · guest 42.1%3:00 · the host 9.1% · guest 90.9%3:00 · the host 9.1% · guest 90.9%6:00 · the host 26.4% · guest 73.6%6:00 · the host 26.4% · guest 73.6%9:00 · the host 28.7% · guest 71.3%9:00 · the host 28.7% · guest 71.3%12:00 · the host 10.5% · guest 89.5%12:00 · the host 10.5% · guest 89.5%15:00 · the host 27.8% · guest 72.2%15:00 · the host 27.8% · guest 72.2%18:00 · the host 14.2% · guest 85.8%18:00 · the host 14.2% · guest 85.8%21:00 · the host 10.9% · guest 89.1%21:00 · the host 10.9% · guest 89.1%24:00 · the host 32.3% · guest 67.7%24:00 · the host 32.3% · guest 67.7%27:00 · the host 18.8% · guest 81.2%27:00 · the host 18.8% · guest 81.2%30:00 · the host 3% · guest 97%30:00 · the host 3% · guest 97%33:00 · the host 25.5% · guest 74.5%33:00 · the host 25.5% · guest 74.5%
Sharpest disagreement ▶ 4:46 Kieran dismisses timing and geography as primary success factors

Kieran firmly refutes Michael's suggestion that outside variables like seasonality or geography account for Kickstarter success, stating those factors made little statistical difference.

Hardest push from the host ▶ 30:53 Michael challenges optimization limits citing Demand Media

Michael challenges the core premise of language optimization, questioning whether algorithmic tuning inevitably results in bland, homogenized content like Demand Media.

Biggest teaching moment ▶ 27:36 Kieran highlights gendered evaluation language in performance reviews

Kieran educates the hosts on empirical findings from performance reviews, revealing how identical traits like aggressiveness are praised in men but criticized as abrasiveness in women.

The host holds their own ▶ 16:12 Sonal contextualizes text analytics within 30-year NLP history

Sonal demonstrates deep domain knowledge by discussing the 30-year progression of natural language processing and contrasting early data collection limits with modern machine learning corpora.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Predicting Success: Applying Machine Learning to Kickstarter Text 2613 Sonal and Michael prompt Kieran on her Kickstarter study findings, with Michael probing on how the model accounts for potential confounding variables like seasonality and geography. Kieran explains how counterintuitive text features like typography mix and word count drive predictive accuracy over traditional design assumptions.
Text Analytics for Job Listings and Shifting Vocabulary Trends 3512 Sonal brings up external research on loan default predictions tied to specific word choices, while Michael asks how software maintains signal when best practices become widely adopted. Kieran details how language impact evolves, citing how the term big data shifted from positive to neutral in tech job posts over time.
Corporate Jargon, Authorship Intent, and Emerging Industry Terms 4411 Sonal reflects on lexicography and how dynamic text data surpasses static dictionaries in capturing changing language usage. Kieran highlights how original authorship outperforms patched job descriptions and details how gateway corporate jargon like synergy negatively impacts listings.
Training Models with Outcome Signals and the Evolution of NLP 5501 Sonal frames the discussion around the 30-year evolution of natural language processing and asks whether domain-specific training corpora or universal principles apply across verticals. Kieran clarifies that outcome signal data, rather than raw text volume, is the critical bottleneck in predictive modeling.
Real-Time Editing Feedback, Resumes, and Gender Bias in Reviews 4601 Sonal connects Kieran's findings on gender bias in reviews and resumes to academic debates on qualitative versus quantitative methodologies. Kieran details how terms like abrasive and aggressive are applied unevenly across gender lines in performance appraisals.
Language Optimization Feedback Loops and Podcast Conclusion 5416 Michael challenges the premise of language optimization by bringing up Demand Media and questioning whether algorithmic optimization leads to bland, formulaic content. Kieran addresses the challenge by explaining that dynamic learning systems adjust as overused phrases lose signal.

Statements from this episode (14)

Assertion Not checkable as stated
Snyder: Textio predicted Kickstarter fundraising success with over 90% accuracy
“We got over 90% predictive on minute zero of a project as to whether it was going to hit its fundraising goal based solely on things like how long is the text and what kind of fonts are you using and how many headings do you have.”
Kieran Snyder Jan 2, 2019 ▶ 2:59
Insight
Snyder: Longer project descriptions perform better on Kickstarter
“Longer is better where Kickstarter is is concerned kind of counterintuitive.”
Kieran Snyder Jan 2, 2019 ▶ 3:29
Insight
Snyder: Kickstarter campaigns succeed when formatted like a "ransom note"
“You want it to look like a ransom note, so you want to mix and match types. You want lots and lots of headings.”
Kieran Snyder Jan 2, 2019 ▶ 3:48
Insight
Snyder: Kickstarter success depends on text structure rather than idea quality
“The quality of your idea doesn't matter just looking at the content aspects we could predict.”
Kieran Snyder Jan 2, 2019 ▶ 4:12
Assertion Not checkable as stated
Snyder: Mentioning off-street parking harms higher-priced home listings
“So we saw when we were prototyping out the real estate stuff that if you say off street parking that really moves the needle for low income homes, but for high income homes in terms of the number of people who go to your open house and then the eventual sale p…”
Kieran Snyder Jan 2, 2019 ▶ 7:15
Assertion Not checkable as stated
Snyder: Textio identified 25,000 job description phrases affecting recruiting outcomes
“In jobs, it matters hugely. You know, we, we've identified at this point over 25,000 unique phrases that move the needle on how many people will apply for your job, what demographics, how qualified they are.”
Kieran Snyder Jan 2, 2019 ▶ 7:43
Assertion Not checkable as stated
Snyder: "Big data" in job posts lost its positive impact by 2015
“So my favorite example of this is the phrase big data. So a year and a half ago, if you use the phrase big data in a tech job listing, it was positive. You know, it was seen as compelling and cutting edge. In June of 2015, it's not negative, but it's totally n…”
Kieran Snyder Jan 2, 2019 ▶ 8:58
Assertion Not checkable as stated
Snyder: The word "synergy" torpedoes job listing candidate response rates
“The biggest, you know, the, one of the very common, we call it a gateway term that kind of torpedoes your listing is the word synergy. But it's a gateway term because when people include synergy, they're also significantly more likely to include, you know, val…”
Kieran Snyder Jan 2, 2019 ▶ 11:04
Assertion Not checkable as stated
Snyder: 'People analytics' has replaced 'workforce analytics' in effective job postings
“Turns out workforce analytics is no longer a good phrase to use. You want to use people analytics.”
Kieran Snyder Jan 2, 2019 ▶ 13:11
Assertion Not checkable as stated
Snyder: Text is the single largest output produced by any business
“In businesses, whatever your business is, text is actually the thing you produce the most of.”
Kieran Snyder Jan 2, 2019 ▶ 18:36
Assertion Not publicly verifiable
Snyder: "Fast-paced environment" in job listings statistically reduces female applicants
“So the difference between fast-paced environment and rapidly moving environment, it's almost head-scratchingly tiny, but statistically, one of them draws many fewer women to apply.”
Kieran Snyder Jan 2, 2019 ▶ 25:44
Assertion Supported
Snyder: 'Abrasive' appeared in 17 female performance reviews, zero male reviews
“The word abrasive, which has been talked about since then ended up, you know, being used in 17 out of a couple hundred women's reviews and zero times in, in men's reviews, right?”
Kieran Snyder Jan 2, 2019 ▶ 27:51
Assertion Supported
Snyder: Tech Resumes Show Systematic Gender Differences in Formatting
“Women's resumes tended to tell a story. They were written in prose. They didn't use bullets nearly as much. They included executive summaries. They included detailed statements of their Personal interests that were twice as long as what men tended to include. …”
Kieran Snyder Jan 2, 2019 ▶ 28:55
Insight
Snyder: Overused language optimization patterns lose effectiveness, forcing marketing innovation
“If everybody tries to glom onto the same patterns, they're no longer effective. Someone is gonna figure out, as with any marketer, someone is gonna figure out how to do it better, and they're gonna introduce the next Pattern for success.”
Kieran Snyder Jan 2, 2019 ▶ 31:27
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.