Jan 2, 2019 · 21m · a16z

a16z Podcast | The Product Edge in Machine Learning Startups

Jensen Harris · 7m spoken AJ Shankar · 6m spoken Steven Sinofsky · 4m spoken Sonal Chokshi · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Sonal Chokshi and moderator Steven Sinofsky discuss how machine learning startups can establish a competitive edge with Textio co-founder Jensen Harris and Everlaw co-founder AJ Shankar. The guests explain how focusing on domain-specific data quality, full-stack enterprise software, and human-AI collaboration enables startups to beat tech giants.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 6.1% of the talking time here. How this is scored →

The host as informed peer 4.8 Guest teaching 4.2 Guest disagreement 1.2 The host pushing back 2.6
05100:0010:0020:001:17–4:59 · The host as informed peer 4/10 Why Niche Machine Learning Startups Can Defeat Tech Giants Steven asks why big tech platforms with existing data advantages do not monopolize e-discovery, prompting AJ to reframe the problem around domain-specific isolation and statistical ML. Steven demonstrates industry familiarity by summarizing architectural choices like launching Torch VMs, while AJ corrects common assumptions about neural networks.4:59–8:07 · The host as informed peer 4/10 Prioritizing Domain Data Quality Over Mass Dataset Volume Jensen and AJ explain why domain-tuned, high-quality data matters more than massive dataset volume, contrasting academic benchmarks with actual customer problems. Steven framing the discussion around 'navigating the idea maze' shows solid venture context while keeping the dynamic collaborative.8:07–11:07 · The host as informed peer 5/10 Building Full-Stack Enterprise Products Around Machine Learning Engines Jensen details Textio's full-stack enterprise requirements, including a custom editor achieving 300 millisecond latency. Steven actively validates the operational realities of SaaS product development, keeping the discussion focused on product-market fit.11:07–15:40 · The host as informed peer 6/10 Navigating Machine Learning Hype and Optimizing Data Cleaning Pipelines Steven explicitly challenges AJ when he gives a vague response about open-source tools, demanding specific details. AJ and Jensen elaborate on data-cleaning pipelines, email deduplication signal loss, and AWS Athena usage.15:40–21:12 · The host as informed peer 5/10 Leveraging Modern Cloud Infrastructure for Rapid Machine Learning Iteration Steven raises a realistic counter-argument about LinkedIn holding all job description data, prompting Jensen to explain why curated outcome data beats raw volume. AJ concludes with key insights on user interface trust and explainability in legal AI.1:17–4:59 · Guest teaching 5/10 Why Niche Machine Learning Startups Can Defeat Tech Giants Steven asks why big tech platforms with existing data advantages do not monopolize e-discovery, prompting AJ to reframe the problem around domain-specific isolation and statistical ML. Steven demonstrates industry familiarity by summarizing architectural choices like launching Torch VMs, while AJ corrects common assumptions about neural networks.4:59–8:07 · Guest teaching 4/10 Prioritizing Domain Data Quality Over Mass Dataset Volume Jensen and AJ explain why domain-tuned, high-quality data matters more than massive dataset volume, contrasting academic benchmarks with actual customer problems. Steven framing the discussion around 'navigating the idea maze' shows solid venture context while keeping the dynamic collaborative.8:07–11:07 · Guest teaching 3/10 Building Full-Stack Enterprise Products Around Machine Learning Engines Jensen details Textio's full-stack enterprise requirements, including a custom editor achieving 300 millisecond latency. Steven actively validates the operational realities of SaaS product development, keeping the discussion focused on product-market fit.11:07–15:40 · Guest teaching 5/10 Navigating Machine Learning Hype and Optimizing Data Cleaning Pipelines Steven explicitly challenges AJ when he gives a vague response about open-source tools, demanding specific details. AJ and Jensen elaborate on data-cleaning pipelines, email deduplication signal loss, and AWS Athena usage.15:40–21:12 · Guest teaching 4/10 Leveraging Modern Cloud Infrastructure for Rapid Machine Learning Iteration Steven raises a realistic counter-argument about LinkedIn holding all job description data, prompting Jensen to explain why curated outcome data beats raw volume. AJ concludes with key insights on user interface trust and explainability in legal AI.1:17–4:59 · Guest disagreement 1/10 Why Niche Machine Learning Startups Can Defeat Tech Giants Steven asks why big tech platforms with existing data advantages do not monopolize e-discovery, prompting AJ to reframe the problem around domain-specific isolation and statistical ML. Steven demonstrates industry familiarity by summarizing architectural choices like launching Torch VMs, while AJ corrects common assumptions about neural networks.4:59–8:07 · Guest disagreement 1/10 Prioritizing Domain Data Quality Over Mass Dataset Volume Jensen and AJ explain why domain-tuned, high-quality data matters more than massive dataset volume, contrasting academic benchmarks with actual customer problems. Steven framing the discussion around 'navigating the idea maze' shows solid venture context while keeping the dynamic collaborative.8:07–11:07 · Guest disagreement 1/10 Building Full-Stack Enterprise Products Around Machine Learning Engines Jensen details Textio's full-stack enterprise requirements, including a custom editor achieving 300 millisecond latency. Steven actively validates the operational realities of SaaS product development, keeping the discussion focused on product-market fit.11:07–15:40 · Guest disagreement 2/10 Navigating Machine Learning Hype and Optimizing Data Cleaning Pipelines Steven explicitly challenges AJ when he gives a vague response about open-source tools, demanding specific details. AJ and Jensen elaborate on data-cleaning pipelines, email deduplication signal loss, and AWS Athena usage.15:40–21:12 · Guest disagreement 1/10 Leveraging Modern Cloud Infrastructure for Rapid Machine Learning Iteration Steven raises a realistic counter-argument about LinkedIn holding all job description data, prompting Jensen to explain why curated outcome data beats raw volume. AJ concludes with key insights on user interface trust and explainability in legal AI.1:17–4:59 · The host pushing back 2/10 Why Niche Machine Learning Startups Can Defeat Tech Giants Steven asks why big tech platforms with existing data advantages do not monopolize e-discovery, prompting AJ to reframe the problem around domain-specific isolation and statistical ML. Steven demonstrates industry familiarity by summarizing architectural choices like launching Torch VMs, while AJ corrects common assumptions about neural networks.4:59–8:07 · The host pushing back 2/10 Prioritizing Domain Data Quality Over Mass Dataset Volume Jensen and AJ explain why domain-tuned, high-quality data matters more than massive dataset volume, contrasting academic benchmarks with actual customer problems. Steven framing the discussion around 'navigating the idea maze' shows solid venture context while keeping the dynamic collaborative.8:07–11:07 · The host pushing back 1/10 Building Full-Stack Enterprise Products Around Machine Learning Engines Jensen details Textio's full-stack enterprise requirements, including a custom editor achieving 300 millisecond latency. Steven actively validates the operational realities of SaaS product development, keeping the discussion focused on product-market fit.11:07–15:40 · The host pushing back 5/10 Navigating Machine Learning Hype and Optimizing Data Cleaning Pipelines Steven explicitly challenges AJ when he gives a vague response about open-source tools, demanding specific details. AJ and Jensen elaborate on data-cleaning pipelines, email deduplication signal loss, and AWS Athena usage.15:40–21:12 · The host pushing back 3/10 Leveraging Modern Cloud Infrastructure for Rapid Machine Learning Iteration Steven raises a realistic counter-argument about LinkedIn holding all job description data, prompting Jensen to explain why curated outcome data beats raw volume. AJ concludes with key insights on user interface trust and explainability in legal AI.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 42.8% · guest 57.2%0:00 · the host 42.8% · guest 57.2%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 18:51 Rejecting LinkedIn's Scale Advantage

Jensen directly pushes back against Steven's counter-argument that LinkedIn's scale controls job descriptions, arguing that small curated outcome data is far more valuable.

Hardest push from the host ▶ 11:55 Calling Out Generic Answer

Steven explicitly rejects AJ's evasive answer about tech stacks ('That's a terrible answer... didn't offer any specifics'), forcing him to detail their actual data cleaning tools.

Biggest teaching moment ▶ 3:49 Demystifying Deep Learning vs Regression

AJ corrects the mainstream hype around neural networks, explaining to the host and audience that practical ML often relies on statistical regression.

The host holds their own ▶ 4:28 Synthesizing ML Architecture Strategy

Steven demonstrates technical grasp by translating AJ's points into actionable advice against prematurely spinning up Torch VMs before understanding the domain problem.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Why Niche Machine Learning Startups Can Defeat Tech Giants 4512 Steven asks why big tech platforms with existing data advantages do not monopolize e-discovery, prompting AJ to reframe the problem around domain-specific isolation and statistical ML. Steven demonstrates industry familiarity by summarizing architectural choices like launching Torch VMs, while AJ corrects common assumptions about neural networks.
Prioritizing Domain Data Quality Over Mass Dataset Volume 4412 Jensen and AJ explain why domain-tuned, high-quality data matters more than massive dataset volume, contrasting academic benchmarks with actual customer problems. Steven framing the discussion around 'navigating the idea maze' shows solid venture context while keeping the dynamic collaborative.
Building Full-Stack Enterprise Products Around Machine Learning Engines 5311 Jensen details Textio's full-stack enterprise requirements, including a custom editor achieving 300 millisecond latency. Steven actively validates the operational realities of SaaS product development, keeping the discussion focused on product-market fit.
Navigating Machine Learning Hype and Optimizing Data Cleaning Pipelines 6525 Steven explicitly challenges AJ when he gives a vague response about open-source tools, demanding specific details. AJ and Jensen elaborate on data-cleaning pipelines, email deduplication signal loss, and AWS Athena usage.
Leveraging Modern Cloud Infrastructure for Rapid Machine Learning Iteration 5413 Steven raises a realistic counter-argument about LinkedIn holding all job description data, prompting Jensen to explain why curated outcome data beats raw volume. AJ concludes with key insights on user interface trust and explainability in legal AI.

Statements from this episode (9)

Assertion Not checkable as stated
AJ Shankar: Most machine learning implementations do not use deep learning
“Most ML implementations are not deep learning. They're not neural networks. There's a ton of implementations that are incredibly valuable that don't involve networks.”
AJ Shankar Jan 2, 2019 ▶ 3:58
Assertion Supported
AJ Shankar: IBM Watson is mostly regression, not an all-solving brain
“What IBM calls Watson, which is not some brain that just solves every problem, but a whole agglomeration of different techniques is largely regression, which is an incredibly powerful technique.”
AJ Shankar Jan 2, 2019 ▶ 4:04
Prediction Not checkable as stated
Jensen Harris: Machine learning algorithms will rapidly become commodities
“I think all of the algorithmic stuff in machine learning is going to be commodity. Like, there are, like, 20 places in the world where they're inventing new algorithms, and that's, you know, educational institutions and huge companies, and that's really import…”
Jensen Harris Jan 2, 2019 ▶ 6:08
Insight
Jensen Harris: Startups beat Google and Microsoft on tailored security policies
“A huge advantage that a startup has that some big company trying to do the same thing doesn't have, which is we can tailor our security policies, tailor the way that we handle the data, the way that we sanitize the data, and the way that we use the data in a v…”
Jensen Harris Jan 2, 2019 ▶ 9:14
Insight
Jensen Harris: Early ML startups need thousands, not billions, of data points
“We found in our sort of earliest days, like our first six months that we didn't have to, you know, really ingest tens of billions of things then. What we really needed was, you know, tens of thousands or hundreds of thousands of really good pieces of data.”
Jensen Harris Jan 2, 2019 ▶ 10:01
Opinion
AJ Shankar: ML startups hype algorithms for fundraising, not customer value
“Typically when companies are hyping up their machine learning component, they're, Might be doing it more to raise money than they are to provide a value to the customer.”
AJ Shankar Jan 2, 2019 ▶ 11:16
Disclosure
Jensen Harris: Textio uses off-the-shelf Python ML libraries, focusing on data pipelines
“Our actual, you know, machine learning algorithms and the core NLP stuff we do is the standard sort of Python libraries that You know, you can go download and use, but we have put an enormous amount of time into our data processing pipeline.”
Jensen Harris Jan 2, 2019 ▶ 14:04
Insight
Jensen Harris: Early ML startups should use cloud providers, not custom infrastructure
“Yeah, you don't need to custom build something, and you shouldn't spend any of your time working on that. Like, you should figure out what cloud platform you're using, whether it's AWS, or whether it's Azure, or something else. They all have built-in ML servic…”
Jensen Harris Jan 2, 2019 ▶ 16:34
Insight
AJ Shankar: AI products must act as transparent partners, not black boxes
“Yeah, I mean, the key thing you want to do is, is present the AI as a partner in the human, in the humans, in the people's endeavors, you know, what they're trying to do. This is something that's going to help you, you're going to work with it, and when you do…”
AJ Shankar Jan 2, 2019 ▶ 20:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.