Sep 15, 2023 · 31m · 20vc

20VC: The Biggest AI Leaders on What Matters More; Model Size or Data Size & Where Does The Value in AI Accrue; to Startups or to Incumbents

Harry Stebbings · 10m spoken Richard Socher · 7m spoken Douwe Kiela · 3m spoken Alex Lebrun · 2m spoken Emad Mostaque · 1m spoken Cris Valenzuela · 1m spoken Noam Shazeer · 1m spoken Tomasz Tunguz · 1m spoken Sarah Guo · 39s spoken Clément Delangue · 36s spoken
0:00 / 0:00

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

This compilation episode of The TwentyVC Podcast brings together top artificial intelligence founders and investors to analyze whether compute scaling or data volume drives AI performance, and whether long-term value will accrue to fast-moving startups or enterprise incumbents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 35.3% of the talking time here. How this is scored →

Harry as informed peer 4.9 Guest teaching 4.6 Guest disagreement 2.9 Harry pushing back 3.6
05100:0010:0020:0030:003:19–6:25 · Harry as informed peer 4/10 Noam Shazeer on Compute and Model Scale Harry Stebbings prompts Noam Shazeer and Cris Valenzuela on compute constraints and model verticalization versus horizontal consolidation. Cris gently rejects the premise of a single model ruling the ecosystem, comparing AI tools to e-commerce diversity.6:25–13:19 · Harry as informed peer 5/10 Richard Socher on Parameters, Data Distribution, and Retrieval Backends Harry raises venture community skepticism about AI startups being thin wrappers around base models. Richard Socher provides an extensive technical masterclass on parameter scaling, world-knowledge encoding, and search retrieval backends to dismantle the thin-wrapper dismissal.13:20–16:24 · Harry as informed peer 4/10 Douwe Kiela on Data Efficiency and Optimal Compute Allocation Douwe Kiela clarifies compute optimality and sample efficiency trade-offs when training smaller models on larger data. Emad Mostaque adds provocative commentary on national datasets and algorithmic bias.16:28–20:46 · Harry as informed peer 5/10 Sarah Guo on Startup Execution Speed vs. Incumbent Data Moats Sarah Guo, Clem Delangue, and Douwe Kiela explore startup velocity versus incumbent distribution. Harry presses Douwe on why venture capitalists reject startups lacking proprietary data moats, and Douwe explains LLM data efficiency.20:51–22:52 · Harry as informed peer 4/10 Richard Socher on Distribution Battles and the Innovator's Dilemma Harry asks Richard Socher about the race between startup distribution and incumbent innovation. Richard articulates Google's Innovator's Dilemma around protecting a five-hundred-million-dollar-a-day advertising business.22:56–25:40 · Harry as informed peer 6/10 Alex LeBrun on Feature Layering vs. New Product Paradigms When Alex LeBrun argues incumbents are too slow, Harry immediately pushes back with concrete counterexamples including Adobe, Notion, and Navan. Alex concedes the point but dismisses incumbent implementations as merely sprinkling magic AI dust onto legacy products.25:47–28:08 · Harry as informed peer 6/10 Tom Tunguz on Execution Moats and Startup Agility Tomasz Tunguz discusses execution moats before Emad Mostaque claims foundation models will heavily consolidate. Harry counters Emad's skepticism about application layers by citing Tom Tunguz's quantitative market-cap distribution data.3:19–6:25 · Guest teaching 3/10 Noam Shazeer on Compute and Model Scale Harry Stebbings prompts Noam Shazeer and Cris Valenzuela on compute constraints and model verticalization versus horizontal consolidation. Cris gently rejects the premise of a single model ruling the ecosystem, comparing AI tools to e-commerce diversity.6:25–13:19 · Guest teaching 7/10 Richard Socher on Parameters, Data Distribution, and Retrieval Backends Harry raises venture community skepticism about AI startups being thin wrappers around base models. Richard Socher provides an extensive technical masterclass on parameter scaling, world-knowledge encoding, and search retrieval backends to dismantle the thin-wrapper dismissal.13:20–16:24 · Guest teaching 5/10 Douwe Kiela on Data Efficiency and Optimal Compute Allocation Douwe Kiela clarifies compute optimality and sample efficiency trade-offs when training smaller models on larger data. Emad Mostaque adds provocative commentary on national datasets and algorithmic bias.16:28–20:46 · Guest teaching 5/10 Sarah Guo on Startup Execution Speed vs. Incumbent Data Moats Sarah Guo, Clem Delangue, and Douwe Kiela explore startup velocity versus incumbent distribution. Harry presses Douwe on why venture capitalists reject startups lacking proprietary data moats, and Douwe explains LLM data efficiency.20:51–22:52 · Guest teaching 4/10 Richard Socher on Distribution Battles and the Innovator's Dilemma Harry asks Richard Socher about the race between startup distribution and incumbent innovation. Richard articulates Google's Innovator's Dilemma around protecting a five-hundred-million-dollar-a-day advertising business.22:56–25:40 · Guest teaching 4/10 Alex LeBrun on Feature Layering vs. New Product Paradigms When Alex LeBrun argues incumbents are too slow, Harry immediately pushes back with concrete counterexamples including Adobe, Notion, and Navan. Alex concedes the point but dismisses incumbent implementations as merely sprinkling magic AI dust onto legacy products.25:47–28:08 · Guest teaching 4/10 Tom Tunguz on Execution Moats and Startup Agility Tomasz Tunguz discusses execution moats before Emad Mostaque claims foundation models will heavily consolidate. Harry counters Emad's skepticism about application layers by citing Tom Tunguz's quantitative market-cap distribution data.3:19–6:25 · Guest disagreement 2/10 Noam Shazeer on Compute and Model Scale Harry Stebbings prompts Noam Shazeer and Cris Valenzuela on compute constraints and model verticalization versus horizontal consolidation. Cris gently rejects the premise of a single model ruling the ecosystem, comparing AI tools to e-commerce diversity.6:25–13:19 · Guest disagreement 2/10 Richard Socher on Parameters, Data Distribution, and Retrieval Backends Harry raises venture community skepticism about AI startups being thin wrappers around base models. Richard Socher provides an extensive technical masterclass on parameter scaling, world-knowledge encoding, and search retrieval backends to dismantle the thin-wrapper dismissal.13:20–16:24 · Guest disagreement 3/10 Douwe Kiela on Data Efficiency and Optimal Compute Allocation Douwe Kiela clarifies compute optimality and sample efficiency trade-offs when training smaller models on larger data. Emad Mostaque adds provocative commentary on national datasets and algorithmic bias.16:28–20:46 · Guest disagreement 2/10 Sarah Guo on Startup Execution Speed vs. Incumbent Data Moats Sarah Guo, Clem Delangue, and Douwe Kiela explore startup velocity versus incumbent distribution. Harry presses Douwe on why venture capitalists reject startups lacking proprietary data moats, and Douwe explains LLM data efficiency.20:51–22:52 · Guest disagreement 2/10 Richard Socher on Distribution Battles and the Innovator's Dilemma Harry asks Richard Socher about the race between startup distribution and incumbent innovation. Richard articulates Google's Innovator's Dilemma around protecting a five-hundred-million-dollar-a-day advertising business.22:56–25:40 · Guest disagreement 5/10 Alex LeBrun on Feature Layering vs. New Product Paradigms When Alex LeBrun argues incumbents are too slow, Harry immediately pushes back with concrete counterexamples including Adobe, Notion, and Navan. Alex concedes the point but dismisses incumbent implementations as merely sprinkling magic AI dust onto legacy products.25:47–28:08 · Guest disagreement 4/10 Tom Tunguz on Execution Moats and Startup Agility Tomasz Tunguz discusses execution moats before Emad Mostaque claims foundation models will heavily consolidate. Harry counters Emad's skepticism about application layers by citing Tom Tunguz's quantitative market-cap distribution data.3:19–6:25 · Harry pushing back 3/10 Noam Shazeer on Compute and Model Scale Harry Stebbings prompts Noam Shazeer and Cris Valenzuela on compute constraints and model verticalization versus horizontal consolidation. Cris gently rejects the premise of a single model ruling the ecosystem, comparing AI tools to e-commerce diversity.6:25–13:19 · Harry pushing back 4/10 Richard Socher on Parameters, Data Distribution, and Retrieval Backends Harry raises venture community skepticism about AI startups being thin wrappers around base models. Richard Socher provides an extensive technical masterclass on parameter scaling, world-knowledge encoding, and search retrieval backends to dismantle the thin-wrapper dismissal.13:20–16:24 · Harry pushing back 2/10 Douwe Kiela on Data Efficiency and Optimal Compute Allocation Douwe Kiela clarifies compute optimality and sample efficiency trade-offs when training smaller models on larger data. Emad Mostaque adds provocative commentary on national datasets and algorithmic bias.16:28–20:46 · Harry pushing back 3/10 Sarah Guo on Startup Execution Speed vs. Incumbent Data Moats Sarah Guo, Clem Delangue, and Douwe Kiela explore startup velocity versus incumbent distribution. Harry presses Douwe on why venture capitalists reject startups lacking proprietary data moats, and Douwe explains LLM data efficiency.20:51–22:52 · Harry pushing back 2/10 Richard Socher on Distribution Battles and the Innovator's Dilemma Harry asks Richard Socher about the race between startup distribution and incumbent innovation. Richard articulates Google's Innovator's Dilemma around protecting a five-hundred-million-dollar-a-day advertising business.22:56–25:40 · Harry pushing back 6/10 Alex LeBrun on Feature Layering vs. New Product Paradigms When Alex LeBrun argues incumbents are too slow, Harry immediately pushes back with concrete counterexamples including Adobe, Notion, and Navan. Alex concedes the point but dismisses incumbent implementations as merely sprinkling magic AI dust onto legacy products.25:47–28:08 · Harry pushing back 5/10 Tom Tunguz on Execution Moats and Startup Agility Tomasz Tunguz discusses execution moats before Emad Mostaque claims foundation models will heavily consolidate. Harry counters Emad's skepticism about application layers by citing Tom Tunguz's quantitative market-cap distribution data.

speaking balance: gold is Harry, purple is the guest (3 minute bins)

0:00 · Harry 100% · guest 0%0:00 · Harry 100% · guest 0%3:00 · Harry 31.9% · guest 68.1%3:00 · Harry 31.9% · guest 68.1%6:00 · Harry 21% · guest 79%6:00 · Harry 21% · guest 79%9:00 · Harry 11.9% · guest 88.1%9:00 · Harry 11.9% · guest 88.1%12:00 · Harry 11.4% · guest 88.6%12:00 · Harry 11.4% · guest 88.6%15:00 · Harry 19.1% · guest 80.9%15:00 · Harry 19.1% · guest 80.9%18:00 · Harry 20.2% · guest 79.8%18:00 · Harry 20.2% · guest 79.8%21:00 · Harry 17% · guest 83%21:00 · Harry 17% · guest 83%24:00 · Harry 14.7% · guest 85.3%24:00 · Harry 14.7% · guest 85.3%27:00 · Harry 76.2% · guest 23.8%27:00 · Harry 76.2% · guest 23.8%30:00 · Harry 100% · guest 0%30:00 · Harry 100% · guest 0%
Sharpest disagreement ▶ 23:39 Alex LeBrun dismisses incumbent AI features as superficial magic dust

Alex forcefully rejects the idea that incumbent feature additions represent true disruption, arguing that companies like Notion and Google Docs will eventually be destroyed by entirely new paradigms.

Hardest push from Harry ▶ 23:25 Harry refuses the premise that incumbents are moving slowly

Harry directly interrupts and challenges Alex LeBrun's assertion that incumbents are universally slow by presenting specific, fast-moving companies like Notion, Adobe, and Navan.

Biggest teaching moment ▶ 6:39 Richard Socher explains parameter scale and world knowledge encoding

Richard uses a vivid linear regression comparison and next-token prediction example across geography to educate the host on why billions of parameters are fundamentally required to capture nuanced knowledge.

Harry holds his own ▶ 27:17 Harry cites Tom Tunguz data to challenge Emad's thin layer claim

Harry counters Emad Mostaque's dismissal of application-layer value by citing specific data from Tom Tunguz showing that application layers capture trillions of dollars across diverse companies.

the scores for every segment, with the reasoning behind each
ChapterTopicHarry as informed peerGuest teachingGuest disagreementHarry pushing backWhy
Noam Shazeer on Compute and Model Scale 4323 Harry Stebbings prompts Noam Shazeer and Cris Valenzuela on compute constraints and model verticalization versus horizontal consolidation. Cris gently rejects the premise of a single model ruling the ecosystem, comparing AI tools to e-commerce diversity.
Richard Socher on Parameters, Data Distribution, and Retrieval Backends 5724 Harry raises venture community skepticism about AI startups being thin wrappers around base models. Richard Socher provides an extensive technical masterclass on parameter scaling, world-knowledge encoding, and search retrieval backends to dismantle the thin-wrapper dismissal.
Douwe Kiela on Data Efficiency and Optimal Compute Allocation 4532 Douwe Kiela clarifies compute optimality and sample efficiency trade-offs when training smaller models on larger data. Emad Mostaque adds provocative commentary on national datasets and algorithmic bias.
Sarah Guo on Startup Execution Speed vs. Incumbent Data Moats 5523 Sarah Guo, Clem Delangue, and Douwe Kiela explore startup velocity versus incumbent distribution. Harry presses Douwe on why venture capitalists reject startups lacking proprietary data moats, and Douwe explains LLM data efficiency.
Richard Socher on Distribution Battles and the Innovator's Dilemma 4422 Harry asks Richard Socher about the race between startup distribution and incumbent innovation. Richard articulates Google's Innovator's Dilemma around protecting a five-hundred-million-dollar-a-day advertising business.
Alex LeBrun on Feature Layering vs. New Product Paradigms 6456 When Alex LeBrun argues incumbents are too slow, Harry immediately pushes back with concrete counterexamples including Adobe, Notion, and Navan. Alex concedes the point but dismisses incumbent implementations as merely sprinkling magic AI dust onto legacy products.
Tom Tunguz on Execution Moats and Startup Agility 6445 Tomasz Tunguz discusses execution moats before Emad Mostaque claims foundation models will heavily consolidate. Harry counters Emad's skepticism about application layers by citing Tom Tunguz's quantitative market-cap distribution data.

Statements from this episode (25)

Insight
Noam Shazeer: Compute and Model Size Constrain AI Progress More Than Data
“Yeah, probably the size of the model is the bigger challenge. We can get a lot of data, but actually the number one thing that's important is how much computation you do to train it.”
Noam Shazeer Sep 15, 2023 ▶ 3:32
Disclosure
Character.ai Spent $2 Million in Compute to Train Its Serving Model
“The model we're serving now, we trained last summer and spent about two million dollars worth of compute cycles doing it.”
Noam Shazeer Sep 15, 2023 ▶ 4:12
Prediction Not checkable as stated
Chris Valenzuela: No Single Monolithic AI Model Will Dominate the Market
“I don't think there's going to be a single model to rule them all.”
Cris Valenzuela Sep 15, 2023 ▶ 5:30
Insight
Chris Valenzuela: AI Models Are Not Moats; Talent and Speed Win
“I think models are not a mode. Models eventually don't matter. What matters most is the people building those models. And how fast can you change and learn from those models?”
Cris Valenzuela Sep 15, 2023 ▶ 6:11
Insight
Richard Socher: Enterprise Incumbents Hold a Key Advantage in Proprietary Data
“So it's complicated in the sense that unsupervised data, just raw internet text is easily accessible, but there's still a lot of data sets out there that are not out there. They're actually stored in a private databases. And indeed, if you want to answer custo…”
Richard Socher Sep 15, 2023 ▶ 9:08
Insight
Richard Socher: AI Startups Can Build Moats Through Distribution and Partnerships
“Turns out you can have moat other than your backend AI model, right? It's distribution, it's partnerships. Your sales funnel processes and so on.”
Richard Socher Sep 15, 2023 ▶ 11:50
Prediction Not checkable as stated
Douwe Kiela: AI Model Parameter Sizes Will Stop Growing Rapidly
“I think Sam Altman had this interesting quote where he was saying that he thought models would stop growing in size. GPT-IV kind of hit this ceiling. I think that's probably right, but not really because size doesn't matter. It's just that data size matters ev…”
Douwe Kiela Sep 15, 2023 ▶ 13:29
Insight
Douwe Kiela: Training Smaller AI Models on More Data Increases Efficiency
“If you train a smaller model on more data for longer, then you get a better model. So you get more bang for your buck if you train it on more data rather than having more parameters.”
Douwe Kiela Sep 15, 2023 ▶ 13:50
Disclosure
Emad Mostaque: Why I Signed the Open Letter for an AI Pause
“The reason I signed that letter, because I think there's a six month pause to get all of our shit together before things go completely insane, and next year this is everywhere, and everyone's investing in everything, and it's just absolute chaos.”
Emad Mostaque Sep 15, 2023 ▶ 15:08
Prediction Not checkable as stated
Emad Mostaque: Every Nation Will Need Sovereign Datasets and Open Models
“As part of that, every nation will need their own data sets, which again, have from broadcaster data. They will need their own open models that can stimulate innovation internally as well.”
Emad Mostaque Sep 15, 2023 ▶ 15:35
Insight
Emad Mostaque: Unbiased AI Models Do Not Exist
“There is no such thing as an unbiased model.”
Emad Mostaque Sep 15, 2023 ▶ 15:48
Prediction Didn’t hold up
Emad Mostaque: No Current AI Models Will Be Used in a Year
“The reality is no models that are out today will be used in a year.”
Emad Mostaque Sep 15, 2023 ▶ 16:21
Insight
Sarah Guo: Speed Is the Only Real Advantage Startups Have
“Classically, the only real advantage startups have is speed, and speed actually might matter more than ever when the environment seems to be moving at warp speed”
Sarah Guo Sep 15, 2023 ▶ 16:53
Opinion
Sarah Guo: Incumbent Data Moats in AI Are Overhyped
“On the incumbent advantage side, much ado has been made about this idea of a data moat. But honestly, there's a lot of data out there and entrepreneurs are incredibly creative about collecting it and increasingly about generating it.”
Sarah Guo Sep 15, 2023 ▶ 17:12
Opinion
Clément Delangue: Deep AI Startups Can Execute 100x Better Than Incumbents
“I think if you were thinking about AI as AI APIs, I agree. If you're thinking of AI as a more radical paradigm switch to build technology, right? And if you think about an AI startup as a company that is actually training models, creating new architectures, op…”
Clément Delangue Sep 15, 2023 ▶ 17:36
Assertion Supported
Douwe Kiela: Meta's Original LLaMA Was Trained Entirely on Open Data
“So the LAMA model was not trained on any proprietary data. It was just trained on open data on the web.”
Douwe Kiela Sep 15, 2023 ▶ 18:27
Insight
Douwe Kiela: Deep Tech AI Startups Must Build Data Flywheels
“If you want to build a deep tech AI startup, then you really want to get a big data flywheel going. You want to start with a lot of data and then have a way to generate lots more data and that data is going to be your moat.”
Douwe Kiela Sep 15, 2023 ▶ 19:34
Prediction Not checkable as stated
Douwe Kiela: GPT-4 Will Disrupt Data Annotators Like Mechanical Turk
“One of the use cases I've been seeing now for GPT-IV is actually that people are using it to generate data and then they're training on that data with cheaper models. So GPT-IV might end up disrupting, not knowledge workers necessarily, but it might just disru…”
Douwe Kiela Sep 15, 2023 ▶ 20:19
Assertion Partly supported
Richard Socher: Bing and Google Copied YouChat Months After Its Launch
“We've also seen Bing and Google copy what we have launched late last year with you chat. They've copied us in the sense that we've launched it earlier and then they launched something very similar three to four months after.”
Richard Socher Sep 15, 2023 ▶ 21:29
Insight
Richard Socher: Google's Ad Revenue Creates an Innovator's Dilemma Against AI
“That kind of big change will be hard for Google too, because they make five hundred million dollars a day with privacy invading advertisements on that page. And so you don't just willy nilly change most of that page and you get rid of the five, six ads that ar…”
Richard Socher Sep 15, 2023 ▶ 21:50
Prediction Open · timeframe Sep 2028
Alex LeBrun: AI Paradigms Will Eventually Displace Notion and Google Docs
“Probably something will Come and destroy Notion and Google Docs and all of them with a totally new paradigm that is made possible by AI.”
Alex LeBrun Sep 15, 2023 ▶ 24:12
Assertion Not checkable as stated
Alex LeBrun: LLM Capabilities Could Not Be Predicted Before Scaling
“The reason is nobody could predict that LLMs would be so useful and powerful before you try to run at this scale.”
Alex LeBrun Sep 15, 2023 ▶ 25:01
Prediction Open · timeframe Sep 2028
Emad Mostaque: Only Five Foundation Model Companies Will Survive by 2028
“I think that there's only going to be five or six foundation model companies in the world in three years, five years.”
Emad Mostaque Sep 15, 2023 ▶ 27:37
Assertion Not checkable as stated
Emad Mostaque: Google Spends $20 Billion per Year on AI
“Google spend twenty billion dollars a year on AI.”
Emad Mostaque Sep 15, 2023 ▶ 28:03
Assertion Partly supported
Emad Mostaque: DeepMind's Annual Salary Budget Is $1.2 Billion
“DeepMind's salary budget is 1.2 billion a year.”
Emad Mostaque Sep 15, 2023 ▶ 28:06
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.