Feb 24, 2025 · 1h 18m · news

Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings

Steeve Morin · 57m spoken Harry Stebbings · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In an insightful interview with Harry Stebbings, ZML founder Steeve Morin critiques the inefficiencies of Nvidia's hardware monopoly and explains why the future of AI belongs to cost-effective, hardware-agnostic inference architectures and algorithmic optimization over brute-force scaling. He highlights Google as the ultimate 'sleeping giant' due to its complete ownership of products, data, and custom TPUs, while offering a realistic perspective on how physical, geopolitical, and financial constraints are driving the next generation of highly efficient AI models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 14.7% of the talking time here. How this is scored →

Harry as informed peer 3.5 Guest teaching 4.9 Guest disagreement 2.8 Harry pushing back 2.8
05100:0020:0040:001:00:000:59–3:08 · Harry as informed peer 2/10 Overview of ZML and Agnostic Infrastructure Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs.3:08–6:24 · Harry as informed peer 3/10 Multi-Hardware Usage and Cost Efficiency Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics.6:24–11:18 · Harry as informed peer 3/10 The H100 Market Bubble and Collateral Risk Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips.11:18–14:02 · Harry as informed peer 1/10 Why GPUs Are Not Built for AI Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks.14:02–16:57 · Harry as informed peer 4/10 Chip Market Structure and Profit Margins Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks.16:57–19:32 · Harry as informed peer 2/10 Infrastructure Needs: Training vs. Inference Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing.19:32–22:57 · Harry as informed peer 3/10 The Inference Auto-Scaling and Overprovisioning Problem Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments.22:57–25:00 · Harry as informed peer 4/10 The Risk of Preemptive Buying and Compute Oversupply Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters.25:00–27:46 · Harry as informed peer 6/10 Nvidia's Position in the Inference Market Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability.27:46–30:11 · Harry as informed peer 3/10 SRAM vs. HBM: Breaking Down Chip Architectures Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers.30:11–34:53 · Harry as informed peer 5/10 The Promise of Latent Space Reasoning Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end.34:53–38:02 · Harry as informed peer 3/10 The Inevitability of Inference Dominance Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff.38:02–41:34 · Harry as informed peer 4/10 AMD's GTM Challenges and Switching Barriers Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration.41:34–43:52 · Harry as informed peer 4/10 Making Switching Costs Zero: ZML's Core Mission Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk.43:52–48:31 · Harry as informed peer 4/10 The Truth About TPUs and Top-Down Integration Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput.48:31–52:34 · Harry as informed peer 5/10 What Everyone Gets Wrong About Inference Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position.52:34–55:07 · Harry as informed peer 4/10 Scaling Laws vs. Algorithmic Efficiency Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus.55:07–58:12 · Harry as informed peer 2/10 Moving Beyond Transformers: Yann LeCun's JEPA Thesis Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy.58:12–1:02:29 · Harry as informed peer 4/10 Distillation, Synthetic Data, and Small Model Advantages Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example.1:02:37–1:05:40 · Harry as informed peer 2/10 Model Efficiency, Size, and Context Windows Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks.1:05:40–1:07:59 · Harry as informed peer 4/10 The Efficiency Drive Toward Smaller Models Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering.1:07:59–1:10:02 · Harry as informed peer 4/10 Brand Power and the Future of AI Providers Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness.1:10:02–1:13:25 · Harry as informed peer 6/10 European Regulations, Mistral, and Funding Constraints Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement.1:13:25–1:17:23 · Harry as informed peer 3/10 Quick-Fire Round: Latency Reasoning and Infrastructure Shifts In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang.0:59–3:08 · Guest teaching 3/10 Overview of ZML and Agnostic Infrastructure Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs.3:08–6:24 · Guest teaching 4/10 Multi-Hardware Usage and Cost Efficiency Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics.6:24–11:18 · Guest teaching 5/10 The H100 Market Bubble and Collateral Risk Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips.11:18–14:02 · Guest teaching 6/10 Why GPUs Are Not Built for AI Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks.14:02–16:57 · Guest teaching 5/10 Chip Market Structure and Profit Margins Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks.16:57–19:32 · Guest teaching 4/10 Infrastructure Needs: Training vs. Inference Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing.19:32–22:57 · Guest teaching 4/10 The Inference Auto-Scaling and Overprovisioning Problem Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments.22:57–25:00 · Guest teaching 4/10 The Risk of Preemptive Buying and Compute Oversupply Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters.25:00–27:46 · Guest teaching 4/10 Nvidia's Position in the Inference Market Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability.27:46–30:11 · Guest teaching 6/10 SRAM vs. HBM: Breaking Down Chip Architectures Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers.30:11–34:53 · Guest teaching 6/10 The Promise of Latent Space Reasoning Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end.34:53–38:02 · Guest teaching 4/10 The Inevitability of Inference Dominance Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff.38:02–41:34 · Guest teaching 5/10 AMD's GTM Challenges and Switching Barriers Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration.41:34–43:52 · Guest teaching 4/10 Making Switching Costs Zero: ZML's Core Mission Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk.43:52–48:31 · Guest teaching 7/10 The Truth About TPUs and Top-Down Integration Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput.48:31–52:34 · Guest teaching 6/10 What Everyone Gets Wrong About Inference Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position.52:34–55:07 · Guest teaching 5/10 Scaling Laws vs. Algorithmic Efficiency Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus.55:07–58:12 · Guest teaching 6/10 Moving Beyond Transformers: Yann LeCun's JEPA Thesis Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy.58:12–1:02:29 · Guest teaching 5/10 Distillation, Synthetic Data, and Small Model Advantages Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example.1:02:37–1:05:40 · Guest teaching 5/10 Model Efficiency, Size, and Context Windows Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks.1:05:40–1:07:59 · Guest teaching 4/10 The Efficiency Drive Toward Smaller Models Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering.1:07:59–1:10:02 · Guest teaching 4/10 Brand Power and the Future of AI Providers Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness.1:10:02–1:13:25 · Guest teaching 5/10 European Regulations, Mistral, and Funding Constraints Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement.1:13:25–1:17:23 · Guest teaching 6/10 Quick-Fire Round: Latency Reasoning and Infrastructure Shifts In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang.0:59–3:08 · Guest disagreement 1/10 Overview of ZML and Agnostic Infrastructure Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs.3:08–6:24 · Guest disagreement 1/10 Multi-Hardware Usage and Cost Efficiency Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics.6:24–11:18 · Guest disagreement 2/10 The H100 Market Bubble and Collateral Risk Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips.11:18–14:02 · Guest disagreement 2/10 Why GPUs Are Not Built for AI Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks.14:02–16:57 · Guest disagreement 2/10 Chip Market Structure and Profit Margins Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks.16:57–19:32 · Guest disagreement 1/10 Infrastructure Needs: Training vs. Inference Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing.19:32–22:57 · Guest disagreement 3/10 The Inference Auto-Scaling and Overprovisioning Problem Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments.22:57–25:00 · Guest disagreement 2/10 The Risk of Preemptive Buying and Compute Oversupply Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters.25:00–27:46 · Guest disagreement 4/10 Nvidia's Position in the Inference Market Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability.27:46–30:11 · Guest disagreement 1/10 SRAM vs. HBM: Breaking Down Chip Architectures Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers.30:11–34:53 · Guest disagreement 4/10 The Promise of Latent Space Reasoning Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end.34:53–38:02 · Guest disagreement 5/10 The Inevitability of Inference Dominance Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff.38:02–41:34 · Guest disagreement 3/10 AMD's GTM Challenges and Switching Barriers Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration.41:34–43:52 · Guest disagreement 1/10 Making Switching Costs Zero: ZML's Core Mission Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk.43:52–48:31 · Guest disagreement 3/10 The Truth About TPUs and Top-Down Integration Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput.48:31–52:34 · Guest disagreement 4/10 What Everyone Gets Wrong About Inference Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position.52:34–55:07 · Guest disagreement 3/10 Scaling Laws vs. Algorithmic Efficiency Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus.55:07–58:12 · Guest disagreement 3/10 Moving Beyond Transformers: Yann LeCun's JEPA Thesis Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy.58:12–1:02:29 · Guest disagreement 3/10 Distillation, Synthetic Data, and Small Model Advantages Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example.1:02:37–1:05:40 · Guest disagreement 2/10 Model Efficiency, Size, and Context Windows Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks.1:05:40–1:07:59 · Guest disagreement 3/10 The Efficiency Drive Toward Smaller Models Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering.1:07:59–1:10:02 · Guest disagreement 3/10 Brand Power and the Future of AI Providers Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness.1:10:02–1:13:25 · Guest disagreement 6/10 European Regulations, Mistral, and Funding Constraints Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement.1:13:25–1:17:23 · Guest disagreement 4/10 Quick-Fire Round: Latency Reasoning and Infrastructure Shifts In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang.0:59–3:08 · Harry pushing back 1/10 Overview of ZML and Agnostic Infrastructure Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs.3:08–6:24 · Harry pushing back 4/10 Multi-Hardware Usage and Cost Efficiency Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics.6:24–11:18 · Harry pushing back 4/10 The H100 Market Bubble and Collateral Risk Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips.11:18–14:02 · Harry pushing back 1/10 Why GPUs Are Not Built for AI Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks.14:02–16:57 · Harry pushing back 4/10 Chip Market Structure and Profit Margins Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks.16:57–19:32 · Harry pushing back 2/10 Infrastructure Needs: Training vs. Inference Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing.19:32–22:57 · Harry pushing back 2/10 The Inference Auto-Scaling and Overprovisioning Problem Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments.22:57–25:00 · Harry pushing back 3/10 The Risk of Preemptive Buying and Compute Oversupply Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters.25:00–27:46 · Harry pushing back 5/10 Nvidia's Position in the Inference Market Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability.27:46–30:11 · Harry pushing back 2/10 SRAM vs. HBM: Breaking Down Chip Architectures Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers.30:11–34:53 · Harry pushing back 4/10 The Promise of Latent Space Reasoning Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end.34:53–38:02 · Harry pushing back 2/10 The Inevitability of Inference Dominance Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff.38:02–41:34 · Harry pushing back 4/10 AMD's GTM Challenges and Switching Barriers Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration.41:34–43:52 · Harry pushing back 4/10 Making Switching Costs Zero: ZML's Core Mission Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk.43:52–48:31 · Harry pushing back 4/10 The Truth About TPUs and Top-Down Integration Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput.48:31–52:34 · Harry pushing back 2/10 What Everyone Gets Wrong About Inference Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position.52:34–55:07 · Harry pushing back 3/10 Scaling Laws vs. Algorithmic Efficiency Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus.55:07–58:12 · Harry pushing back 1/10 Moving Beyond Transformers: Yann LeCun's JEPA Thesis Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy.58:12–1:02:29 · Harry pushing back 3/10 Distillation, Synthetic Data, and Small Model Advantages Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example.1:02:37–1:05:40 · Harry pushing back 1/10 Model Efficiency, Size, and Context Windows Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks.1:05:40–1:07:59 · Harry pushing back 2/10 The Efficiency Drive Toward Smaller Models Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering.1:07:59–1:10:02 · Harry pushing back 3/10 Brand Power and the Future of AI Providers Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness.1:10:02–1:13:25 · Harry pushing back 5/10 European Regulations, Mistral, and Funding Constraints Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement.1:13:25–1:17:23 · Harry pushing back 2/10 Quick-Fire Round: Latency Reasoning and Infrastructure Shifts In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang.

speaking balance: gold is Harry, purple is the guest (3 minute bins)

0:00 · Harry 26.1% · guest 73.9%0:00 · Harry 26.1% · guest 73.9%3:00 · Harry 15% · guest 85%3:00 · Harry 15% · guest 85%6:00 · Harry 18% · guest 82%6:00 · Harry 18% · guest 82%9:00 · Harry 8.4% · guest 91.6%9:00 · Harry 8.4% · guest 91.6%12:00 · Harry 11.8% · guest 88.2%12:00 · Harry 11.8% · guest 88.2%15:00 · Harry 12.3% · guest 87.7%15:00 · Harry 12.3% · guest 87.7%18:00 · Harry 11% · guest 89%18:00 · Harry 11% · guest 89%21:00 · Harry 16.7% · guest 83.3%21:00 · Harry 16.7% · guest 83.3%24:00 · Harry 15.6% · guest 84.4%24:00 · Harry 15.6% · guest 84.4%27:00 · Harry 7.5% · guest 92.5%27:00 · Harry 7.5% · guest 92.5%30:00 · Harry 19% · guest 81%30:00 · Harry 19% · guest 81%33:00 · Harry 11.2% · guest 88.8%33:00 · Harry 11.2% · guest 88.8%36:00 · Harry 17.5% · guest 82.5%36:00 · Harry 17.5% · guest 82.5%39:00 · Harry 12.5% · guest 87.5%39:00 · Harry 12.5% · guest 87.5%42:00 · Harry 11.6% · guest 88.4%42:00 · Harry 11.6% · guest 88.4%45:00 · Harry 12.6% · guest 87.4%45:00 · Harry 12.6% · guest 87.4%48:00 · Harry 14.8% · guest 85.2%48:00 · Harry 14.8% · guest 85.2%51:00 · Harry 18.6% · guest 81.4%51:00 · Harry 18.6% · guest 81.4%54:00 · Harry 7.1% · guest 92.9%54:00 · Harry 7.1% · guest 92.9%57:00 · Harry 20.1% · guest 79.9%57:00 · Harry 20.1% · guest 79.9%1:00:00 · Harry 15.6% · guest 84.4%1:00:00 · Harry 15.6% · guest 84.4%1:03:00 · Harry 14.6% · guest 85.4%1:03:00 · Harry 14.6% · guest 85.4%1:06:00 · Harry 15.6% · guest 84.4%1:06:00 · Harry 15.6% · guest 84.4%1:09:00 · Harry 18.2% · guest 81.8%1:09:00 · Harry 18.2% · guest 81.8%1:12:00 · Harry 10.5% · guest 89.5%1:12:00 · Harry 10.5% · guest 89.5%1:15:00 · Harry 21.5% · guest 78.5%1:15:00 · Harry 21.5% · guest 78.5%1:18:00 · Harry 0% · guest 0%1:18:00 · Harry 0% · guest 0%
Sharpest disagreement ▶ 1:11:36 Steeve mocks Stargate as an 'American car of AI'

Steeve forcefully rejects the hype surrounding the $500B Stargate project, calling it inefficient, brute-forced vertical scaling that consumes massive gas without being a good car.

Hardest push from Harry ▶ 7:01 Harry challenges Steeve on H100 bubble claims

Harry directly pushes back when Steeve claims the H100 market is a bubble, demanding to know why it isn't legitimate value given current demand.

Biggest teaching moment ▶ 45:11 Steeve breaks down Nvidia vs AMD inference memory math

Steeve delivers a concrete technical schooling on GPU RAM limits, demonstrating why 8 H100s only fit 2 models while AMD cards fit 1 model per GPU, yielding 4x throughput.

Harry holds his own ▶ 25:00 Harry pits Jensen Huang against Jonathan Ross

Harry demonstrates strong domain awareness by contrasting Jensen Huang's 40% inference revenue claim directly against Groq founder Jonathan Ross's counter-claims to test the guest.

the scores for every segment, with the reasoning behind each
ChapterTopicHarry as informed peerGuest teachingGuest disagreementHarry pushing backWhy
Overview of ZML and Agnostic Infrastructure 2311 Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs.
Multi-Hardware Usage and Cost Efficiency 3414 Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics.
The H100 Market Bubble and Collateral Risk 3524 Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips.
Why GPUs Are Not Built for AI 1621 Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks.
Chip Market Structure and Profit Margins 4524 Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks.
Infrastructure Needs: Training vs. Inference 2412 Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing.
The Inference Auto-Scaling and Overprovisioning Problem 3432 Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments.
The Risk of Preemptive Buying and Compute Oversupply 4423 Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters.
Nvidia's Position in the Inference Market 6445 Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability.
SRAM vs. HBM: Breaking Down Chip Architectures 3612 Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers.
The Promise of Latent Space Reasoning 5644 Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end.
The Inevitability of Inference Dominance 3452 Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff.
AMD's GTM Challenges and Switching Barriers 4534 Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration.
Making Switching Costs Zero: ZML's Core Mission 4414 Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk.
The Truth About TPUs and Top-Down Integration 4734 Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput.
What Everyone Gets Wrong About Inference 5642 Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position.
Scaling Laws vs. Algorithmic Efficiency 4533 Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus.
Moving Beyond Transformers: Yann LeCun's JEPA Thesis 2631 Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy.
Distillation, Synthetic Data, and Small Model Advantages 4533 Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example.
Model Efficiency, Size, and Context Windows 2521 Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks.
The Efficiency Drive Toward Smaller Models 4432 Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering.
Brand Power and the Future of AI Providers 4433 Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness.
European Regulations, Mistral, and Funding Constraints 6565 Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement.
Quick-Fire Round: Latency Reasoning and Infrastructure Shifts 3642 In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang.

Statements from this episode (57)

Opinion
Morin: Nvidia forces developers to care about unnecessary details like CUDA
“The thing with Nvidia is that they spend a lot of energy making you care about stuff you shouldn't care about, and they were very successful.”
Steeve Morin Feb 24, 2025 ▶ 36:41
Prediction Open · timeframe Feb 2030
Morin predicts AI compute will be 95% inference within five years
“In five years, I would say 95% inference, five percent training.”
Steeve Morin Feb 24, 2025 ▶ 0:15
Opinion
Morin: Google is the sleeping giant of the AI race
“Google has, like, you know, Android, Google Docs, Whatever, they have everything, they can sprinkle everywhere. This is the sleeping giant in my mind.”
Steeve Morin Feb 24, 2025 ▶ 0:24
Insight
Morin: Closed-source AI models are actually complex backend constellations
“At least if you look at close source model, they're not really models. They're more like backend, right? And there are a lot of tricks that you feel like you're talking to one model, but ultimately you're talking to a constellation, an assembly of backends tha…”
Steeve Morin Feb 24, 2025 ▶ 2:07
Prediction Not checkable as stated
Morin: Running standalone AI model weights will eventually become obsolete
“Models in the sense of, you know, getting, you know, weights and running them is something that is ultimately going away because you know, in favor of like full blown backends, right? You feel like you're talking to a model, but ultimately you're talking to an…”
Steeve Morin Feb 24, 2025 ▶ 2:41
Prediction Not checkable as stated
Morin: Future AI backend APIs will run locally in enterprise clouds
“The thing is, that API will be running locally, right? Locally, I mean, in your own, you know, cloud, you know, instances, and so on.”
Steeve Morin Feb 24, 2025 ▶ 2:59
Assertion Contradicted
Morin: Switching from Nvidia to AMD offers 4x spend efficiency
“A simple example is if you know, switch from Nvidia to AMD on a seven TB model, you can get four times better efficiency, right? In terms of spend.”
Steeve Morin Feb 24, 2025 ▶ 3:44
Opinion
Morin: Nvidia is far from the most efficient AI hardware platform
“But it's by far not the most efficient platform. And arguably, even in terms of software, it's not the best software platform.”
Steeve Morin Feb 24, 2025 ▶ 6:08
Prediction Not checkable as stated
Morin: The market bubble around Nvidia H100 GPUs will burst
“There's going to be a need for inference. Very hard to say whether it will be worth, you know, everybody's money to do it on H 100. That is a bubble that I think will blow some time.”
Steeve Morin Feb 24, 2025 ▶ 6:44
Assertion Partly supported
Morin: Nvidia H100 costs 5x A100 price for 2x inference speed
“H 100 comes along and inference is it's worth five times the price. And it may be runs twice in terms of performance on inference. That is on training. It's a lot better, but on inference, it's like maybe twice as fast when it actually, when it came out, it ra…”
Steeve Morin Feb 24, 2025 ▶ 7:20
Opinion
Morin: AI agents and reasoning will disrupt Nvidia's chip dominance
“Ultimately, the two things that could really, very much shake the industry, the chip industry, in my opinion, is our agents and reasoning.”
Steeve Morin Feb 24, 2025 ▶ 8:23
Assertion Partly supported
Morin: Nvidia Blackwell chips suffered surface bending causing cooling issues
“For Blackwell, they assembled two chips. But the surface was so big that the chip started to, you know , wave, like, I don't know the English word, but like, you know, started to bend a bit, which further perpetuated the problem because it then didn't make con…”
Steeve Morin Feb 24, 2025 ▶ 10:32
Insight
Morin: GPUs are a clever workaround, not natively built for AI
“GPUs are, you know, are a good trick for AI, but they're not built for AI.”
Steeve Morin Feb 24, 2025 ▶ 11:13
Assertion Supported
Morin: Groq and Cerebras beat GPUs via on-chip data storage
“Actually, that's why Grok achieves, ah, not Grok, but Grok, Cerebras, and all these folks, they achieve very high performance single stream is because the data is right in the chip that doesn't have to get it from memory, which is slow, which GPU has to do.”
Steeve Morin Feb 24, 2025 ▶ 12:49
Insight
Morin: Nvidia won AI training via Mellanox interconnects, not raw compute
“The reason probably Nvidia won, at least in the training space, is because of Mellanox, right? Not because of the raw compute.”
Steeve Morin Feb 24, 2025 ▶ 13:17
Assertion Partly supported
Morin details profit margins across TSMC, Nvidia, and cloud providers
“Nvidia, like a TSMC sells you at 60% margin. Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon that takes, let's say a 30% margin.”
Steeve Morin Feb 24, 2025 ▶ 15:30
Assertion Not checkable as stated
Morin: Google TPUs lack commercial success outside of Google
“They are very much successful inside of Google, but not much outside of Google, let's say, right?”
Steeve Morin Feb 24, 2025 ▶ 16:21
Insight
Morin: Interconnect dependency is the core difference between training and inference
“In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of…”
Steeve Morin Feb 24, 2025 ▶ 17:58
Assertion Not checkable as stated
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Steeve Morin Feb 24, 2025 ▶ 21:35
Insight
Morin: Cloud compute hoarding creates fake AI GPU scarcity
“So in, in the case of, you know, Amazon or Google, that would be buying reserved compute, which you're not going to use because if you buy it on demand, you will get tremendously ripped off. So that creates this like face scarcity of compute because that peopl…”
Steeve Morin Feb 24, 2025 ▶ 22:35
Assertion Supported
Morin: Nvidia Blackwell chip shipments are delayed and orders are canceled
“Blackwell is late and orders are getting canceled.”
Steeve Morin Feb 24, 2025 ▶ 23:08
Disclosure
Morin: Cold emails offering discounted GPU compute began in late 2024
“I'm getting cold emails for, you know, discounts, you know, from services I never heard about. It's, and it, and I started getting these emails probably around October, November.”
Steeve Morin Feb 24, 2025 ▶ 23:45
Prediction Open · timeframe Dec 2025
Morin: AI chip oversupply will lead to GPUs selling at 30% value
“I very much worry there will be an oversupply of these chips. The problem is, is that, you know, remember, the chips are the collateral. So, you know, somewhere, you know, in the US or whatever, there's going to be a data center with like a thousand GPUs that …”
Steeve Morin Feb 24, 2025 ▶ 24:31
Prediction Not checkable as stated
Morin: Nvidia will remain dominant in AI inference due to availability
“The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Availability that is right. So I think Nvidia is used to stay at least if not for the H-one hundre…”
Steeve Morin Feb 24, 2025 ▶ 25:33
Prediction Open · timeframe Feb 2028
Morin: Etched and Visor will bring high-speed inference chips at lower prices
“So my bet is, I think there will be, you know, chips on the market that do that at much lower price. And there's two companies I see going in that direction. One is called Etched. And the other one is called Visor.”
Steeve Morin Feb 24, 2025 ▶ 29:09
Assertion Contradicted
Morin: GPUs cannot deliver latent space AI reasoning at scale
“Fundamentally, GPUs cannot deliver, deliver this, plain and simple at scale.”
Steeve Morin Feb 24, 2025 ▶ 31:24
Insight
Morin: Scaling SRAM is a dead end for AI hardware
“No, SRAM, this will not deliver. It's a dead end in terms of scaling SRAM means scaling the surface mean you get, you know, depreciating problems.”
Steeve Morin Feb 24, 2025 ▶ 32:18
Prediction Held up
Morin: Compute-in-memory will be the next frontier in AI hardware
“So this is the next frontier, and the idea is that instead of, like, transferring the data between external memory and the CPU and do the compute there, you actually, you know, bring the CPU to the memory and you do everything. It's very, you know, it's crazy …”
Steeve Morin Feb 24, 2025 ▶ 33:48
Prediction Not checkable as stated
Morin: Nvidia may lose market dominance in AI inference and training
“I think that there's a shot that they don't.”
Steeve Morin Feb 24, 2025 ▶ 35:25
Assertion Not checkable as stated
Morin: Microsoft's AMD chip deployments made OpenAI inference profitable
“Microsoft comes along and buys it all, makes, by the way, OpenAI, or at least on the inference side, puts OpenAI in the green because of the efficiency gains.”
Steeve Morin Feb 24, 2025 ▶ 39:44
Assertion Contradicted
Morin: Apple purchased 100,000 Trainium AI chips from Amazon
“Let's take Amazon, for instance, with Tranium. Apple just came and said, Hey, we're going to buy a 100,000 of them.”
Steeve Morin Feb 24, 2025 ▶ 40:30
Insight
Morin: Being 7x better on cost won't get customers off Nvidia
“I know for a fact that being seven times better and whatever, take whatever metric you want. Whether it's spend, whether it's whatever. It's not enough to get people to switch. People will choose nothing over something.”
Steeve Morin Feb 24, 2025 ▶ 40:59
Insight
Morin: Bottom-up AI infrastructure strategies fail because developers do not care
“I think that if you are doing it bottom up, infra to applications, you will lose because nobody will care. As they don't today, right? If you look at TPUs, they're available, they're great. Nobody cares.”
Steeve Morin Feb 24, 2025 ▶ 43:37
Opinion
Morin: Google TPUs have more mature software and compute than Nvidia
“But AMD can do training, but it's also, but in terms of maturity, the, by far the most mature software and compute is TPUs, and then it's Nvidia.”
Steeve Morin Feb 24, 2025 ▶ 44:10
Assertion Contradicted
Morin: Doubling GPUs in AI inference yields only 10% performance gain
“If you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount o…”
Steeve Morin Feb 24, 2025 ▶ 45:22
Assertion Contradicted
Morin: AMD GPUs achieve 4x inference throughput over Nvidia setups
“If you run on AMD, well, there's enough memory inside the GPU to run one model per card. So you get, you know, eight GPUs, eight times the throughput, while on the other hand, you get eight GPUs, two, you know, two, maybe two and a half times the throughput. S…”
Steeve Morin Feb 24, 2025 ▶ 46:00
Assertion Supported
Morin: AMD AI chips are 30% cheaper than Nvidia GPUs
“These chips are 30% cheaper than Nvidia's.”
Steeve Morin Feb 24, 2025 ▶ 47:07
Opinion
Morin: Nvidia stock remains a better buy than AMD due to supply
“Probably I would go today, at least I would go with Nvidia still.”
Steeve Morin Feb 24, 2025 ▶ 48:08
Assertion Not checkable as stated
Morin: Nvidia gives zero discounts even on tens of thousands of GPUs
“I talked to a lot of people that build data centers and I tell them, you know, they're, mind you, these people like buy tens of thousands of GPUs. And I asked them, Hey, do you get at least a discount or something? And they're like, no. The only thing we get i…”
Steeve Morin Feb 24, 2025 ▶ 51:56
Assertion Partly supported
Morin: xAI cluster is four 25k GPU networks, not 100k unified
“The XAI cluster. It's not a 100,000 GPUs. It is four times 25,000.”
Steeve Morin Feb 24, 2025 ▶ 53:00
Insight
Morin: Non-transformer models fundamentally alter LLM compute requirements
“In the case of LLMs, for instance, you have these, what's called non transformer models that changes fundamentally the compute requirements.”
Steeve Morin Feb 24, 2025 ▶ 54:59
Opinion
Morin: Bullish on Yann LeCun's JEPA thesis over traditional LLMs
“As in LLMs are at that end, what we need is something that understands the world fundamentally, and this is the, it's JEPA thesis, it's called. I'm very bullish on this, but it's very frontier.”
Steeve Morin Feb 24, 2025 ▶ 55:58
Assertion Supported
Morin: Distilled smaller AI models can outperform their larger base models
“Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the bigger model through distillation.”
Steeve Morin Feb 24, 2025 ▶ 1:01:48
Assertion Open · timeframe Feb 2026
Morin: Google DeepMind has abandoned model fine-tuning for context windows
“You talk to people at DeepMind And they don't even fine tune anymore. Because they have such, you know, what's called big context window”
Steeve Morin Feb 24, 2025 ▶ 1:02:43
Opinion
Morin: RAG is a dirty workaround limited by context size
“It's a bit dirty because, of course, you know, you are limited by the amount of data you can input, right?”
Steeve Morin Feb 24, 2025 ▶ 1:04:51
Insight
Morin: AI developers will always choose smaller models if performance matches
“What really pushes model sizes are the efficiency rather than specializing. So meaning that if you can do the same performance with a smaller model that is fine tuned with rag or whatever, then you'll do it with a smaller because again, less is better.”
Steeve Morin Feb 24, 2025 ▶ 1:06:21
Prediction Not checkable as stated
Morin: AI model market will resemble car makers, not winner-take-all
“Is mental model in terms of model providers, ah, they'll be like car makers. Right? There's no win or tickle. Everybody will have their own.”
Steeve Morin Feb 24, 2025 ▶ 1:08:45
Opinion
Morin: DeepSeek's impact was largely overblown by media narrative and drama
“Yes, Deep Seek made a very good, you know, made waves, but it was, you know, waves that were amplified by the media and the narrative and the drama, right?”
Steeve Morin Feb 24, 2025 ▶ 1:09:05
Assertion Supported
Morin: China's domestic AI chips currently match Nvidia's A100 capability
“They're a bit late in terms of, you know, ASIC. There are like A-one-hundred level”
Steeve Morin Feb 24, 2025 ▶ 1:09:27
Prediction Not checkable as stated
Morin: Hardware bans will force China to out-innovate Western AI long-term
“They are constrained, so they are bound to Bound to do better. They can just not buy their way into better compute. So I think it hinders their success, but I think it's short term to think that way.”
Steeve Morin Feb 24, 2025 ▶ 1:09:47
Opinion
Morin: Founders should focus on success before worrying about European AI regulation
“No, I don't care. I have zero, I, this is something I makes me wonder sometimes. I understand the narrative and so on, but I am absolutely not fearful. Let's be successful first and then we'll talk about the politics.”
Steeve Morin Feb 24, 2025 ▶ 1:10:10
Opinion
Morin: Written-off narratives about Mistral AI's demise are baseless FUD
“They are very competent. So I don't know. I think it's easy to spread FUD. There's a lot of FUD going around, especially about regulation and everything. But here's the thing. I look around me and I don't see You know, what I read. I am hardly convinced about,…”
Steeve Morin Feb 24, 2025 ▶ 1:10:55
Opinion
Morin: Stargate data center is inefficient vertical scaling for AI
“It is a vertical scaling. And as you know, my days are spent on efficiency. So I look at these things as being like, all right, this is a bigger, you know, this is an American car of AI. It's big. It consumes a lot of gas, but ultimately, you know, it's not a …”
Steeve Morin Feb 24, 2025 ▶ 1:12:11
Insight
Morin: Talent and energy are the primary bottlenecks in AI
“What is ultimately the number, the, probably the two limiting factor today is talent. And energy. That's it.”
Steeve Morin Feb 24, 2025 ▶ 1:12:49
Prediction Not checkable as stated
Morin: Latency reasoning is 2025's fundamental AI infrastructure shift
“Latency reasoning. Definitely. This year. So, you know, as I was saying, like, the shift from throughput, so how speed my answers to how long it takes for my answer complete to appear. That is probably one of the fundamental, like this year, right?”
Steeve Morin Feb 24, 2025 ▶ 1:13:38
Insight
Morin: AI startups must avoid reselling compute and verticalize on product
“Probably the number one thing I would say is do not resell compute if you can. A lot of, you know, AI startups That are building on top of AI are trying to make a margin, you know, on top of a very big cake. And ultimately what they sell is compute. If you loo…”
Steeve Morin Feb 24, 2025 ▶ 1:14:21
Assertion Not checkable as stated
Morin: Nvidia intentionally smoothed H100 deliveries to prevent revenue spikes
“The supply of H 100 was actually a smooth out over the year so that they decided so that they didn't have like a big, you know, spike in deliveries and then a quarter less, right?”
Steeve Morin Feb 24, 2025 ▶ 1:16:30

Shorts cut from this episode

▶ Why NVIDIA will NOT be sustainable 🤔 · 20VC with Harry Steb (@0:00) ▶ Why Google is the biggest buy right now 💰 · 20VC with Harry (@0:20) ▶ Why you are getting SCREWED by the margin stacks of AI 😨 · (@15:34) ▶ OpenAI’s big problem 😯 · 20VC with Harry Stebbings (@0:11) ▶ Who will win the AI arms race? 😯 · 20VC with Harry Stebbing (@0:00) ▶ Why Amazon can beat NVIDIA in AI 🤯 · 20VC with Harry Stebbi (@35:37)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.