Feb 24, 2025 · 1h 18m · news
Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262 · 20VC with Harry Stebbings
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In an insightful interview with Harry Stebbings, ZML founder Steeve Morin critiques the inefficiencies of Nvidia's hardware monopoly and explains why the future of AI belongs to cost-effective, hardware-agnostic inference architectures and algorithmic optimization over brute-force scaling. He highlights Google as the ultimate 'sleeping giant' due to its complete ownership of products, data, and custom TPUs, while offering a realistic perspective on how physical, geopolitical, and financial constraints are driving the next generation of highly efficient AI models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 14.7% of the talking time here. How this is scored →
speaking balance: gold is Harry, purple is the guest (3 minute bins)
Steeve forcefully rejects the hype surrounding the $500B Stargate project, calling it inefficient, brute-forced vertical scaling that consumes massive gas without being a good car.
Hardest push from Harry ▶ 7:01 Harry challenges Steeve on H100 bubble claimsHarry directly pushes back when Steeve claims the H100 market is a bubble, demanding to know why it isn't legitimate value given current demand.
Biggest teaching moment ▶ 45:11 Steeve breaks down Nvidia vs AMD inference memory mathSteeve delivers a concrete technical schooling on GPU RAM limits, demonstrating why 8 H100s only fit 2 models while AMD cards fit 1 model per GPU, yielding 4x throughput.
Harry holds his own ▶ 25:00 Harry pits Jensen Huang against Jonathan RossHarry demonstrates strong domain awareness by contrasting Jensen Huang's 40% inference revenue claim directly against Groq founder Jonathan Ross's counter-claims to test the guest.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Harry as informed peer | Guest teaching | Guest disagreement | Harry pushing back | Why |
|---|---|---|---|---|---|---|
| Overview of ZML and Agnostic Infrastructure | 2 | 3 | 1 | 1 | Harry opens with foundational questions asking for an overview of ZML and multi-model backend usage. Steeve explains that models are increasingly abstractions for constellations of backend APIs. | |
| Multi-Hardware Usage and Cost Efficiency | 3 | 4 | 1 | 4 | Harry pushes Steeve on why companies do not switch to AMD if it offers a 4x cost efficiency. Steeve explains PyTorch and CUDA ecosystem lock-in alongside supply resale dynamics. | |
| The H100 Market Bubble and Collateral Risk | 3 | 5 | 2 | 4 | Harry asks why the H100 market might be a bubble, challenging Steeve's skepticism. Steeve details financial amortization risks and the physical limits of Blackwell chips. | |
| Why GPUs Are Not Built for AI | 1 | 6 | 2 | 1 | Harry self-deprecatingly asks basic questions about why GPUs are not natively built for AI. Steeve provides a detailed historical breakdown of GPGPU tricks and interconnect bottlenecks. | |
| Chip Market Structure and Profit Margins | 4 | 5 | 2 | 4 | Harry probes profit margin dynamics across chip providers, suggesting competition should lower margins. Steeve breaks down TSMC, Nvidia, and cloud provider margin stacks. | |
| Infrastructure Needs: Training vs. Inference | 2 | 4 | 1 | 2 | Harry asks for clarification on why training needs 'more is more' while inference favors 'less is better'. Steeve uses an intuitive analogy comparing single fine artwork to mass manufacturing. | |
| The Inference Auto-Scaling and Overprovisioning Problem | 3 | 4 | 3 | 2 | Harry synthesizes Steeve's points on duct-taped production inference and overprovisioning. Steeve criticizes the lack of proper auto-scaling in current AI deployments. | |
| The Risk of Preemptive Buying and Compute Oversupply | 4 | 4 | 2 | 3 | Harry asks forward-looking questions regarding hardware obsolescence and compute oversupply overhang. Steeve confirms seeing cold emails offering heavily discounted GPU clusters. | |
| Nvidia's Position in the Inference Market | 6 | 4 | 4 | 5 | Harry pits statements from Jensen Huang against Jonathan Ross from Groq to test Nvidia's inference position. Steeve rejects the premise of Ross's extreme claim based on real-world chip availability. | |
| SRAM vs. HBM: Breaking Down Chip Architectures | 3 | 6 | 1 | 2 | Steeve breaks down the technical differences between SRAM and HBM memory architectures using numerical model sizing. Harry asks about unit cost levers. | |
| The Promise of Latent Space Reasoning | 5 | 6 | 4 | 4 | Harry cites Groq's founder regarding HBM dominance, prompting Steeve to explain latent space reasoning and dismiss SRAM scaling as a dead end. | |
| The Inevitability of Inference Dominance | 3 | 4 | 5 | 2 | Harry asks Steeve to estimate the future value split between training and inference. Steeve aggressively dismisses Nvidia's CUDA moat as artificial marketing fluff. | |
| AMD's GTM Challenges and Switching Barriers | 4 | 5 | 3 | 4 | Harry shares his personal stock performance contrast between Nvidia and AMD while asking about switching barriers. Steeve highlights that even 7x efficiency gains fail to trigger enterprise migration. | |
| Making Switching Costs Zero: ZML's Core Mission | 4 | 4 | 1 | 4 | Harry asks how ZML achieves zero buy-in switching and specifically pushes Steeve on who carries the underlying hardware risk. | |
| The Truth About TPUs and Top-Down Integration | 4 | 7 | 3 | 4 | Harry asks Steeve to explain the exact mechanics of how Microsoft avoids losing money on inference via AMD. Steeve delivers a precise mathematical breakdown of GPU RAM and throughput. | |
| What Everyone Gets Wrong About Inference | 5 | 6 | 4 | 2 | Harry cites specific Capex figures for Meta and Microsoft. Steeve lays out his Product, Data, Compute triangle framework, explaining why Google holds the dominant position. | |
| Scaling Laws vs. Algorithmic Efficiency | 4 | 5 | 3 | 3 | Harry asks when the industry will pivot from brute-force scaling to algorithmic efficiency. Steeve explains physical networking limits in superclusters like xAI's Colossus. | |
| Moving Beyond Transformers: Yann LeCun's JEPA Thesis | 2 | 6 | 3 | 1 | Steeve details Yann LeCun's JEPA thesis and energy-based world models. Harry asks why Steeve finds it compelling, leading to an intuitive physical energy analogy. | |
| Distillation, Synthetic Data, and Small Model Advantages | 4 | 5 | 3 | 3 | Harry asks whether model distillation is ethically wrong or just open source in a new form. Steeve dismisses complaints against distillation using a Star Wars training data example. | |
| Model Efficiency, Size, and Context Windows | 2 | 5 | 2 | 1 | Harry asks for a baseline explanation of Retrieval-Augmented Generation (RAG). Steeve explains vector embeddings while characterizing current RAG implementations as clever but dirty tricks. | |
| The Efficiency Drive Toward Smaller Models | 4 | 4 | 3 | 2 | Harry prompts Steeve on Chinese AI innovation following DeepSeek. Steeve argues that hardware export constraints served as the driver for superior efficiency engineering. | |
| Brand Power and the Future of AI Providers | 4 | 4 | 3 | 3 | Harry inquires whether DeepSeek represents a existential threat to ChatGPT. Steeve cites consumer brand moats using a personal anecdote about his mother's awareness. | |
| European Regulations, Mistral, and Funding Constraints | 6 | 5 | 6 | 5 | Harry directly confronts Steeve with rumors that Mistral lacks capital to compete. Steeve defends European competitiveness and vehemently mocks the $500B Stargate announcement. | |
| Quick-Fire Round: Latency Reasoning and Infrastructure Shifts | 3 | 6 | 4 | 2 | In the quick-fire round, Steeve provides insider details regarding Blackwell silicon design defects, characterising its thermal and bending issues as a massive roadblock for Jensen Huang. |