Mar 24, 2025 · 1h 14m · 20vc
Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this 20VC podcast interview, Cerebras CEO Andrew Feldman details how his company's revolutionary wafer-scale chip architecture overcomes traditional GPU bottlenecks to challenge Nvidia's dominance, while analyzing the geopolitical, economic, and technical realities of AI's shift toward utility.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 16.1% of the talking time here. How this is scored →
speaking balance: gold is Harry, purple is the guest (3 minute bins)
Feldman flatly rejects Stebbings' premise that the industry is far along across compute, algorithms, and data, stating 'I think they're wrong' and insisting early-stage progress remains across all pillars.
Hardest push from Harry ▶ 51:18 Challenging IPO timing and S-1 disclosure risksStebbings explicitly challenges Feldman on filing an S-1 public disclosure, calling it preemptive and pointing out that it grants competitors asymmetric financial and strategic information.
Biggest teaching moment ▶ 42:00 Debunking CUDA lock-in in AI inferenceFeldman directly corrects popular narrative around Nvidia's CUDA moat, demonstrating that in inference there is zero vendor lock-in and switching providers takes ten keystrokes.
Harry holds his own ▶ 7:47 Citing Nvidia's revenue breakdown in inferenceStebbings demonstrates sharp industry domain knowledge by citing specific revenue figures, pointing out that 40% of Nvidia's revenue relies on HBM chips for inference despite HBM's latency limitations.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Harry as informed peer | Guest teaching | Guest disagreement | Harry pushing back | Why |
|---|---|---|---|---|---|---|
| Welcome and the Genesis of Cerebras in 2015 | 1 | 4 | 1 | 0 | Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication. | |
| Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off | 2 | 6 | 1 | 1 | Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale. | |
| Why the Market Continues to Rely on HBM-Based GPUs | 3 | 4 | 2 | 3 | Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads. | |
| Overcoming the "Impossible" Silicon Yield Challenge | 1 | 6 | 1 | 2 | Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture. | |
| Prioritizing Speed and the Broadband Internet Analogy | 2 | 4 | 1 | 1 | Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts. | |
| The Formula for Inference and the Shift to Utility | 2 | 4 | 1 | 1 | Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility. | |
| Energy Demands and Data Center Construction Sophistication | 4 | 5 | 5 | 4 | Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders. | |
| Reducing Inference Costs and the Future of Algorithmic Efficiency | 3 | 5 | 3 | 3 | Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections. | |
| Simulators, Rare Corner Cases, and Synthetic Data | 3 | 5 | 5 | 3 | Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars. | |
| Jevon's Paradox and the Exponential Expansion of Compute | 4 | 4 | 4 | 3 | Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation. | |
| Hardware Value Distribution and the CUDA Moat Myth | 4 | 6 | 4 | 3 | Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes. | |
| Future AI Chip Market Shares and Timelines | 4 | 4 | 2 | 2 | Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent. | |
| Cerebras' Cashflow Positive Position and the G42 Relationship | 4 | 3 | 3 | 5 | Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters. | |
| The S-1 Filing and the Strategic Advantages of Going Public | 4 | 3 | 4 | 6 | Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology. | |
| Export Controls, China, and Ethical Leadership | 3 | 4 | 2 | 2 | Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions. | |
| AI Policy and the Chinese Technical Threat | 4 | 5 | 4 | 5 | Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses. | |
| Industrial Policy, Infrastructure, and Competitor Inspiration | 3 | 4 | 2 | 2 | Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining. | |
| Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling | 3 | 3 | 2 | 2 | During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong. | |
| Serial Entrepreneurship, Experience, and Edge Computing | 3 | 5 | 3 | 3 | Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications. |