Mar 24, 2025 · 1h 14m · 20vc

Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings

Andrew Feldman · 51m spoken Harry Stebbings · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this 20VC podcast interview, Cerebras CEO Andrew Feldman details how his company's revolutionary wafer-scale chip architecture overcomes traditional GPU bottlenecks to challenge Nvidia's dominance, while analyzing the geopolitical, economic, and technical realities of AI's shift toward utility.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 16.1% of the talking time here. How this is scored →

Harry as informed peer 3.0 Guest teaching 4.4 Guest disagreement 2.6 Harry pushing back 2.7
05100:0015:0030:0045:001:00:000:39–4:07 · Harry as informed peer 1/10 Welcome and the Genesis of Cerebras in 2015 Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication.4:07–7:47 · Harry as informed peer 2/10 Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale.7:47–10:16 · Harry as informed peer 3/10 Why the Market Continues to Rely on HBM-Based GPUs Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads.10:16–14:43 · Harry as informed peer 1/10 Overcoming the "Impossible" Silicon Yield Challenge Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture.14:43–17:43 · Harry as informed peer 2/10 Prioritizing Speed and the Broadband Internet Analogy Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts.17:43–20:21 · Harry as informed peer 2/10 The Formula for Inference and the Shift to Utility Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility.20:21–24:38 · Harry as informed peer 4/10 Energy Demands and Data Center Construction Sophistication Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders.24:38–29:00 · Harry as informed peer 3/10 Reducing Inference Costs and the Future of Algorithmic Efficiency Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections.29:00–32:13 · Harry as informed peer 3/10 Simulators, Rare Corner Cases, and Synthetic Data Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars.32:13–38:01 · Harry as informed peer 4/10 Jevon's Paradox and the Exponential Expansion of Compute Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation.38:01–45:14 · Harry as informed peer 4/10 Hardware Value Distribution and the CUDA Moat Myth Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes.45:14–48:18 · Harry as informed peer 4/10 Future AI Chip Market Shares and Timelines Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent.48:18–51:17 · Harry as informed peer 4/10 Cerebras' Cashflow Positive Position and the G42 Relationship Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters.51:17–55:29 · Harry as informed peer 4/10 The S-1 Filing and the Strategic Advantages of Going Public Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology.55:29–58:44 · Harry as informed peer 3/10 Export Controls, China, and Ethical Leadership Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions.58:44–1:02:53 · Harry as informed peer 4/10 AI Policy and the Chinese Technical Threat Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses.1:02:53–1:05:23 · Harry as informed peer 3/10 Industrial Policy, Infrastructure, and Competitor Inspiration Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining.1:05:23–1:09:13 · Harry as informed peer 3/10 Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong.1:09:13–1:12:36 · Harry as informed peer 3/10 Serial Entrepreneurship, Experience, and Edge Computing Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications.0:39–4:07 · Guest teaching 4/10 Welcome and the Genesis of Cerebras in 2015 Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication.4:07–7:47 · Guest teaching 6/10 Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale.7:47–10:16 · Guest teaching 4/10 Why the Market Continues to Rely on HBM-Based GPUs Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads.10:16–14:43 · Guest teaching 6/10 Overcoming the "Impossible" Silicon Yield Challenge Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture.14:43–17:43 · Guest teaching 4/10 Prioritizing Speed and the Broadband Internet Analogy Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts.17:43–20:21 · Guest teaching 4/10 The Formula for Inference and the Shift to Utility Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility.20:21–24:38 · Guest teaching 5/10 Energy Demands and Data Center Construction Sophistication Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders.24:38–29:00 · Guest teaching 5/10 Reducing Inference Costs and the Future of Algorithmic Efficiency Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections.29:00–32:13 · Guest teaching 5/10 Simulators, Rare Corner Cases, and Synthetic Data Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars.32:13–38:01 · Guest teaching 4/10 Jevon's Paradox and the Exponential Expansion of Compute Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation.38:01–45:14 · Guest teaching 6/10 Hardware Value Distribution and the CUDA Moat Myth Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes.45:14–48:18 · Guest teaching 4/10 Future AI Chip Market Shares and Timelines Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent.48:18–51:17 · Guest teaching 3/10 Cerebras' Cashflow Positive Position and the G42 Relationship Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters.51:17–55:29 · Guest teaching 3/10 The S-1 Filing and the Strategic Advantages of Going Public Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology.55:29–58:44 · Guest teaching 4/10 Export Controls, China, and Ethical Leadership Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions.58:44–1:02:53 · Guest teaching 5/10 AI Policy and the Chinese Technical Threat Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses.1:02:53–1:05:23 · Guest teaching 4/10 Industrial Policy, Infrastructure, and Competitor Inspiration Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining.1:05:23–1:09:13 · Guest teaching 3/10 Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong.1:09:13–1:12:36 · Guest teaching 5/10 Serial Entrepreneurship, Experience, and Edge Computing Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications.0:39–4:07 · Guest disagreement 1/10 Welcome and the Genesis of Cerebras in 2015 Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication.4:07–7:47 · Guest disagreement 1/10 Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale.7:47–10:16 · Guest disagreement 2/10 Why the Market Continues to Rely on HBM-Based GPUs Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads.10:16–14:43 · Guest disagreement 1/10 Overcoming the "Impossible" Silicon Yield Challenge Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture.14:43–17:43 · Guest disagreement 1/10 Prioritizing Speed and the Broadband Internet Analogy Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts.17:43–20:21 · Guest disagreement 1/10 The Formula for Inference and the Shift to Utility Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility.20:21–24:38 · Guest disagreement 5/10 Energy Demands and Data Center Construction Sophistication Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders.24:38–29:00 · Guest disagreement 3/10 Reducing Inference Costs and the Future of Algorithmic Efficiency Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections.29:00–32:13 · Guest disagreement 5/10 Simulators, Rare Corner Cases, and Synthetic Data Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars.32:13–38:01 · Guest disagreement 4/10 Jevon's Paradox and the Exponential Expansion of Compute Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation.38:01–45:14 · Guest disagreement 4/10 Hardware Value Distribution and the CUDA Moat Myth Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes.45:14–48:18 · Guest disagreement 2/10 Future AI Chip Market Shares and Timelines Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent.48:18–51:17 · Guest disagreement 3/10 Cerebras' Cashflow Positive Position and the G42 Relationship Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters.51:17–55:29 · Guest disagreement 4/10 The S-1 Filing and the Strategic Advantages of Going Public Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology.55:29–58:44 · Guest disagreement 2/10 Export Controls, China, and Ethical Leadership Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions.58:44–1:02:53 · Guest disagreement 4/10 AI Policy and the Chinese Technical Threat Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses.1:02:53–1:05:23 · Guest disagreement 2/10 Industrial Policy, Infrastructure, and Competitor Inspiration Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining.1:05:23–1:09:13 · Guest disagreement 2/10 Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong.1:09:13–1:12:36 · Guest disagreement 3/10 Serial Entrepreneurship, Experience, and Edge Computing Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications.0:39–4:07 · Harry pushing back 0/10 Welcome and the Genesis of Cerebras in 2015 Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication.4:07–7:47 · Harry pushing back 1/10 Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale.7:47–10:16 · Harry pushing back 3/10 Why the Market Continues to Rely on HBM-Based GPUs Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads.10:16–14:43 · Harry pushing back 2/10 Overcoming the "Impossible" Silicon Yield Challenge Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture.14:43–17:43 · Harry pushing back 1/10 Prioritizing Speed and the Broadband Internet Analogy Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts.17:43–20:21 · Harry pushing back 1/10 The Formula for Inference and the Shift to Utility Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility.20:21–24:38 · Harry pushing back 4/10 Energy Demands and Data Center Construction Sophistication Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders.24:38–29:00 · Harry pushing back 3/10 Reducing Inference Costs and the Future of Algorithmic Efficiency Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections.29:00–32:13 · Harry pushing back 3/10 Simulators, Rare Corner Cases, and Synthetic Data Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars.32:13–38:01 · Harry pushing back 3/10 Jevon's Paradox and the Exponential Expansion of Compute Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation.38:01–45:14 · Harry pushing back 3/10 Hardware Value Distribution and the CUDA Moat Myth Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes.45:14–48:18 · Harry pushing back 2/10 Future AI Chip Market Shares and Timelines Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent.48:18–51:17 · Harry pushing back 5/10 Cerebras' Cashflow Positive Position and the G42 Relationship Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters.51:17–55:29 · Harry pushing back 6/10 The S-1 Filing and the Strategic Advantages of Going Public Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology.55:29–58:44 · Harry pushing back 2/10 Export Controls, China, and Ethical Leadership Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions.58:44–1:02:53 · Harry pushing back 5/10 AI Policy and the Chinese Technical Threat Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses.1:02:53–1:05:23 · Harry pushing back 2/10 Industrial Policy, Infrastructure, and Competitor Inspiration Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining.1:05:23–1:09:13 · Harry pushing back 2/10 Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong.1:09:13–1:12:36 · Harry pushing back 3/10 Serial Entrepreneurship, Experience, and Edge Computing Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications.

speaking balance: gold is Harry, purple is the guest (3 minute bins)

0:00 · Harry 26.9% · guest 73.1%0:00 · Harry 26.9% · guest 73.1%3:00 · Harry 17.3% · guest 82.7%3:00 · Harry 17.3% · guest 82.7%6:00 · Harry 12.3% · guest 87.7%6:00 · Harry 12.3% · guest 87.7%9:00 · Harry 16.8% · guest 83.2%9:00 · Harry 16.8% · guest 83.2%12:00 · Harry 7.8% · guest 92.2%12:00 · Harry 7.8% · guest 92.2%15:00 · Harry 6.4% · guest 93.6%15:00 · Harry 6.4% · guest 93.6%18:00 · Harry 21.8% · guest 78.2%18:00 · Harry 21.8% · guest 78.2%21:00 · Harry 16.9% · guest 83.1%21:00 · Harry 16.9% · guest 83.1%24:00 · Harry 21.2% · guest 78.8%24:00 · Harry 21.2% · guest 78.8%27:00 · Harry 11.6% · guest 88.4%27:00 · Harry 11.6% · guest 88.4%30:00 · Harry 13.6% · guest 86.4%30:00 · Harry 13.6% · guest 86.4%33:00 · Harry 14.7% · guest 85.3%33:00 · Harry 14.7% · guest 85.3%36:00 · Harry 23.5% · guest 76.5%36:00 · Harry 23.5% · guest 76.5%39:00 · Harry 10% · guest 90%39:00 · Harry 10% · guest 90%42:00 · Harry 10.3% · guest 89.7%42:00 · Harry 10.3% · guest 89.7%45:00 · Harry 20.9% · guest 79.1%45:00 · Harry 20.9% · guest 79.1%48:00 · Harry 17.4% · guest 82.6%48:00 · Harry 17.4% · guest 82.6%51:00 · Harry 28% · guest 72%51:00 · Harry 28% · guest 72%54:00 · Harry 30.2% · guest 69.8%54:00 · Harry 30.2% · guest 69.8%57:00 · Harry 9.3% · guest 90.7%57:00 · Harry 9.3% · guest 90.7%1:00:00 · Harry 9.6% · guest 90.4%1:00:00 · Harry 9.6% · guest 90.4%1:03:00 · Harry 5.1% · guest 94.9%1:03:00 · Harry 5.1% · guest 94.9%1:06:00 · Harry 11.8% · guest 88.2%1:06:00 · Harry 11.8% · guest 88.2%1:09:00 · Harry 21.9% · guest 78.1%1:09:00 · Harry 21.9% · guest 78.1%1:12:00 · Harry 19.1% · guest 80.9%1:12:00 · Harry 19.1% · guest 80.9%
Sharpest disagreement ▶ 29:50 Direct rejection of host consensus premise

Feldman flatly rejects Stebbings' premise that the industry is far along across compute, algorithms, and data, stating 'I think they're wrong' and insisting early-stage progress remains across all pillars.

Hardest push from Harry ▶ 51:18 Challenging IPO timing and S-1 disclosure risks

Stebbings explicitly challenges Feldman on filing an S-1 public disclosure, calling it preemptive and pointing out that it grants competitors asymmetric financial and strategic information.

Biggest teaching moment ▶ 42:00 Debunking CUDA lock-in in AI inference

Feldman directly corrects popular narrative around Nvidia's CUDA moat, demonstrating that in inference there is zero vendor lock-in and switching providers takes ten keystrokes.

Harry holds his own ▶ 7:47 Citing Nvidia's revenue breakdown in inference

Stebbings demonstrates sharp industry domain knowledge by citing specific revenue figures, pointing out that 40% of Nvidia's revenue relies on HBM chips for inference despite HBM's latency limitations.

the scores for every segment, with the reasoning behind each
ChapterTopicHarry as informed peerGuest teachingGuest disagreementHarry pushing backWhy
Welcome and the Genesis of Cerebras in 2015 1410 Stebbings opens with broad background questions and openly framing the episode as a learning experience for himself. Feldman explains the technical origin of Cerebras and the data movement challenges of matrix multiplication.
Wafer-Scale Architecture and the SRAM vs. HBM Trade-Off 2611 Stebbings asks whether a single chip architecture can serve fine-tuning, training, and inference. Feldman gently reframes the computer architecture trade-offs, explaining the differences between high-capacity slow HBM and low-capacity fast SRAM at wafer scale.
Why the Market Continues to Rely on HBM-Based GPUs 3423 Stebbings cites that 40 percent of Nvidia revenue comes from inference on HBM chips to challenge why the market relies on them if HBM is slow. Feldman explains how graphics legacy architectures became weaknesses in dedicated AI workloads.
Overcoming the "Impossible" Silicon Yield Challenge 1612 Stebbings interrupts to ask basic technical definitions regarding silicon yield and why it was deemed impossible to solve. Feldman provides a comprehensive explanation using a cookie-cutter analogy and redundant tile architecture.
Prioritizing Speed and the Broadband Internet Analogy 2411 Stebbings asks how customers prioritize speed versus cost and efficiency. Feldman details batch latency versus interactive mode requirements, drawing comparisons to early broadband internet shifts.
The Formula for Inference and the Shift to Utility 2411 Stebbings prompts Feldman to recount an equation for inference shared prior to the recording. Feldman outlines how AI shifted in late 2024 from a novel tech demo into daily operational workflow utility.
Energy Demands and Data Center Construction Sophistication 4554 Stebbings quotes Groq CEO Jonathan Ross's assertion that data center buildouts are being done by tourists. Feldman directly counters Ross's quote, asserting that early project operators like Crusoe and TeraWulf are highly sophisticated builders.
Reducing Inference Costs and the Future of Algorithmic Efficiency 3533 Stebbings asks how to reconcile claims that scaling laws are plateauing with claims of massive remaining efficiency gains. Feldman explains GPU hardware underutilization (93 percent wasted) and sparsity in network connections.
Simulators, Rare Corner Cases, and Synthetic Data 3553 Stebbings synthesizes the common refrain that compute, algorithms, and data are already far along. Feldman flatly rejects the premise twice, arguing that AI development remains in its early stages across all three pillars.
Jevon's Paradox and the Exponential Expansion of Compute 4443 Stebbings brings up Jevon's paradox and Satya Nadella's commentary, leading to banter about venture capitalists citing 19th-century economists. Feldman highlights DeepSeek's engineering accomplishments and discusses the optics of model distillation.
Hardware Value Distribution and the CUDA Moat Myth 4643 Stebbings probes whether defensibility lies in hardware and questions Nvidia's CUDA lock-in. Feldman thoroughly dismantles the CUDA moat myth in inference, explaining that switching hardware requires ten keystrokes.
Future AI Chip Market Shares and Timelines 4422 Stebbings sets up a structured market dynamics comparison between an Uber winner-take-most structure versus an AWS cloud shared structure. Feldman projects Nvidia's market share will normalize to 50 to 60 percent.
Cerebras' Cashflow Positive Position and the G42 Relationship 4335 Stebbings directly probes Cerebras' revenue risk due to customer concentration with G42. Feldman acknowledges it as both a strength and weakness while describing the operational muscles built from deploying massive clusters.
The S-1 Filing and the Strategic Advantages of Going Public 4346 Stebbings challenges Feldman on filing an S-1 public disclosure, calling it preemptive and noting it hands competitors asymmetric information. Feldman fires back that Cerebras possesses asymmetric technology.
Export Controls, China, and Ethical Leadership 3422 Stebbings asks about hardware export control implementation and effectiveness. Feldman highlights policy complexity and unintended consequences, such as US venture capital backing Chinese EDA startups in response to sanctions.
AI Policy and the Chinese Technical Threat 4545 Stebbings questions Feldman on why Cerebras voluntarily chose not to sell to China if export controls are ineffective. Feldman explains his personal ethical heuristic known as the 'mother test' regarding military and surveillance uses.
Industrial Policy, Infrastructure, and Competitor Inspiration 3422 Stebbings asks about China's industrial policy achievements such as special economic zones in Shenzhen. Feldman discusses how Western nations can draw inspiration from competitor execution and regulatory streamlining.
Quick-Fire: Middle East Peace, Legacy GPUs, and Liquid Cooling 3322 During the quick-fire segment, Stebbings asks about core beliefs, AI threats, and past mistakes. Feldman reveals he fought his chief system architect on water cooling for years before realizing he was wrong.
Serial Entrepreneurship, Experience, and Edge Computing 3533 Stebbings questions the value of serial entrepreneurship after four previous startups. Feldman outlines where experienced operational leadership matters in complex hardware manufacturing versus consumer social applications.

Statements from this episode (44)

Assertion Partly supported
Feldman: GPUs operate at only 5% to 7% utilization during inference
“In a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93% wasted.”
Andrew Feldman Mar 24, 2025 ▶ 25:34
Prediction Open · timeframe Mar 2030
Feldman: AI will rely less on transformers in 3 to 5 years
“We won't be as dependent on transformers in three years or five years as we are now.”
Andrew Feldman Mar 24, 2025 ▶ 34:05
Opinion
Feldman: GPU off-chip memory architecture can be beaten in inference
“The fundamental architecture of the GPU with off-chip memory is not great for inference. Now, they will continue to do well in inference, but it can be beaten, and I think they know it.”
Andrew Feldman Mar 24, 2025 ▶ 0:17
Disclosure
Feldman: Cerebras is my fifth startup and first time underestimating market size
“You know, this is my fifth startup, and the first time I underestimated the size of the market by a lot.”
Andrew Feldman Mar 24, 2025 ▶ 2:01
Insight
Feldman: AI's core matrix math calculations are trivial to design
“First, the underlying calculation is trivial. It's a matrix multiplication, and an FMAC can be developed by any second year electrical engineering student.”
Andrew Feldman Mar 24, 2025 ▶ 3:15
Insight
Feldman: Data movement, not computation, is the core bottleneck in AI chips
“The hard part with AI work is results and intermediate results have to be moved a lot. And therein is the most complicated part. They have to be moved to memory and from memory. And they have to be broken up and moved among GPUs.”
Andrew Feldman Mar 24, 2025 ▶ 3:36
Assertion Supported
Feldman: AI Fine-Tuning and Full Training Share Identical Computational Workloads
“Is the computational work for training from scratch different from fine tuning? And the answer is, it's not different. It's approximately the same.”
Andrew Feldman Mar 24, 2025 ▶ 4:51
Assertion Supported
Feldman: Generating One Word on a 70B AI Model Moves 140GB of Data
“In generative inference, you have to move all the weights from memory to compute to generate a single word. And you have to move them again to generate the next word, and again. So if you have a seventy billion parameter model, not a giant model, and each weig…”
Andrew Feldman Mar 24, 2025 ▶ 5:22
Assertion Supported
Feldman: GPU High Bandwidth Memory Offers High Capacity but Is Slow
“They use memory, a memory called HBM, the type of DRAM, and it is phenomenal memory. But it is slow. It is slow and high capacity.”
Andrew Feldman Mar 24, 2025 ▶ 6:25
Assertion Supported
Feldman: Running DeepSeek 671B on Standard SRAM Chips Requires Up to 8,000 Chips
“If you build a normal size chip with SRAM and you want to do a four hundred billion parameter model and inference, you might need 4000 chips. Or if you want to do a DeepSeq six 71, you might need six or 8000 chips.”
Andrew Feldman Mar 24, 2025 ▶ 7:13
Assertion Open · timeframe Mar 2025
Feldman: Cerebras inference has been the fastest platform since August 2024 launch
“And every day since August 26th when we launched Inference, our way has been the fastest way across a whole set of models tested by artificial analysis and others.”
Andrew Feldman Mar 24, 2025 ▶ 10:05
Assertion Partly supported
Feldman: Cerebras Chip Yield Matches or Exceeds Traditional Smaller Chips
“In fact, we invented techniques that allow us to yield as well or better than others who are building much smaller jobs.”
Andrew Feldman Mar 24, 2025 ▶ 11:20
Assertion Supported
Feldman: Cerebras Is First to Yield Whole Wafers in Computing History
“And that had never been done in a computer before, and that's at the heart of our architecture. That allowed us to yield and deliver whole wafers. Nobody had ever been able to do that in the seventy-year history of our industry.”
Andrew Feldman Mar 24, 2025 ▶ 14:14
Insight
Feldman: Milliseconds of latency destroy user attention in interactive AI
“What we know is that in interactive mode, milliseconds matter. In interactive mode, what Urs Holtz over at Google years ago showed was that you can destroy your user's attention With milliseconds of delay. So being the fastest matters everything in that domain…”
Andrew Feldman Mar 24, 2025 ▶ 16:04
Insight
Feldman: AI inference speed will unlock new industries like broadband did
“If you remember Blockbuster, right. First, I mean, let's look at the history of that. You're exactly right. First, we used to drive to Blockbuster to get a DVD, right, or a video. Then Netflix was mailing them to us, right? And then we got broadband. And there…”
Andrew Feldman Mar 24, 2025 ▶ 17:13
Opinion
Feldman: ChatGPT was a UI invention, not a technical innovation
“ChatGPT was not really a technical innovation. It was a user interface invention.”
Andrew Feldman Mar 24, 2025 ▶ 19:12
Prediction Not checkable as stated
Feldman: AI inference market will grow over 100x in five years
“I think we're way over a hundred times bigger.”
Andrew Feldman Mar 24, 2025 ▶ 20:38
Assertion Not checkable as stated
Feldman: Former Bitcoin miners like TeraWulf and Crusoe lead AI data center buildouts
“I think the guys who were there early were some of the Bitcoin mining companies, like Tara Wolf, the guys at Crusoe, and others, ah, guys in Europe, ah, ah, They were early in building buildings near low cost power in order to run compute that used a lot of po…”
Andrew Feldman Mar 24, 2025 ▶ 23:50
Assertion Supported
Feldman: OpenAI o1 demonstrates scaling laws remain fully functional for inference
“OpenAI's work on O-one shows me that the scaling laws, certainly for inference, are fully functional, right? The more compute you put on inference, the better answer you get.”
Andrew Feldman Mar 24, 2025 ▶ 27:11
Prediction Open · timeframe Mar 2030
Feldman: In five years, almost all AI training data will be synthetic
“Almost all synthetic.”
Andrew Feldman Mar 24, 2025 ▶ 30:18
Assertion Open
Feldman: Cheaper Compute Has Never Shrunk the Computing Market
“There are no, there are very few examples in our industry, actually none in compute in 50 years, in which by making things cheaper and faster, the market got smaller. Market always gets bigger. Always.”
Andrew Feldman Mar 24, 2025 ▶ 33:43
Insight
Feldman: DeepSeek proved AI frontier models don't need thousands of employees
“And what DeepSeek showed us is you don't need 5000 people and, you know, billions of dollars of gear. You can do it with 200 smart people.”
Andrew Feldman Mar 24, 2025 ▶ 35:01
Opinion
Feldman: DeepSeek's open-source release had unprecedented immediate impact
“I think there are a few examples of an open source anything having this sort of immediate impact that model had. Right? I mean, that, that model had a giant impact in a technical community of really smart people, and there are very few examples of other open s…”
Andrew Feldman Mar 24, 2025 ▶ 37:15
Assertion Supported
Feldman: Nvidia Has Zero CUDA Software Lock-In for AI Inference
“In inference, it's not real at all. There's no CUDA lock-in in inference. None. Well, you can move from OpenAI on an NVIDIA GPU to Cerebrus to Fireworks Service on something else to Together to perplexity with 10 keystrokes. I mean, anybody who actually uses A…”
Andrew Feldman Mar 24, 2025 ▶ 42:10
Assertion Supported
Feldman: Intel Kept 75% Market Share Despite a Decade of Failures
“Intel has made, until hiring LibBoo, prior to that, nearly a decade of Catastrophic decisions. Right? And they still own 80% of the x-eighty-six market. 75% of the market. Right? AMD's worked up to like 25% or 30%. And after a decade of screwing up.”
Andrew Feldman Mar 24, 2025 ▶ 44:02
Prediction Open · timeframe Mar 2030
Feldman: Nvidia's AI chip market share will fall to 50-60% in 5 years
“I think you know, in five years from now, NVIDIA is going to have 60. Somewhere between 50 and 60% of the market, right? I think right now they have approximately all of it. I think they will come down over time.”
Andrew Feldman Mar 24, 2025 ▶ 45:43
Prediction Open · timeframe Mar 2030
Feldman: Chip providers will hold far more market value than model companies
“In the five year timeframe, Yes.”
Andrew Feldman Mar 24, 2025 ▶ 47:02
Insight
Feldman: Negative gross margins show a tech company sells commodities
“Traditionally, your gross margins were a measure of your technical differentiation, right? And I think if you're running a negative gross margin business, I think you're It speaks for itself. You're selling commodity. You're not, ah, your value creation isn't …”
Andrew Feldman Mar 24, 2025 ▶ 48:37
Assertion Contradicted
Feldman: Cerebras deployed more exaflops than anyone except Nvidia and AMD
“We, we've deployed tens of exaflops of compute vastly more than anybody else that, that, that isn't AMD or NVIDIA, right?”
Andrew Feldman Mar 24, 2025 ▶ 49:55
Assertion Supported
Feldman: AI startup valuations match historically public-market-only levels
“The valuations that Anthropic and OpenAI and some of the others are getting are historically public market only valuations.”
Andrew Feldman Mar 24, 2025 ▶ 51:42
Prediction Held up
Feldman: Cerebras will be among the first AI hardware companies to go public
“We think that we will be among the first in the category.”
Andrew Feldman Mar 24, 2025 ▶ 52:30
Prediction Open · timeframe Mar 2027
Feldman: Cerebras will land several G42-scale deals within 24 months
“That's a good question. Several.”
Andrew Feldman Mar 24, 2025 ▶ 53:05
Assertion Partly supported
Feldman: Cerebras' deal with G42 was estimated at over $1 billion
“I mean, when we announced it was some estimated it was north of a billion.”
Andrew Feldman Mar 24, 2025 ▶ 53:11
Insight
Feldman: Nvidia customer delivery delays create major opening for competitors
“And so I think there is a real opportunity in the potential for NVIDIA Customer Unhappiness, for sure. For those of us who are competing with them. I mean, if you can't get your gear, you may as well test somebody else's. And that, that's a huge opening.”
Andrew Feldman Mar 24, 2025 ▶ 55:11
Assertion Open · timeframe Mar 2026
Feldman: DeepSeek likely accessed AI chips hosted in Singapore
“It turns out that they probably did use chips in Singapore.”
Andrew Feldman Mar 24, 2025 ▶ 56:35
Assertion Contradicted
Feldman: US export controls led US VCs to fund Chinese EDA startups
“You, ah, sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market. And so U.S. Venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools, right?”
Andrew Feldman Mar 24, 2025 ▶ 57:49
Opinion
Feldman: Current US administration is better for AI than prior administration
“I think the past administration lined itself up against big tech. And that, that, that was a mistake. AI is also in a different place, so it's easier to be for it. Right? It's less scary now than it was in, I don't know, 21. Right? It, we sort of have a better…”
Andrew Feldman Mar 24, 2025 ▶ 59:01
Disclosure
Feldman: Cerebras rejected China AI chip deal over ethical concerns
“And I asked myself that, and I came to believe that, that the deal on the table was, wouldn't be used for good, and I wasn't comfortable with that, and I wouldn't have been able to explain to my mother. And that's a moral compass, and I think that's, ah, ah.”
Andrew Feldman Mar 24, 2025 ▶ 1:00:37
Opinion
Feldman: Western tech leaders fundamentally underestimate China's AI capabilities
“A hundred percent. I think, and it is one of the most obvious and frequent errors in judgment. Is that you underestimate the other side. I think, ah, you have to look carefully at what they're doing, and their investment in infrastructure has been extraordinar…”
Andrew Feldman Mar 24, 2025 ▶ 1:01:28
Prediction Didn’t hold up
Feldman: AI adoption will equal cell phone penetration within two years
“I do think that Within a year or two. AI's penetration will be approximately the same as telephones, cell phones.”
Andrew Feldman Mar 24, 2025 ▶ 1:07:02
Disclosure
Feldman Resisted Water-Cooled Chip Design at Cerebras in 2016
“In, ah, 2016, JP, one of our co-founders and chief system architect, Laid out a plan that would have us doing water cooling, and for our systems, and nobody else was doing it, and I fought so hard, and I was so wrong. JP was right.”
Andrew Feldman Mar 24, 2025 ▶ 1:07:35
Insight
Feldman: Experience matters far more than naivety in complex hardware startups
“I think if you are in a business in which running a business is a benefit, then experience matters a great deal. I think if you are in a business in which you look like Your customer. There was a reason why social, social networks were started by people right …”
Andrew Feldman Mar 24, 2025 ▶ 1:09:40
Prediction Not checkable as stated
Feldman: Sub-milliwatt sensor chips will sell massive volume and power robotics
“I'd say in the chip world, the sub milliwatt, really tiny, tiny little chips that live next to sensors that do inference. These are tiny little things that will only send back useful data. Is a, an extremely interesting market, and they will sell enormous volu…”
Andrew Feldman Mar 24, 2025 ▶ 1:11:54
Prediction Not checkable as stated
Feldman: Cerebras tech will enable novel disease therapeutics within five years
“I think in three to five years, I would like our technology to have been used to solve two important societal problems. I would like it to be used to have found a therapeutic for, ah, an affliction that, that impacts more than a million people a year. I would …”
Andrew Feldman Mar 24, 2025 ▶ 1:12:56

Shorts cut from this episode

▶ Intel’s Indestructible Moat · 20VC with Harry Stebbings (@44:11) ▶ “NVIDIA can be BEATEN, and they know it” 😱 · 20VC with Harr (@0:00) ▶ "America’s Power Is in the Wrong Places" ⚡️ · 20VC with Harr (@22:03)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.