Dec 23, 2024 · 1h 29m · bg2-pod

AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod

Dylan Patel · 59m spoken Brad Gerstner · 15m spoken Bill Gurley · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of BG², host Brad Gerstner and co-host Bill Gurley interview Dylan Patel, founder of SemiAnalysis, to analyze the state of the global AI semiconductor landscape, hyperscaler data center power constraints, NVIDIA's competitive moats, custom ASIC developments, and the shifting economics of AI compute.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Brad and Bill hold 28.4% of the talking time here. How this is scored →

Brad and Bill as informed peer 6.0 Guest teaching 6.1 Guest disagreement 1.8 Brad and Bill pushing back 1.2
05100:0020:0040:001:00:001:20:001:52–4:15 · Brad and Bill as informed peer 4/10 Dylan Patel's Origin Story in Semiconductors Brad introduces Dylan Patel and asks about his background. Dylan explains fixing his Xbox as an 8-year-old and reading semiconductor earnings reports as an intern. The tone is warm and conversational.4:15–6:58 · Brad and Bill as informed peer 5/10 NVIDIA Market Dominance and Google's TPU Exception Bill asks about NVIDIA's workload share, assuming Google workloads are non-LLM. Dylan clarifies that Google runs 30% of production workloads and has run transformers since 2018 via BERT.6:58–10:58 · Brad and Bill as informed peer 6/10 The 'Three-Headed Dragon': Why NVIDIA Dominates Bill and Brad ask about NVIDIA's competitive moat across software and hardware. Dylan outlines the 'three-headed dragon' (hardware, software, networking) and notes Google did rack-scale TPUs back in 2018 with Broadcom before NVIDIA's Blackwell.10:58–13:12 · Brad and Bill as informed peer 5/10 NVIDIA's Supply Chain Strategy & Annual Release Cadence Brad and Bill explore how NVIDIA's annual hardware release cadence defends their lead. Dylan explains how Jensen operates on a paranoid, short-term deployment cycle that relentlessly pushes supply chain partners.13:12–17:18 · Brad and Bill as informed peer 6/10 Inference vs. Training Moats and Performance TCO Bill asks about CUDA and training versus inference moats. Dylan distinguishes researcher-driven training from production inference where Microsoft hand-optimizes AMD hardware for fixed models.17:18–21:58 · Brad and Bill as informed peer 7/10 Data Center Buildout: $1 Trillion AI Workloads and CPU Replacement Thesis Brad shows Jensen's chart claiming a $1T CPU data center replacement cycle. Dylan pushes back slightly, pointing out mainframes and classic CPUs continue growing, but notes upgrading dense modern CPUs effectively creates power headroom for AI servers.21:58–28:16 · Brad and Bill as informed peer 6/10 Analyzing Satya Nadella's Power and Data Center Bottleneck Comments Bill and Brad bring up Ilya Sutskever's claim that internet pre-training data is tapped out. Dylan explains how synthetic data generation and functional verification provide a brand-new scaling vector.28:16–34:00 · Brad and Bill as informed peer 6/10 Functional Verification: Where Synthetic Data Works Best Brad asks why hyperscalers are building 300k chip clusters if pre-training is stalling. Dylan explains that logarithmic scaling laws and massive synthetic rollout filtering still require massive compute capacity.34:00–36:26 · Brad and Bill as informed peer 5/10 Objective vs. Subjective Grading in AI Training Models Bill asks for examples of where synthetic data works versus fails. Dylan distinguishes objective domains like code compilation and math from subjective domains like art and email writing.36:26–40:57 · Brad and Bill as informed peer 6/10 Hyperscaler CapEx Acceleration and Multi-Gigawatt Data Center Buildouts Dylan notes Wall Street capex estimates are far too low based on SemiAnalysis tracking multi-gigawatt buildouts. Bill asks about distributed training, and Dylan highlights Microsoft's fiber investments enabling split workloads.40:57–44:03 · Brad and Bill as informed peer 8/10 Evaluating the NVIDIA vs. Cisco 2000 Bubble Comparison Brad forcefully challenges the CNBC Cisco 2000 comparison by citing NVIDIA's 30x PE versus Cisco's 120x PE. Dylan agrees, emphasizing that current capex comes from profitable hyperscalers rather than debt-financed telecom speculation.44:03–50:14 · Brad and Bill as informed peer 6/10 Inference-Time Reasoning Economics and the 50x Compute Cost Increase Dylan details the economics of inference-time reasoning, demonstrating that generating 10,000 hidden reasoning tokens plus smaller batch sizes increases compute cost per user query by 50x.50:14–53:12 · Brad and Bill as informed peer 6/10 Enterprise Value Proposition and ROI of Reasoning Models Brad and Dylan discuss how enterprises will gladly pay a 50x token premium because replacing or augmenting a $300k white-collar engineer provides enormous ROI.53:12–1:00:35 · Brad and Bill as informed peer 7/10 Model Bifurcation: Frontier Model Premium vs. Small Model Commoditization Bill compares AI infrastructure commoditization to Sun/Oracle in dot-com. Dylan explains the bifurcation: small open-source models face brutal price wars, while frontier models command high margins.1:00:35–1:05:39 · Brad and Bill as informed peer 6/10 The High Bandwidth Memory (HBM) Boom & Market Dynamics Brad asks about High Bandwidth Memory (HBM). Dylan explains that HBM from SK Hynix is NVIDIA's largest COGS item above TSMC wafer costs and details Samsung's struggles in high-end HBM.1:05:39–1:10:11 · Brad and Bill as informed peer 6/10 Alternative AI Hardware: AMD's Prospects, MI300X, and Software Bottlenecks Bill asks how AMD is performing. Dylan bluntly states AMD has 'no clue how to do software' and refuses to invest in internal dogfooding supercomputers, leaving them dependent on hyperscaler assistance.1:10:11–1:13:28 · Brad and Bill as informed peer 6/10 Google TPU Architecture, System Maturity, and Cloud Rental Strategy Dylan explains Google's TPU system maturity and liquid cooling advantages, but notes Google fails to commercialize TPUs because of restrictive internal software, opaque cloud pricing, and high internal demand.1:13:28–1:17:33 · Brad and Bill as informed peer 6/10 Amazon's Trainium 2 ('Amazon Basics TPU') and the 400k Chip Cluster Dylan nicknames Amazon Trainium 2 the 'Amazon Basics TPU,' highlighting its cost-effectiveness in HBM per dollar and Anthropic's planned 400k-chip supercomputer cluster.1:17:33–1:20:01 · Brad and Bill as informed peer 6/10 Broadcom & Custom ASIC Landscape for 2025/2026 Brad asks about Broadcom's run-up versus NVIDIA. Dylan details Broadcom's custom ASIC pipeline with Meta, Apple, and OpenAI, alongside their critical networking switch dominance.1:20:01–1:23:38 · Brad and Bill as informed peer 6/10 2025–2026 Semiconductor Market Outlook & NeoCloud Consolidation Dylan forecasts massive 2025 capex growth followed by a potential 2026 reckoning, while predicting severe consolidation among the 80 existing NeoClouds as GPU rental prices collapse.1:23:38–1:28:05 · Brad and Bill as informed peer 7/10 Revenue Growth Trajectory, Model Scaling & Capital Competition Brad and Bill push back on endless spending by citing Satya's discipline of tying capex to real revenue. Dylan argues that fear of being outscaled by competitors will drive hyperscalers to spend ahead anyway.1:52–4:15 · Guest teaching 2/10 Dylan Patel's Origin Story in Semiconductors Brad introduces Dylan Patel and asks about his background. Dylan explains fixing his Xbox as an 8-year-old and reading semiconductor earnings reports as an intern. The tone is warm and conversational.4:15–6:58 · Guest teaching 7/10 NVIDIA Market Dominance and Google's TPU Exception Bill asks about NVIDIA's workload share, assuming Google workloads are non-LLM. Dylan clarifies that Google runs 30% of production workloads and has run transformers since 2018 via BERT.6:58–10:58 · Guest teaching 6/10 The 'Three-Headed Dragon': Why NVIDIA Dominates Bill and Brad ask about NVIDIA's competitive moat across software and hardware. Dylan outlines the 'three-headed dragon' (hardware, software, networking) and notes Google did rack-scale TPUs back in 2018 with Broadcom before NVIDIA's Blackwell.10:58–13:12 · Guest teaching 5/10 NVIDIA's Supply Chain Strategy & Annual Release Cadence Brad and Bill explore how NVIDIA's annual hardware release cadence defends their lead. Dylan explains how Jensen operates on a paranoid, short-term deployment cycle that relentlessly pushes supply chain partners.13:12–17:18 · Guest teaching 7/10 Inference vs. Training Moats and Performance TCO Bill asks about CUDA and training versus inference moats. Dylan distinguishes researcher-driven training from production inference where Microsoft hand-optimizes AMD hardware for fixed models.17:18–21:58 · Guest teaching 6/10 Data Center Buildout: $1 Trillion AI Workloads and CPU Replacement Thesis Brad shows Jensen's chart claiming a $1T CPU data center replacement cycle. Dylan pushes back slightly, pointing out mainframes and classic CPUs continue growing, but notes upgrading dense modern CPUs effectively creates power headroom for AI servers.21:58–28:16 · Guest teaching 7/10 Analyzing Satya Nadella's Power and Data Center Bottleneck Comments Bill and Brad bring up Ilya Sutskever's claim that internet pre-training data is tapped out. Dylan explains how synthetic data generation and functional verification provide a brand-new scaling vector.28:16–34:00 · Guest teaching 6/10 Functional Verification: Where Synthetic Data Works Best Brad asks why hyperscalers are building 300k chip clusters if pre-training is stalling. Dylan explains that logarithmic scaling laws and massive synthetic rollout filtering still require massive compute capacity.34:00–36:26 · Guest teaching 6/10 Objective vs. Subjective Grading in AI Training Models Bill asks for examples of where synthetic data works versus fails. Dylan distinguishes objective domains like code compilation and math from subjective domains like art and email writing.36:26–40:57 · Guest teaching 7/10 Hyperscaler CapEx Acceleration and Multi-Gigawatt Data Center Buildouts Dylan notes Wall Street capex estimates are far too low based on SemiAnalysis tracking multi-gigawatt buildouts. Bill asks about distributed training, and Dylan highlights Microsoft's fiber investments enabling split workloads.40:57–44:03 · Guest teaching 4/10 Evaluating the NVIDIA vs. Cisco 2000 Bubble Comparison Brad forcefully challenges the CNBC Cisco 2000 comparison by citing NVIDIA's 30x PE versus Cisco's 120x PE. Dylan agrees, emphasizing that current capex comes from profitable hyperscalers rather than debt-financed telecom speculation.44:03–50:14 · Guest teaching 8/10 Inference-Time Reasoning Economics and the 50x Compute Cost Increase Dylan details the economics of inference-time reasoning, demonstrating that generating 10,000 hidden reasoning tokens plus smaller batch sizes increases compute cost per user query by 50x.50:14–53:12 · Guest teaching 6/10 Enterprise Value Proposition and ROI of Reasoning Models Brad and Dylan discuss how enterprises will gladly pay a 50x token premium because replacing or augmenting a $300k white-collar engineer provides enormous ROI.53:12–1:00:35 · Guest teaching 6/10 Model Bifurcation: Frontier Model Premium vs. Small Model Commoditization Bill compares AI infrastructure commoditization to Sun/Oracle in dot-com. Dylan explains the bifurcation: small open-source models face brutal price wars, while frontier models command high margins.1:00:35–1:05:39 · Guest teaching 7/10 The High Bandwidth Memory (HBM) Boom & Market Dynamics Brad asks about High Bandwidth Memory (HBM). Dylan explains that HBM from SK Hynix is NVIDIA's largest COGS item above TSMC wafer costs and details Samsung's struggles in high-end HBM.1:05:39–1:10:11 · Guest teaching 7/10 Alternative AI Hardware: AMD's Prospects, MI300X, and Software Bottlenecks Bill asks how AMD is performing. Dylan bluntly states AMD has 'no clue how to do software' and refuses to invest in internal dogfooding supercomputers, leaving them dependent on hyperscaler assistance.1:10:11–1:13:28 · Guest teaching 6/10 Google TPU Architecture, System Maturity, and Cloud Rental Strategy Dylan explains Google's TPU system maturity and liquid cooling advantages, but notes Google fails to commercialize TPUs because of restrictive internal software, opaque cloud pricing, and high internal demand.1:13:28–1:17:33 · Guest teaching 7/10 Amazon's Trainium 2 ('Amazon Basics TPU') and the 400k Chip Cluster Dylan nicknames Amazon Trainium 2 the 'Amazon Basics TPU,' highlighting its cost-effectiveness in HBM per dollar and Anthropic's planned 400k-chip supercomputer cluster.1:17:33–1:20:01 · Guest teaching 6/10 Broadcom & Custom ASIC Landscape for 2025/2026 Brad asks about Broadcom's run-up versus NVIDIA. Dylan details Broadcom's custom ASIC pipeline with Meta, Apple, and OpenAI, alongside their critical networking switch dominance.1:20:01–1:23:38 · Guest teaching 7/10 2025–2026 Semiconductor Market Outlook & NeoCloud Consolidation Dylan forecasts massive 2025 capex growth followed by a potential 2026 reckoning, while predicting severe consolidation among the 80 existing NeoClouds as GPU rental prices collapse.1:23:38–1:28:05 · Guest teaching 6/10 Revenue Growth Trajectory, Model Scaling & Capital Competition Brad and Bill push back on endless spending by citing Satya's discipline of tying capex to real revenue. Dylan argues that fear of being outscaled by competitors will drive hyperscalers to spend ahead anyway.1:52–4:15 · Guest disagreement 1/10 Dylan Patel's Origin Story in Semiconductors Brad introduces Dylan Patel and asks about his background. Dylan explains fixing his Xbox as an 8-year-old and reading semiconductor earnings reports as an intern. The tone is warm and conversational.4:15–6:58 · Guest disagreement 2/10 NVIDIA Market Dominance and Google's TPU Exception Bill asks about NVIDIA's workload share, assuming Google workloads are non-LLM. Dylan clarifies that Google runs 30% of production workloads and has run transformers since 2018 via BERT.6:58–10:58 · Guest disagreement 2/10 The 'Three-Headed Dragon': Why NVIDIA Dominates Bill and Brad ask about NVIDIA's competitive moat across software and hardware. Dylan outlines the 'three-headed dragon' (hardware, software, networking) and notes Google did rack-scale TPUs back in 2018 with Broadcom before NVIDIA's Blackwell.10:58–13:12 · Guest disagreement 1/10 NVIDIA's Supply Chain Strategy & Annual Release Cadence Brad and Bill explore how NVIDIA's annual hardware release cadence defends their lead. Dylan explains how Jensen operates on a paranoid, short-term deployment cycle that relentlessly pushes supply chain partners.13:12–17:18 · Guest disagreement 2/10 Inference vs. Training Moats and Performance TCO Bill asks about CUDA and training versus inference moats. Dylan distinguishes researcher-driven training from production inference where Microsoft hand-optimizes AMD hardware for fixed models.17:18–21:58 · Guest disagreement 3/10 Data Center Buildout: $1 Trillion AI Workloads and CPU Replacement Thesis Brad shows Jensen's chart claiming a $1T CPU data center replacement cycle. Dylan pushes back slightly, pointing out mainframes and classic CPUs continue growing, but notes upgrading dense modern CPUs effectively creates power headroom for AI servers.21:58–28:16 · Guest disagreement 2/10 Analyzing Satya Nadella's Power and Data Center Bottleneck Comments Bill and Brad bring up Ilya Sutskever's claim that internet pre-training data is tapped out. Dylan explains how synthetic data generation and functional verification provide a brand-new scaling vector.28:16–34:00 · Guest disagreement 2/10 Functional Verification: Where Synthetic Data Works Best Brad asks why hyperscalers are building 300k chip clusters if pre-training is stalling. Dylan explains that logarithmic scaling laws and massive synthetic rollout filtering still require massive compute capacity.34:00–36:26 · Guest disagreement 1/10 Objective vs. Subjective Grading in AI Training Models Bill asks for examples of where synthetic data works versus fails. Dylan distinguishes objective domains like code compilation and math from subjective domains like art and email writing.36:26–40:57 · Guest disagreement 2/10 Hyperscaler CapEx Acceleration and Multi-Gigawatt Data Center Buildouts Dylan notes Wall Street capex estimates are far too low based on SemiAnalysis tracking multi-gigawatt buildouts. Bill asks about distributed training, and Dylan highlights Microsoft's fiber investments enabling split workloads.40:57–44:03 · Guest disagreement 1/10 Evaluating the NVIDIA vs. Cisco 2000 Bubble Comparison Brad forcefully challenges the CNBC Cisco 2000 comparison by citing NVIDIA's 30x PE versus Cisco's 120x PE. Dylan agrees, emphasizing that current capex comes from profitable hyperscalers rather than debt-financed telecom speculation.44:03–50:14 · Guest disagreement 2/10 Inference-Time Reasoning Economics and the 50x Compute Cost Increase Dylan details the economics of inference-time reasoning, demonstrating that generating 10,000 hidden reasoning tokens plus smaller batch sizes increases compute cost per user query by 50x.50:14–53:12 · Guest disagreement 1/10 Enterprise Value Proposition and ROI of Reasoning Models Brad and Dylan discuss how enterprises will gladly pay a 50x token premium because replacing or augmenting a $300k white-collar engineer provides enormous ROI.53:12–1:00:35 · Guest disagreement 2/10 Model Bifurcation: Frontier Model Premium vs. Small Model Commoditization Bill compares AI infrastructure commoditization to Sun/Oracle in dot-com. Dylan explains the bifurcation: small open-source models face brutal price wars, while frontier models command high margins.1:00:35–1:05:39 · Guest disagreement 2/10 The High Bandwidth Memory (HBM) Boom & Market Dynamics Brad asks about High Bandwidth Memory (HBM). Dylan explains that HBM from SK Hynix is NVIDIA's largest COGS item above TSMC wafer costs and details Samsung's struggles in high-end HBM.1:05:39–1:10:11 · Guest disagreement 3/10 Alternative AI Hardware: AMD's Prospects, MI300X, and Software Bottlenecks Bill asks how AMD is performing. Dylan bluntly states AMD has 'no clue how to do software' and refuses to invest in internal dogfooding supercomputers, leaving them dependent on hyperscaler assistance.1:10:11–1:13:28 · Guest disagreement 1/10 Google TPU Architecture, System Maturity, and Cloud Rental Strategy Dylan explains Google's TPU system maturity and liquid cooling advantages, but notes Google fails to commercialize TPUs because of restrictive internal software, opaque cloud pricing, and high internal demand.1:13:28–1:17:33 · Guest disagreement 2/10 Amazon's Trainium 2 ('Amazon Basics TPU') and the 400k Chip Cluster Dylan nicknames Amazon Trainium 2 the 'Amazon Basics TPU,' highlighting its cost-effectiveness in HBM per dollar and Anthropic's planned 400k-chip supercomputer cluster.1:17:33–1:20:01 · Guest disagreement 1/10 Broadcom & Custom ASIC Landscape for 2025/2026 Brad asks about Broadcom's run-up versus NVIDIA. Dylan details Broadcom's custom ASIC pipeline with Meta, Apple, and OpenAI, alongside their critical networking switch dominance.1:20:01–1:23:38 · Guest disagreement 2/10 2025–2026 Semiconductor Market Outlook & NeoCloud Consolidation Dylan forecasts massive 2025 capex growth followed by a potential 2026 reckoning, while predicting severe consolidation among the 80 existing NeoClouds as GPU rental prices collapse.1:23:38–1:28:05 · Guest disagreement 2/10 Revenue Growth Trajectory, Model Scaling & Capital Competition Brad and Bill push back on endless spending by citing Satya's discipline of tying capex to real revenue. Dylan argues that fear of being outscaled by competitors will drive hyperscalers to spend ahead anyway.1:52–4:15 · Brad and Bill pushing back 0/10 Dylan Patel's Origin Story in Semiconductors Brad introduces Dylan Patel and asks about his background. Dylan explains fixing his Xbox as an 8-year-old and reading semiconductor earnings reports as an intern. The tone is warm and conversational.4:15–6:58 · Brad and Bill pushing back 1/10 NVIDIA Market Dominance and Google's TPU Exception Bill asks about NVIDIA's workload share, assuming Google workloads are non-LLM. Dylan clarifies that Google runs 30% of production workloads and has run transformers since 2018 via BERT.6:58–10:58 · Brad and Bill pushing back 1/10 The 'Three-Headed Dragon': Why NVIDIA Dominates Bill and Brad ask about NVIDIA's competitive moat across software and hardware. Dylan outlines the 'three-headed dragon' (hardware, software, networking) and notes Google did rack-scale TPUs back in 2018 with Broadcom before NVIDIA's Blackwell.10:58–13:12 · Brad and Bill pushing back 1/10 NVIDIA's Supply Chain Strategy & Annual Release Cadence Brad and Bill explore how NVIDIA's annual hardware release cadence defends their lead. Dylan explains how Jensen operates on a paranoid, short-term deployment cycle that relentlessly pushes supply chain partners.13:12–17:18 · Brad and Bill pushing back 2/10 Inference vs. Training Moats and Performance TCO Bill asks about CUDA and training versus inference moats. Dylan distinguishes researcher-driven training from production inference where Microsoft hand-optimizes AMD hardware for fixed models.17:18–21:58 · Brad and Bill pushing back 2/10 Data Center Buildout: $1 Trillion AI Workloads and CPU Replacement Thesis Brad shows Jensen's chart claiming a $1T CPU data center replacement cycle. Dylan pushes back slightly, pointing out mainframes and classic CPUs continue growing, but notes upgrading dense modern CPUs effectively creates power headroom for AI servers.21:58–28:16 · Brad and Bill pushing back 1/10 Analyzing Satya Nadella's Power and Data Center Bottleneck Comments Bill and Brad bring up Ilya Sutskever's claim that internet pre-training data is tapped out. Dylan explains how synthetic data generation and functional verification provide a brand-new scaling vector.28:16–34:00 · Brad and Bill pushing back 2/10 Functional Verification: Where Synthetic Data Works Best Brad asks why hyperscalers are building 300k chip clusters if pre-training is stalling. Dylan explains that logarithmic scaling laws and massive synthetic rollout filtering still require massive compute capacity.34:00–36:26 · Brad and Bill pushing back 1/10 Objective vs. Subjective Grading in AI Training Models Bill asks for examples of where synthetic data works versus fails. Dylan distinguishes objective domains like code compilation and math from subjective domains like art and email writing.36:26–40:57 · Brad and Bill pushing back 1/10 Hyperscaler CapEx Acceleration and Multi-Gigawatt Data Center Buildouts Dylan notes Wall Street capex estimates are far too low based on SemiAnalysis tracking multi-gigawatt buildouts. Bill asks about distributed training, and Dylan highlights Microsoft's fiber investments enabling split workloads.40:57–44:03 · Brad and Bill pushing back 1/10 Evaluating the NVIDIA vs. Cisco 2000 Bubble Comparison Brad forcefully challenges the CNBC Cisco 2000 comparison by citing NVIDIA's 30x PE versus Cisco's 120x PE. Dylan agrees, emphasizing that current capex comes from profitable hyperscalers rather than debt-financed telecom speculation.44:03–50:14 · Brad and Bill pushing back 1/10 Inference-Time Reasoning Economics and the 50x Compute Cost Increase Dylan details the economics of inference-time reasoning, demonstrating that generating 10,000 hidden reasoning tokens plus smaller batch sizes increases compute cost per user query by 50x.50:14–53:12 · Brad and Bill pushing back 1/10 Enterprise Value Proposition and ROI of Reasoning Models Brad and Dylan discuss how enterprises will gladly pay a 50x token premium because replacing or augmenting a $300k white-collar engineer provides enormous ROI.53:12–1:00:35 · Brad and Bill pushing back 2/10 Model Bifurcation: Frontier Model Premium vs. Small Model Commoditization Bill compares AI infrastructure commoditization to Sun/Oracle in dot-com. Dylan explains the bifurcation: small open-source models face brutal price wars, while frontier models command high margins.1:00:35–1:05:39 · Brad and Bill pushing back 1/10 The High Bandwidth Memory (HBM) Boom & Market Dynamics Brad asks about High Bandwidth Memory (HBM). Dylan explains that HBM from SK Hynix is NVIDIA's largest COGS item above TSMC wafer costs and details Samsung's struggles in high-end HBM.1:05:39–1:10:11 · Brad and Bill pushing back 2/10 Alternative AI Hardware: AMD's Prospects, MI300X, and Software Bottlenecks Bill asks how AMD is performing. Dylan bluntly states AMD has 'no clue how to do software' and refuses to invest in internal dogfooding supercomputers, leaving them dependent on hyperscaler assistance.1:10:11–1:13:28 · Brad and Bill pushing back 1/10 Google TPU Architecture, System Maturity, and Cloud Rental Strategy Dylan explains Google's TPU system maturity and liquid cooling advantages, but notes Google fails to commercialize TPUs because of restrictive internal software, opaque cloud pricing, and high internal demand.1:13:28–1:17:33 · Brad and Bill pushing back 1/10 Amazon's Trainium 2 ('Amazon Basics TPU') and the 400k Chip Cluster Dylan nicknames Amazon Trainium 2 the 'Amazon Basics TPU,' highlighting its cost-effectiveness in HBM per dollar and Anthropic's planned 400k-chip supercomputer cluster.1:17:33–1:20:01 · Brad and Bill pushing back 1/10 Broadcom & Custom ASIC Landscape for 2025/2026 Brad asks about Broadcom's run-up versus NVIDIA. Dylan details Broadcom's custom ASIC pipeline with Meta, Apple, and OpenAI, alongside their critical networking switch dominance.1:20:01–1:23:38 · Brad and Bill pushing back 1/10 2025–2026 Semiconductor Market Outlook & NeoCloud Consolidation Dylan forecasts massive 2025 capex growth followed by a potential 2026 reckoning, while predicting severe consolidation among the 80 existing NeoClouds as GPU rental prices collapse.1:23:38–1:28:05 · Brad and Bill pushing back 2/10 Revenue Growth Trajectory, Model Scaling & Capital Competition Brad and Bill push back on endless spending by citing Satya's discipline of tying capex to real revenue. Dylan argues that fear of being outscaled by competitors will drive hyperscalers to spend ahead anyway.

speaking balance: gold is Brad and Bill, purple is the guest (3 minute bins)

0:00 · Brad and Bill 43.5% · guest 56.5%0:00 · Brad and Bill 43.5% · guest 56.5%3:00 · Brad and Bill 48.2% · guest 51.8%3:00 · Brad and Bill 48.2% · guest 51.8%6:00 · Brad and Bill 47.3% · guest 52.7%6:00 · Brad and Bill 47.3% · guest 52.7%9:00 · Brad and Bill 3.4% · guest 96.6%9:00 · Brad and Bill 3.4% · guest 96.6%12:00 · Brad and Bill 22.7% · guest 77.3%12:00 · Brad and Bill 22.7% · guest 77.3%15:00 · Brad and Bill 31.1% · guest 68.9%15:00 · Brad and Bill 31.1% · guest 68.9%18:00 · Brad and Bill 23% · guest 77%18:00 · Brad and Bill 23% · guest 77%21:00 · Brad and Bill 51.5% · guest 48.5%21:00 · Brad and Bill 51.5% · guest 48.5%24:00 · Brad and Bill 1.6% · guest 98.4%24:00 · Brad and Bill 1.6% · guest 98.4%27:00 · Brad and Bill 28.3% · guest 71.7%27:00 · Brad and Bill 28.3% · guest 71.7%30:00 · Brad and Bill 55.5% · guest 44.5%30:00 · Brad and Bill 55.5% · guest 44.5%33:00 · Brad and Bill 12% · guest 88%33:00 · Brad and Bill 12% · guest 88%36:00 · Brad and Bill 29.2% · guest 70.8%36:00 · Brad and Bill 29.2% · guest 70.8%39:00 · Brad and Bill 36.5% · guest 63.5%39:00 · Brad and Bill 36.5% · guest 63.5%42:00 · Brad and Bill 49.7% · guest 50.3%42:00 · Brad and Bill 49.7% · guest 50.3%45:00 · Brad and Bill 0.8% · guest 99.2%45:00 · Brad and Bill 0.8% · guest 99.2%48:00 · Brad and Bill 18.4% · guest 81.6%48:00 · Brad and Bill 18.4% · guest 81.6%51:00 · Brad and Bill 34.5% · guest 65.5%51:00 · Brad and Bill 34.5% · guest 65.5%54:00 · Brad and Bill 3.6% · guest 96.4%54:00 · Brad and Bill 3.6% · guest 96.4%57:00 · Brad and Bill 24.5% · guest 75.5%57:00 · Brad and Bill 24.5% · guest 75.5%1:00:00 · Brad and Bill 34.8% · guest 65.2%1:00:00 · Brad and Bill 34.8% · guest 65.2%1:03:00 · Brad and Bill 20.8% · guest 79.2%1:03:00 · Brad and Bill 20.8% · guest 79.2%1:06:00 · Brad and Bill 22.7% · guest 77.3%1:06:00 · Brad and Bill 22.7% · guest 77.3%1:09:00 · Brad and Bill 14.6% · guest 85.4%1:09:00 · Brad and Bill 14.6% · guest 85.4%1:12:00 · Brad and Bill 19% · guest 81%1:12:00 · Brad and Bill 19% · guest 81%1:15:00 · Brad and Bill 14.8% · guest 85.2%1:15:00 · Brad and Bill 14.8% · guest 85.2%1:18:00 · Brad and Bill 24.1% · guest 75.9%1:18:00 · Brad and Bill 24.1% · guest 75.9%1:21:00 · Brad and Bill 23.2% · guest 76.8%1:21:00 · Brad and Bill 23.2% · guest 76.8%1:24:00 · Brad and Bill 36% · guest 64%1:24:00 · Brad and Bill 36% · guest 64%1:27:00 · Brad and Bill 85.1% · guest 14.9%1:27:00 · Brad and Bill 85.1% · guest 14.9%
Sharpest disagreement ▶ 1:07:35 Dylan's Blunt Take on AMD's Software Inability

Dylan pulls no punches regarding AMD, directly stating they have no clue how to do software and refuse to build their own dogfooding clusters.

Hardest push from Brad and Bill ▶ 1:26:20 Brad Refuses Capex Disconnection from Revenue

Brad counters Dylan's pure arms-race capex narrative by insisting hyperscalers like Microsoft must anchor annual hardware purchases to actual inference revenue growth.

Biggest teaching moment ▶ 49:00 Dylan Breaks Down the 50x Reasoning Cost Math

Dylan breaks down how reasoning tokens combined with KV-cache context length constraints collapse concurrency, resulting in a 50x per-query cost spike.

Brad and Bill hold their own ▶ 41:20 Brad Dismantles the NVIDIA-Cisco Bubble Comparison

Brad demonstrates sharp market acumen by citing exact valuation multiples and balance sheet mechanics to refute lazy media comparisons to the 2000 dot-com peak.

the scores for every segment, with the reasoning behind each
ChapterTopicBrad and Bill as informed peerGuest teachingGuest disagreementBrad and Bill pushing backWhy
Dylan Patel's Origin Story in Semiconductors 4210 Brad introduces Dylan Patel and asks about his background. Dylan explains fixing his Xbox as an 8-year-old and reading semiconductor earnings reports as an intern. The tone is warm and conversational.
NVIDIA Market Dominance and Google's TPU Exception 5721 Bill asks about NVIDIA's workload share, assuming Google workloads are non-LLM. Dylan clarifies that Google runs 30% of production workloads and has run transformers since 2018 via BERT.
The 'Three-Headed Dragon': Why NVIDIA Dominates 6621 Bill and Brad ask about NVIDIA's competitive moat across software and hardware. Dylan outlines the 'three-headed dragon' (hardware, software, networking) and notes Google did rack-scale TPUs back in 2018 with Broadcom before NVIDIA's Blackwell.
NVIDIA's Supply Chain Strategy & Annual Release Cadence 5511 Brad and Bill explore how NVIDIA's annual hardware release cadence defends their lead. Dylan explains how Jensen operates on a paranoid, short-term deployment cycle that relentlessly pushes supply chain partners.
Inference vs. Training Moats and Performance TCO 6722 Bill asks about CUDA and training versus inference moats. Dylan distinguishes researcher-driven training from production inference where Microsoft hand-optimizes AMD hardware for fixed models.
Data Center Buildout: $1 Trillion AI Workloads and CPU Replacement Thesis 7632 Brad shows Jensen's chart claiming a $1T CPU data center replacement cycle. Dylan pushes back slightly, pointing out mainframes and classic CPUs continue growing, but notes upgrading dense modern CPUs effectively creates power headroom for AI servers.
Analyzing Satya Nadella's Power and Data Center Bottleneck Comments 6721 Bill and Brad bring up Ilya Sutskever's claim that internet pre-training data is tapped out. Dylan explains how synthetic data generation and functional verification provide a brand-new scaling vector.
Functional Verification: Where Synthetic Data Works Best 6622 Brad asks why hyperscalers are building 300k chip clusters if pre-training is stalling. Dylan explains that logarithmic scaling laws and massive synthetic rollout filtering still require massive compute capacity.
Objective vs. Subjective Grading in AI Training Models 5611 Bill asks for examples of where synthetic data works versus fails. Dylan distinguishes objective domains like code compilation and math from subjective domains like art and email writing.
Hyperscaler CapEx Acceleration and Multi-Gigawatt Data Center Buildouts 6721 Dylan notes Wall Street capex estimates are far too low based on SemiAnalysis tracking multi-gigawatt buildouts. Bill asks about distributed training, and Dylan highlights Microsoft's fiber investments enabling split workloads.
Evaluating the NVIDIA vs. Cisco 2000 Bubble Comparison 8411 Brad forcefully challenges the CNBC Cisco 2000 comparison by citing NVIDIA's 30x PE versus Cisco's 120x PE. Dylan agrees, emphasizing that current capex comes from profitable hyperscalers rather than debt-financed telecom speculation.
Inference-Time Reasoning Economics and the 50x Compute Cost Increase 6821 Dylan details the economics of inference-time reasoning, demonstrating that generating 10,000 hidden reasoning tokens plus smaller batch sizes increases compute cost per user query by 50x.
Enterprise Value Proposition and ROI of Reasoning Models 6611 Brad and Dylan discuss how enterprises will gladly pay a 50x token premium because replacing or augmenting a $300k white-collar engineer provides enormous ROI.
Model Bifurcation: Frontier Model Premium vs. Small Model Commoditization 7622 Bill compares AI infrastructure commoditization to Sun/Oracle in dot-com. Dylan explains the bifurcation: small open-source models face brutal price wars, while frontier models command high margins.
The High Bandwidth Memory (HBM) Boom & Market Dynamics 6721 Brad asks about High Bandwidth Memory (HBM). Dylan explains that HBM from SK Hynix is NVIDIA's largest COGS item above TSMC wafer costs and details Samsung's struggles in high-end HBM.
Alternative AI Hardware: AMD's Prospects, MI300X, and Software Bottlenecks 6732 Bill asks how AMD is performing. Dylan bluntly states AMD has 'no clue how to do software' and refuses to invest in internal dogfooding supercomputers, leaving them dependent on hyperscaler assistance.
Google TPU Architecture, System Maturity, and Cloud Rental Strategy 6611 Dylan explains Google's TPU system maturity and liquid cooling advantages, but notes Google fails to commercialize TPUs because of restrictive internal software, opaque cloud pricing, and high internal demand.
Amazon's Trainium 2 ('Amazon Basics TPU') and the 400k Chip Cluster 6721 Dylan nicknames Amazon Trainium 2 the 'Amazon Basics TPU,' highlighting its cost-effectiveness in HBM per dollar and Anthropic's planned 400k-chip supercomputer cluster.
Broadcom & Custom ASIC Landscape for 2025/2026 6611 Brad asks about Broadcom's run-up versus NVIDIA. Dylan details Broadcom's custom ASIC pipeline with Meta, Apple, and OpenAI, alongside their critical networking switch dominance.
2025–2026 Semiconductor Market Outlook & NeoCloud Consolidation 6721 Dylan forecasts massive 2025 capex growth followed by a potential 2026 reckoning, while predicting severe consolidation among the 80 existing NeoClouds as GPU rental prices collapse.
Revenue Growth Trajectory, Model Scaling & Capital Competition 7622 Brad and Bill push back on endless spending by citing Satya's discipline of tying capex to real revenue. Dylan argues that fear of being outscaled by competitors will drive hyperscalers to spend ahead anyway.

Statements from this episode (62)

Assertion Not checkable as stated
Patel: SemiAnalysis tracks all 1,500 global semiconductor fabs
“We track all 1500 fabs in the world. For your purposes, only 50 of them matter, but like, you know, all 1500 fabs around the world.”
Dylan Patel Dec 23, 2024 ▶ 3:01
Assertion Not checkable as stated
Patel: NVIDIA holds 98% of AI workloads outside Google, 70% overall
“So I would say if you ignored Google, it would be over 98%. But then when you bring Google into the mix, it's actually more like 70 because Google is really that large a percentage of AI workloads especially production workloads.”
Dylan Patel Dec 23, 2024 ▶ 4:32
Assertion Supported
Patel: Google has run transformers in Search workloads since 2018
“Google was running transformers even in their search workload since 2018, 2019. The advent of BERT, which was one of the most most well-known, most popular transformers before we got to the GPT madness is, has been their, in their production search workloads f…”
Dylan Patel Dec 23, 2024 ▶ 5:36
Assertion Not checkable as stated
Patel: Apple rents Google's TPU silicon, but GCP AI rentals remain GPUs
“While they do have some customers for their internal silicon externally, such as Apple the vast majority of their external rental business for AI in terms of cloud business is still GPUs.”
Dylan Patel Dec 23, 2024 ▶ 6:42
Opinion
Patel: Every semiconductor company except NVIDIA is terrible at software
“I would say every semiconductor company in the world sucks at software except for NVIDIA, right?”
Dylan Patel Dec 23, 2024 ▶ 7:07
Assertion Not checkable as stated
Patel: NVIDIA ships chips from design to deployment faster than competitors
“They get chips out faster than other people from Thought design to deployed.”
Dylan Patel Dec 23, 2024 ▶ 7:26
Assertion Open · timeframe Dec 2024
Patel: Google built rack-scale AI systems with Broadcom in 2018 before NVIDIA
“Google actually did this alongside Broadcom you know, and they did it before Nvidia, right? You know, today everyone's freaking out about, or not freaking out, but like everyone's like very excited about Nvidia's Blackwell system, right? It is a rack Of GPUs. …”
Dylan Patel Dec 23, 2024 ▶ 9:24
Insight
Patel: Semiconductor companies lack the engineers needed for rack-scale systems
“Building a chip is one thing, but building many chips that connect together, cooling them appropriately, networking them together, making sure that it's reliable at that scale is, is a whole host of problems that semiconductor companies don't have the engineer…”
Dylan Patel Dec 23, 2024 ▶ 10:44
Opinion
Patel: NVIDIA's primary differentiation comes from deep supply chain integration
“I would say for differentiating, NVIDIA has primarily focused on supply chain things, which, you know, might sound like, oh, well like, yeah, they're just like ordering stuff. No, no, no, no. You have to work deeply with the supply chain to build the next gene…”
Dylan Patel Dec 23, 2024 ▶ 11:07
Assertion Not checkable as stated
Patel: NVIDIA's Jensen Huang only plans 12 to 18 months ahead
“Well, the funny thing is a lot of people at NVIDIA will say Jensen doesn't plan more than a year or year and a half out. Because they change things and they'll deploy them out that fast, right? No semi, every other semiconductor company takes years to deploy, …”
Dylan Patel Dec 23, 2024 ▶ 12:58
Insight
Patel: NVIDIA's inference moat relies on hardware rather than software
“NVIDIA's moat in, in inference is actually A lot smaller on software but it's a lot bigger on, hey, they just have the best hardware.”
Dylan Patel Dec 23, 2024 ▶ 13:53
Assertion Partly supported
Patel: NVIDIA is cutting Blackwell margins to compete with custom ASICs
“Like with Blackwell, not only is it way, way, way faster, anywhere from 10 to 15 times on really large models for inference, because they've optimized it for very large language models, they've also decided, hey, we're gonna cut our margin, too, somewhat, beca…”
Dylan Patel Dec 23, 2024 ▶ 14:16
Prediction Not checkable as stated
Patel: Tanking LLM delivery costs will induce demand for compute
“The cost for delivering LLMs is Is tanking, which is going to induce demand, right?”
Dylan Patel Dec 23, 2024 ▶ 15:05
Assertion Supported
Patel: Microsoft runs GPT models on AMD hardware for inference
“Microsoft has deployed GPT-style models on On other competitors' hardware, such as AMD they've, and some of their own, but mostly AMD, and so they can wring that out with software because they can spend hundreds of engineers, dozens of engineers' hours hundred…”
Dylan Patel Dec 23, 2024 ▶ 16:50
Assertion Supported
Gerstner: Jensen projected $1T in new AI and CPU replacement workloads
“And for the first time, he said, not only are we going to have a trillion dollars of new AI workloads over the course of the next four years he said, but we're also going to have a trillion dollars of CPU replacement, of data center replacement workloads over …”
Brad Gerstner Dec 23, 2024 ▶ 17:31
Assertion Supported
Patel: IBM mainframes increase in volume and revenue every cycle
“IBM mainframe sell more volume and revenue every single cycle.”
Dylan Patel Dec 23, 2024 ▶ 19:02
Assertion Not checkable as stated
Patel: Most AWS data center CPUs are Intel chips from 2015-2020
“The plurality of Amazon's CPUs in their data centers Are 24 core Intel CPUs from, that were manufactured from 2015 to twenty-twenty.”
Dylan Patel Dec 23, 2024 ▶ 20:55
Insight
Patel: Consolidating legacy CPU servers frees up power for AI workloads
“If I just replace, like, six servers with one, I've basically invented power out of thin air, right? I mean, like, you know, in effect, because these old servers, which are six plus years old, or even, you know, they can just be deprecated and put, so with Cap…”
Dylan Patel Dec 23, 2024 ▶ 21:11
Opinion
Gerstner: Data centers and power, not GPUs, are the real bottleneck
“What I think it was more a assessment on the real bottleneck, which is data centers and power as opposed to GPUs because GPUs have come online.”
Brad Gerstner Dec 23, 2024 ▶ 22:13
Assertion Not checkable as stated
Patel: AI training has barely tapped vast video data reserves
“We have barely, barely, barely tapped video data, right? So there is a significant amount of data that's not tapped, it's just video data Is so much more information than written data.”
Dylan Patel Dec 23, 2024 ▶ 24:06
Insight
Patel: Synthetic data generation enables continued AI scaling despite data limits
“You can create data out of thin air almost, right? In certain domains, right? And so this is the whole, the debate around scaling laws is how can we create data?”
Dylan Patel Dec 23, 2024 ▶ 25:25
Assertion Not checkable as stated
Patel: AI industry is in early days of synthetic data
“Where have we gone on synthetic data? Oh, we're still like very early days, right? We've spent tens of millions of dollars maybe on synthetic data.”
Dylan Patel Dec 23, 2024 ▶ 28:08
Insight
Patel: Synthetic training only works in functionally verifiable domains like math
“We can't teach it what good art is. Because we have no way to functionally prove what good art is. We can teach it to write really good software. We can teach it how to do mathematical proofs. We can teach it how to engineer systems, because there are, while t…”
Dylan Patel Dec 23, 2024 ▶ 28:58
Prediction Not checkable as stated
Patel: AI models may improve faster over the next 6-12 months
“We may actually see models improve faster in the next six months to a year than we saw them improve in the last year. Because there's this new axis of synthetic data generation and the amount of compute we can throw at it is, we're still right here in the scal…”
Dylan Patel Dec 23, 2024 ▶ 33:00
Assertion Not checkable as stated
Patel: AI labs have tapped out human-written internet text data
“Humans post on the internet every day, and we've already tapped that out, right? Kind of more or less on a text.”
Dylan Patel Dec 23, 2024 ▶ 35:05
Opinion
Patel: Wall Street hyperscaler CapEx estimates are far too low
“I think when you look at the streets estimates for capex, they're all far too low.”
Dylan Patel Dec 23, 2024 ▶ 36:37
Insight
Patel: Multi-gigawatt buildouts disprove claims that AI scaling is over
“Why is Mark Zuckerberg building a two gigawatt data center in Louisiana? Why is Amazon building these multi gigawatt data centers? Why is Google, why is Microsoft building multiple gigawatt data centers? Plus buying billions and billions of dollars of fiber to…”
Dylan Patel Dec 23, 2024 ▶ 37:39
Insight
Patel: Modern AI training requires more inference compute than weight updates
“In fact, there's more inference in training than there is updating the model weights, because you have to generate hundreds of possibilities And then, oh, you only train on a couple of them, right?”
Dylan Patel Dec 23, 2024 ▶ 39:14
Insight
Patel: AI pre-training gains are becoming logarithmically more expensive
“So, the whole paradigm of training, you know, pre-training is, is, is not slowing down. It's just, it's logarithmically more expensive each, for each generation, for each incremental improvement.”
Dylan Patel Dec 23, 2024 ▶ 40:22
Assertion Supported
Gerstner: NVIDIA trades around 30 P/E compared to Cisco's 2000 peak
“Cisco, you know, 2000 and we'll show it on the pod, but you know, they peaked at like a 120 PE. Right. And yeah, you know, if you look at the fall off that occurred in revenue and in EBITDA, you know, and then it had 70% compression in the priced earnings mult…”
Brad Gerstner Dec 23, 2024 ▶ 41:18
Assertion Supported
Patel: Cisco was fueled by speculative debt whereas NVIDIA is cash-backed
“Cisco's revenue, a lot of it was funded through private-slash-credit investments into building out telecom infrastructure, right? When we look at NVIDIA's revenue sources, very little of it is private-slash-credit, right? And in some cases, yes, it's private-s…”
Dylan Patel Dec 23, 2024 ▶ 42:35
Assertion Contradicted
Patel: Inflation-adjusted private capital at dot-com peak exceeded today's levels
“At the peak of the dot com, you know, especially once you inflation adjust it the private capital entering the space was much larger than it is today”
Dylan Patel Dec 23, 2024 ▶ 43:01
Opinion
Gurley: Corporate America invests more in AI than during the internet boom
“I think corporate America is investing more in AI and with more conviction than they did even in the internet wave also.”
Bill Gurley Dec 23, 2024 ▶ 43:56
Assertion Supported
Patel: GPT-4 cost hundreds of millions to train, generates billions in revenue
“Hundreds of millions of dollars to train GPT-IV. And it's generating billions of dollars of revenue.”
Dylan Patel Dec 23, 2024 ▶ 45:26
Insight
Patel: Reasoning models like OpenAI o1 increase compute costs by 50x
“When I do this with O-one, right, because it's doing that thinking phase of 10,000... It spends a lot of memory on generating this KV cache and reading this KV cache constantly. Now the maximum batch size, i.e. Concurrent users I can have, is a fraction of tha…”
Dylan Patel Dec 23, 2024 ▶ 49:24
Insight
Patel: Expensive reasoning queries unlock pricing power by automating new human tasks
“Yes, the queries are expensive, but they're nothing close to the human, right? And so each level of productivity gain I get each level of capabilities jump is a whole new class of tasks that it can do And therefore I can charge for that. Right. So this is the …”
Dylan Patel Dec 23, 2024 ▶ 51:39
Prediction Held up
Patel: Google and Anthropic will both release reasoning models soon
“There's a Google model that is doing reasoning right now, and it's not released yet, but it's gonna be released soon enough, right? Anthropic is going to release a reasoning model.”
Dylan Patel Dec 23, 2024 ▶ 52:27
Prediction Not checkable as stated
Patel: Reasoning models will see humongous performance gains within a year
“And so this, the performance improvements we'll get out of these models is, is humongous, right? In, in the coming, you know, six months to a year in certain benchmarks where you have functional verifiers.”
Dylan Patel Dec 23, 2024 ▶ 53:01
Assertion Not checkable as stated
Patel: Microsoft earns 50% to 70% gross margins on OpenAI models
“Microsoft's earning 50 to 70% gross margins on OpenAI models, and that's with the profit share they get to get, or the share that they give OpenAI, right?”
Dylan Patel Dec 23, 2024 ▶ 55:53
Assertion Not checkable as stated
Patel: Anthropic showed ~70% gross margins in its most recent round
“Or, you know, Anthropic, similarly, in their most recent round, they were showing, like, 70% gross margins.”
Dylan Patel Dec 23, 2024 ▶ 56:01
Assertion Contradicted
Patel: A quarter of Nvidia shipments could provide Llama-7B to humanity
“If you take one quarter of NVIDIA shipments and you said all of them are going to inference Lama-Seven-B, you can give every single person on earth a hundred tokens per minute, right? Or sorry, a hundred tokens per second. You give every single person on earth…”
Dylan Patel Dec 23, 2024 ▶ 59:42
Assertion Not checkable as stated
Patel: SK Hynix memory is NVIDIA's largest COGS item, surpassing TSMC
“When you look at the cost of goods sold of NVIDIA their highest cost of goods sold is not TSMC, which is a Thing that people don't realize. It's actually HBM memory primarily SK Hynix.”
Dylan Patel Dec 23, 2024 ▶ 1:02:27
Assertion Partly supported
Patel: Samsung holds almost zero share in HBM memory, especially at NVIDIA
“In HBM, Samsung has almost no share, right? Especially at NVIDIA”
Dylan Patel Dec 23, 2024 ▶ 1:03:16
Assertion Not checkable as stated
Patel: Standard high-end server memory yields higher gross margins than HBM
“The gross margins on HBM have not been fantastic. They've been good, but they haven't been fantastic. Actually, regular memory, high-end, like, server memory that is not HBM is actually higher gross margin than HBM.”
Dylan Patel Dec 23, 2024 ▶ 1:05:08
Assertion Partly supported
Patel: AWS Trainium2 offers highest HBM capacity and bandwidth per dollar
“Their whole thing at reInvent, if you really talk to them when they announced Trinium II and our whole post about it and our analysis of it is, like, supply chain-wise, this is, looks, you know, you squint your eyes, this looks like an Amazon Basics TPU, right…”
Dylan Patel Dec 23, 2024 ▶ 1:06:02
Opinion
Patel: AMD lacks software talent and refuses to fund internal GPU clusters
“AMD is really good, but they're missing software. AMD has no clue how to do software, I think. They've got very few developers on it. They won't spend the money to build a GPU cluster for themselves so that they can develop software.”
Dylan Patel Dec 23, 2024 ▶ 1:07:36
Prediction Not checkable as stated
Patel: AMD will see less AI revenue from Microsoft and Meta in 2025
“Yes, I think they'll have they'll have a lot less success with Microsoft than they did this year. And they'll have less success than they did with Meta than they did this year.”
Dylan Patel Dec 23, 2024 ▶ 1:10:01
Assertion Supported
Patel: Google TPU clusters scale up to 8,000 chips today
“Nvidia's talking about GB 200, NVL 72, TPUs go to 8000 today, right?”
Dylan Patel Dec 23, 2024 ▶ 1:11:30
Assertion Not checkable as stated
Patel: Initial NVIDIA GPU cloud deployments suffer a 5% failure rate
“Google's brought in a level of reliability that NVIDIA GPUs don't have. You know, the dirty secret is to go ask people what the reliability rate of GPUs is in the cloud or in a deployment. It's like, oh God, it is not, they're reliable-ish, like, but like, esp…”
Dylan Patel Dec 23, 2024 ▶ 1:11:54
Assertion Not checkable as stated
Patel: Apple accounts for over 70% of Google's TPU rental revenue
“There's only one company accounts for over 70% of Google's revenue from TPUs as far as I understand, and that's Apple.”
Dylan Patel Dec 23, 2024 ▶ 1:14:35
Assertion Not checkable as stated
Patel: Amazon Trainium 2 costs ~$5,000 per chip versus $30,000+ for NVIDIA
“You're not paying, you know, north of 30,000, you know, 40,000 dollars per chip for the server, you're paying, you know, significantly less, right? 5000 dollars per chip, right?”
Dylan Patel Dec 23, 2024 ▶ 1:16:23
Assertion Supported
Patel: Amazon and Anthropic are building a 400,000-chip Trainium supercomputer
“Amazon and Anthropic have decided to, you know, make a 400,000 tranium server supercomputer, right? 400,000 chips, right?”
Dylan Patel Dec 23, 2024 ▶ 1:16:50
Assertion Supported
Patel: Broadcom holds custom ASIC wins with Meta, OpenAI, and Apple
“Broadcom does have multiple custom ASIC wins, right? It's not just Google here. Meta's, Meta's ramping up mostly still for recommendation systems, but their custom chips are gonna get better. You know, there's other players like OpenAI who are making a chip, r…”
Dylan Patel Dec 23, 2024 ▶ 1:18:20
Assertion Supported
Patel: Broadcom is building an NVSwitch competitor for AMD and others
“Broadcom is making a Competitors to that, that they will cede to the market, right? Multiple companies will be using that. Not just, you know, AMD will be using that competitor to NVSwitch, but they're not making it themselves because they don't have the skill…”
Dylan Patel Dec 23, 2024 ▶ 1:19:47
Prediction Not checkable as stated
Patel: Google TPU purchases will pause for six months over space limits
“Like in the next six months there is a bit of a slowdown in Google TPU purchases because they have no data center space. They want more. They just literally have no data center space to put them.”
Dylan Patel Dec 23, 2024 ▶ 1:20:28
Prediction Held up
Patel: Hyperscalers will significantly boost capex in 2025, lifting chip suppliers
“The plans for hyperscalers are pretty firm on, they're gonna spend a crap load more next year, right? And therefore, the ecosystem of networking players, of ASIC vendors, of systems vendors is gonna do well, whether it be NVIDIA or Marvell or Broadcom or AMD o…”
Dylan Patel Dec 23, 2024 ▶ 1:21:20
Prediction Not checkable as stated
Patel: Only 5 to 10 of the 80 GPU NeoClouds will survive
“80 NeoClouds are not going to survive. Maybe five to 10 will. And that's because five of those are sovereign, right? And then the other five are like actually like market competitive.”
Dylan Patel Dec 23, 2024 ▶ 1:22:59
Assertion Not checkable as stated
Patel: Hyperscalers represent 50% to 60% of AI chip revenue
“Roughly you can say hyperscalers are 50 ish percent of revenue, 50 to 60%, and the rest of it is Neo cloud slash sovereign AI.”
Dylan Patel Dec 23, 2024 ▶ 1:23:16
Assertion Supported
Patel: Nvidia Blackwell costs over twice as much to manufacture as Hopper
“The cost to make Blackwell is north of two X that of the cost to make Hopper, right?”
Dylan Patel Dec 23, 2024 ▶ 1:24:07
Prediction Not checkable as stated
Patel: Meta and Microsoft may take free cash flows close to zero
“I think Meta and Microsoft may even take their free cash flows close to zero and just spent.”
Dylan Patel Dec 23, 2024 ▶ 1:24:40
Insight
Gerstner: AI infrastructure capex growth requires matching 30% inference revenue growth
“So, you know, if you think that infrastructure expenses are going to grow at 30% a year, then I think you have to believe that the underlying inference revenues, right, both on the consumer side and the enterprise side are going to grow somewhere in that range…”
Brad Gerstner Dec 23, 2024 ▶ 1:27:27
Assertion Not checkable as stated
Gerstner: Lower-tier AI model makers are abandoning the compute arms race
“I think you already see some of these smaller second and third tier models, changing business model, falling aside, no longer engaged in the arms race.”
Brad Gerstner Dec 23, 2024 ▶ 1:29:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 40 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.