Dec 5, 2023 · 1h 7m · latent-space

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis

Dylan Patel · 54m spoken Shawn Wang · 3m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Dylan Patel, Chief Analyst at SemiAnalysis, joins the Latent Space Podcast to examine the economics, physical bottlenecks, supply chain geopolitics, and silicon architecture underpinning modern AI, detailing how compute disparity shapes strategies for both hyperscalers and resource-constrained developers.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 11.5% of the talking time here. How this is scored →

The hosts as informed peer 4.7 Guest teaching 5.8 Guest disagreement 3.0 The hosts pushing back 1.7
05100:0015:0030:0045:001:00:000:06–2:17 · The hosts as informed peer 4/10 Introductions and the Rise of SemiAnalysis Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention.2:17–5:04 · The hosts as informed peer 5/10 The Economics of AI Infrastructure and Training Costs Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs.5:04–10:54 · The hosts as informed peer 3/10 The GPU Poor vs. GPU Rich Landscape Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs.10:54–14:25 · The hosts as informed peer 5/10 Google TPUs, Framework Ecosystems, and Compilers Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack.14:25–21:38 · The hosts as informed peer 4/10 Compute Utilization Metrics: MFU in Training vs. MBU in Inference Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient.21:38–24:20 · The hosts as informed peer 5/10 Interconnect Scaling and Google's Partnership with Broadcom Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups.24:20–30:07 · The hosts as informed peer 5/10 Strategic Opportunities for GPU-Poor Developers and Startups Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor.30:07–37:46 · The hosts as informed peer 6/10 Model Depreciation, Latency Profiling, and Noam Shazeer's Vision Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds.37:46–46:41 · The hosts as informed peer 5/10 The AI Silicon Landscape and Nvidia's Formidable Moat Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat.46:41–53:49 · The hosts as informed peer 5/10 Frontier Lab Partnerships and Apple's GenAI Conundrum Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs.53:49–58:44 · The hosts as informed peer 4/10 AI Safety Debates and Semiconductor Supply Chain Fragility Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals.58:44–1:07:38 · The hosts as informed peer 5/10 Lightning Round, Research Workflow, and the Genie Question In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck.0:06–2:17 · Guest teaching 3/10 Introductions and the Rise of SemiAnalysis Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention.2:17–5:04 · Guest teaching 5/10 The Economics of AI Infrastructure and Training Costs Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs.5:04–10:54 · Guest teaching 6/10 The GPU Poor vs. GPU Rich Landscape Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs.10:54–14:25 · Guest teaching 6/10 Google TPUs, Framework Ecosystems, and Compilers Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack.14:25–21:38 · Guest teaching 8/10 Compute Utilization Metrics: MFU in Training vs. MBU in Inference Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient.21:38–24:20 · Guest teaching 6/10 Interconnect Scaling and Google's Partnership with Broadcom Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups.24:20–30:07 · Guest teaching 6/10 Strategic Opportunities for GPU-Poor Developers and Startups Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor.30:07–37:46 · Guest teaching 6/10 Model Depreciation, Latency Profiling, and Noam Shazeer's Vision Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds.37:46–46:41 · Guest teaching 8/10 The AI Silicon Landscape and Nvidia's Formidable Moat Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat.46:41–53:49 · Guest teaching 5/10 Frontier Lab Partnerships and Apple's GenAI Conundrum Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs.53:49–58:44 · Guest teaching 7/10 AI Safety Debates and Semiconductor Supply Chain Fragility Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals.58:44–1:07:38 · Guest teaching 4/10 Lightning Round, Research Workflow, and the Genie Question In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck.0:06–2:17 · Guest disagreement 2/10 Introductions and the Rise of SemiAnalysis Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention.2:17–5:04 · Guest disagreement 2/10 The Economics of AI Infrastructure and Training Costs Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs.5:04–10:54 · Guest disagreement 5/10 The GPU Poor vs. GPU Rich Landscape Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs.10:54–14:25 · Guest disagreement 2/10 Google TPUs, Framework Ecosystems, and Compilers Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack.14:25–21:38 · Guest disagreement 4/10 Compute Utilization Metrics: MFU in Training vs. MBU in Inference Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient.21:38–24:20 · Guest disagreement 2/10 Interconnect Scaling and Google's Partnership with Broadcom Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups.24:20–30:07 · Guest disagreement 2/10 Strategic Opportunities for GPU-Poor Developers and Startups Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor.30:07–37:46 · Guest disagreement 4/10 Model Depreciation, Latency Profiling, and Noam Shazeer's Vision Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds.37:46–46:41 · Guest disagreement 5/10 The AI Silicon Landscape and Nvidia's Formidable Moat Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat.46:41–53:49 · Guest disagreement 3/10 Frontier Lab Partnerships and Apple's GenAI Conundrum Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs.53:49–58:44 · Guest disagreement 4/10 AI Safety Debates and Semiconductor Supply Chain Fragility Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals.58:44–1:07:38 · Guest disagreement 1/10 Lightning Round, Research Workflow, and the Genie Question In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck.0:06–2:17 · The hosts pushing back 1/10 Introductions and the Rise of SemiAnalysis Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention.2:17–5:04 · The hosts pushing back 2/10 The Economics of AI Infrastructure and Training Costs Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs.5:04–10:54 · The hosts pushing back 1/10 The GPU Poor vs. GPU Rich Landscape Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs.10:54–14:25 · The hosts pushing back 2/10 Google TPUs, Framework Ecosystems, and Compilers Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack.14:25–21:38 · The hosts pushing back 1/10 Compute Utilization Metrics: MFU in Training vs. MBU in Inference Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient.21:38–24:20 · The hosts pushing back 2/10 Interconnect Scaling and Google's Partnership with Broadcom Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups.24:20–30:07 · The hosts pushing back 1/10 Strategic Opportunities for GPU-Poor Developers and Startups Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor.30:07–37:46 · The hosts pushing back 4/10 Model Depreciation, Latency Profiling, and Noam Shazeer's Vision Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds.37:46–46:41 · The hosts pushing back 2/10 The AI Silicon Landscape and Nvidia's Formidable Moat Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat.46:41–53:49 · The hosts pushing back 2/10 Frontier Lab Partnerships and Apple's GenAI Conundrum Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs.53:49–58:44 · The hosts pushing back 1/10 AI Safety Debates and Semiconductor Supply Chain Fragility Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals.58:44–1:07:38 · The hosts pushing back 1/10 Lightning Round, Research Workflow, and the Genie Question In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 45.7% · guest 54.3%0:00 · the hosts 45.7% · guest 54.3%3:00 · the hosts 18.3% · guest 81.7%3:00 · the hosts 18.3% · guest 81.7%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 22.2% · guest 77.8%9:00 · the hosts 22.2% · guest 77.8%12:00 · the hosts 5.5% · guest 94.5%12:00 · the hosts 5.5% · guest 94.5%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 10.8% · guest 89.2%21:00 · the hosts 10.8% · guest 89.2%24:00 · the hosts 15.6% · guest 84.4%24:00 · the hosts 15.6% · guest 84.4%27:00 · the hosts 2.5% · guest 97.5%27:00 · the hosts 2.5% · guest 97.5%30:00 · the hosts 30.1% · guest 69.9%30:00 · the hosts 30.1% · guest 69.9%33:00 · the hosts 11.6% · guest 88.4%33:00 · the hosts 11.6% · guest 88.4%36:00 · the hosts 14.7% · guest 85.3%36:00 · the hosts 14.7% · guest 85.3%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 3.9% · guest 96.1%45:00 · the hosts 3.9% · guest 96.1%48:00 · the hosts 6.3% · guest 93.7%48:00 · the hosts 6.3% · guest 93.7%51:00 · the hosts 9.5% · guest 90.5%51:00 · the hosts 9.5% · guest 90.5%54:00 · the hosts 13.7% · guest 86.3%54:00 · the hosts 13.7% · guest 86.3%57:00 · the hosts 10.2% · guest 89.8%57:00 · the hosts 10.2% · guest 89.8%1:00:00 · the hosts 29% · guest 71%1:00:00 · the hosts 29% · guest 71%1:03:00 · the hosts 7.6% · guest 92.4%1:03:00 · the hosts 7.6% · guest 92.4%1:06:00 · the hosts 9.6% · guest 90.4%1:06:00 · the hosts 9.6% · guest 90.4%
Sharpest disagreement ▶ 30:55 Dismissing fine-tuning of smaller models

Patel rejects the conventional industry enthusiasm for fine-tuning 7B models, bluntly calling it a useless waste of time compared to utilizing larger models at decent batch sizes.

Hardest push from the hosts ▶ 31:44 Swyx challenges GPT-3.5 sizing and sparse routing

Swyx presses Patel on the exact parameter scale of GPT-3.5 versus Falcon 140B and questions whether the model is smaller or sparsely routed.

Biggest teaching moment ▶ 17:35 Detailed derivation of LLM inference bandwidth requirements

Patel walks through the exact math showing why reading 70 billion parameters at human reading speed requires over 2 TB/s of bandwidth, exposing the inefficiency of standard inference libraries.

The host holds their own ▶ 30:07 Alessio and Swyx analyze model depreciation velocity

Alessio and Swyx frame the economic dilemma of investing heavily in model training when open-source architecture advances make models obsolete within three months.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and the Rise of SemiAnalysis 4321 Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention.
The Economics of AI Infrastructure and Training Costs 5522 Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs.
The GPU Poor vs. GPU Rich Landscape 3651 Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs.
Google TPUs, Framework Ecosystems, and Compilers 5622 Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack.
Compute Utilization Metrics: MFU in Training vs. MBU in Inference 4841 Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient.
Interconnect Scaling and Google's Partnership with Broadcom 5622 Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups.
Strategic Opportunities for GPU-Poor Developers and Startups 5621 Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor.
Model Depreciation, Latency Profiling, and Noam Shazeer's Vision 6644 Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds.
The AI Silicon Landscape and Nvidia's Formidable Moat 5852 Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat.
Frontier Lab Partnerships and Apple's GenAI Conundrum 5532 Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs.
AI Safety Debates and Semiconductor Supply Chain Fragility 4741 Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals.
Lightning Round, Research Workflow, and the Genie Question 5411 In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck.

Statements from this episode (34)

Assertion Supported
Patel: SemiAnalysis reported GPT-4 mixture of experts architecture in January
“Just being clear, I talked about mixture of experts in January, it's just people didn't really notice it.”
Dylan Patel Dec 5, 2023 ▶ 1:08
Opinion
Patel: AWS holds major infrastructure efficiency advantage over Azure and GCP
“Behind the scenes, they've done a lot on the infrastructure that is super custom, That Microsoft, Azure, and Google Cloud just don't even match in terms of efficiency. Like, if you think about the cost to rent out SSD space, so the cost to rent, you know, offe…”
Dylan Patel Dec 5, 2023 ▶ 3:04
Prediction Not checkable as stated
Patel: AI software will have lower human R&D costs but higher COGS than SaaS
“An AI software looks like it's going to be very different in my opinion, right? Like the R&D cost is much lower in terms of people but the cost of goods sold in terms of actually operating the service, I think will be much higher, right?”
Dylan Patel Dec 5, 2023 ▶ 4:12
Opinion
Patel: Frontier AI model training costs are effectively irrelevant
“In my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-IV, right? Like 20,000 A-one hundreds. That's like, I know it sounds like a lot of money.”
Dylan Patel Dec 5, 2023 ▶ 4:34
Assertion Supported
Patel: Nvidia manufactured 400k H100s last quarter and will sell 530k this quarter
“There's 400 to 500,000 being, 400,000 manufactured last quarter, and like five 30,000 this quarter being sold, right, of H-Hundreds”
Dylan Patel Dec 5, 2023 ▶ 7:03
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Dylan Patel Dec 5, 2023 ▶ 7:46
Opinion
Patel: TruthfulQA is a garbage benchmark that gets gamed
“Truthful QA is a garbage benchmark, like, you, like, some of the models that are very high on there, if you use it for five seconds, you're like, this is garbage, right?”
Dylan Patel Dec 5, 2023 ▶ 9:41
Prediction Not checkable as stated
Patel: Google will have more compute than any company by far
“Google is gonna have more compute than any other company in the world, period, by, like, a large, large factor.”
Dylan Patel Dec 5, 2023 ▶ 9:59
Assertion Supported
Patel: Google TPU v5e is half the size of TPU v5p
“TPUv-V-V-E is, like the new one, but it's mostly, mostly an inference chip. It's a small chip. It's a little bit, it's about half the size of a TPUv-V-V. That chip, you know, you can get very good performance on, like, Of Lama, 70 B inference, right?”
Dylan Patel Dec 5, 2023 ▶ 11:49
Opinion
Patel: Hardware vendors and developers are coalescing around OpenAI's Triton
“Likewise, there's OpenAI's Triton, like, what they're trying to do there, and like you know, everyone's really coalescing around Triton. You know, people, you know, third-party hardware vendors”
Dylan Patel Dec 5, 2023 ▶ 13:40
Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Dylan Patel Dec 5, 2023 ▶ 14:40
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”
Dylan Patel Dec 5, 2023 ▶ 18:40
Assertion Supported
Patel: Google is co-designing TPU v5 and TPU v6 with Broadcom
“TPU, you know, all the way through to all, TPU v-five, which is the one they're in production of now, and six, and, you know, all these are all gonna be you know, co-designed with Broadcom.”
Dylan Patel Dec 5, 2023 ▶ 22:25
Assertion Not checkable as stated
Patel: Google targets 2027 to exit Broadcom partnership after 2025 attempt failed
“Their latest target to get away from Broadcom is twenty-twenty-seven, right? But like, you know, that's four years from now. Chip design cycle is four years. So they already tried to get away in twenty-twenty-five, and that failed.”
Dylan Patel Dec 5, 2023 ▶ 23:01
Assertion Supported
Patel: Broadcom is working on Meta's second-generation internal AI chip
“So Broadcom is, is, is, is working with, on Google's chip right now, but of course on Meta's, Meta's internal AI chip, which they're on the second generation of, working on that.”
Dylan Patel Dec 5, 2023 ▶ 23:53
Insight
Patel: AI chip bottlenecks lie in interconnects and memory, not matrix multiplication
“Multiplying tensors is kind of, you know, anyone can, there's a lot of people who've made good matrix multiply units, right? But it's about, like, getting good utilization out of those, and interfacing with the memory, and interfacing with other chips really e…”
Dylan Patel Dec 5, 2023 ▶ 24:06
Assertion Supported
Patel: Running LLaMA-70B at reading speed requires 2.1 TB/s memory bandwidth
“Hey, to run Llama's seventy billion requires two terabytes a second of memory bandwidth, 2.1, at reading, human reading speed.”
Dylan Patel Dec 5, 2023 ▶ 26:06
Prediction Didn’t hold up
Patel: Google will avoid deploying local laptop models to retain control
“I don't think Google is going to deploy a model that I can run on my laptop to help me with code or help me with, you know, XYZ. They're always going to want to run it on the cloud for control.”
Dylan Patel Dec 5, 2023 ▶ 26:44
Assertion Not checkable as stated
Patel: Several companies make tens of millions from adult AI models
“I think there's a couple companies who make, Tens of millions of dollars of revenue from, yeah, from LLMs or diffusion models for porn”
Dylan Patel Dec 5, 2023 ▶ 29:48
Opinion
Patel: Fine-tuning existing small models for cloud use is useless
“Unless, unless you're fine tuning for on device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time, right?”
Dylan Patel Dec 5, 2023 ▶ 30:56
Assertion Not checkable as stated
Patel: No open-source model is currently close to GPT-3.5
“And like, there's nothing open source that is anywhere close to 3.5 yet.”
Dylan Patel Dec 5, 2023 ▶ 31:33
Prediction Not checkable as stated
Patel: Roughly 10 companies have compute to beat GPT-4 within six months
“There's, like, 10 companies that have enough compute in one single data center to be able to beat GPT-IV, right? Like, straight up, like, if not today, within the next six months, right?”
Dylan Patel Dec 5, 2023 ▶ 33:51
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Dylan Patel Dec 5, 2023 ▶ 41:01
Prediction Held up
Patel: AMD MI300 will beat Nvidia H100 on paper within a quarter
“AMD. They have a GPU. MI 300. That will be better than the H 100 in a quarter or so. Now, that says nothing about how hard it is to program it, but at least hardware-wise, on paper, it's better.”
Dylan Patel Dec 5, 2023 ▶ 41:13
Assertion Not checkable as stated
Patel: Nvidia and Google control over 80% of advanced AI silicon manufacturing capacity
“To simplify it, NVIDIA has a little bit more than half, and Google has, like, 30%, right, through Broadcom. So it's like, the total capacity for everyone else is much lower, and they're all sharing it”
Dylan Patel Dec 5, 2023 ▶ 42:01
Assertion Not checkable as stated
Patel: AMD MI300 costs more than twice as much to manufacture as Nvidia H100
“In the case of amd their manufacturing costs for mi-thruhundred or more than twice that of h-one hundred and it only beats h-one hundred by a little bit from you know performance stuff i've seen”
Dylan Patel Dec 5, 2023 ▶ 42:52
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Dylan Patel Dec 5, 2023 ▶ 44:00
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Dylan Patel Dec 5, 2023 ▶ 46:55
Assertion Not checkable as stated
Patel: Microsoft's custom AI chip performs worse than Nvidia H100
“Microsoft's going to announce their chip soon. It's worse performance than the H-H-E-N-H-E-D but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.”
Dylan Patel Dec 5, 2023 ▶ 48:53
Opinion
Patel: US and China can each shut down the other's chip supply chain
“The U.S. Could absolutely shut down the Chinese semiconductor supply chain, they won't, but, and China could absolutely shut down the U.S. One, actually, by the way.”
Dylan Patel Dec 5, 2023 ▶ 55:58
Assertion Partly supported
Patel: All sub-7nm chips depend on tooling from one Austrian monopoly
“There is no chip that is less than seven nanometer that doesn't get touched by this one Austrian company's tool, right? And there is no alternative.”
Dylan Patel Dec 5, 2023 ▶ 56:19
Assertion Supported
Patel: TSMC's Arizona fab still depends on Taiwan for masks and shipping
“TSMC is building a fab in Arizona. It's quite a bit smaller than the fabs in, in, in Taiwan, but even ignoring that, those fabs still have to ship everything to Taiwan back anyways. And also they have to get what's called a mask from Taiwan and get sent to get…”
Dylan Patel Dec 5, 2023 ▶ 56:50
Assertion Contradicted
Patel: Large-scale AI training currently requires a single data center
“Everything that we've seen so far is that large-scale training has to happen in an individual data center with very high-speed networking.”
Dylan Patel Dec 5, 2023 ▶ 1:06:07
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Dylan Patel Dec 5, 2023 ▶ 1:06:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.