Dec 5, 2023 · 1h 7m · latent-space
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Dylan Patel, Chief Analyst at SemiAnalysis, joins the Latent Space Podcast to examine the economics, physical bottlenecks, supply chain geopolitics, and silicon architecture underpinning modern AI, detailing how compute disparity shapes strategies for both hyperscalers and resource-constrained developers.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 11.5% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Patel rejects the conventional industry enthusiasm for fine-tuning 7B models, bluntly calling it a useless waste of time compared to utilizing larger models at decent batch sizes.
Hardest push from the hosts ▶ 31:44 Swyx challenges GPT-3.5 sizing and sparse routingSwyx presses Patel on the exact parameter scale of GPT-3.5 versus Falcon 140B and questions whether the model is smaller or sparsely routed.
Biggest teaching moment ▶ 17:35 Detailed derivation of LLM inference bandwidth requirementsPatel walks through the exact math showing why reading 70 billion parameters at human reading speed requires over 2 TB/s of bandwidth, exposing the inefficiency of standard inference libraries.
The host holds their own ▶ 30:07 Alessio and Swyx analyze model depreciation velocityAlessio and Swyx frame the economic dilemma of investing heavily in model training when open-source architecture advances make models obsolete within three months.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and the Rise of SemiAnalysis | 4 | 3 | 2 | 1 | Swyx introduces Dylan Patel and mentions his own background covering semiconductors. Patel gives a brief, polite correction noting he had discussed Mixture of Experts months earlier before it gained viral attention. | |
| The Economics of AI Infrastructure and Training Costs | 5 | 5 | 2 | 2 | Swyx references hedge fund semi coverage and quotes Patel's provocative thesis on training costs. Patel explains the shifting cost structure of AI software from high R&D to high COGS and infrastructure operating costs. | |
| The GPU Poor vs. GPU Rich Landscape | 3 | 6 | 5 | 1 | Patel delivers an extended analysis on GPU allocations, calling SF startup GPU boasting detached from supply realities. He forcefully criticizes Hugging Face benchmarks like TruthfulQA as garbage and dismisses batch size 1 inference on high-end GPUs. | |
| Google TPUs, Framework Ecosystems, and Compilers | 5 | 6 | 2 | 2 | Alessio brings up previous guest Chris Lattner and asks about PyTorch vs JAX/XLA on Google hardware. Patel details the distinction between TPU v5 and v5e and how lower-level compilers like Triton and Palace fit into the stack. | |
| Compute Utilization Metrics: MFU in Training vs. MBU in Inference | 4 | 8 | 4 | 1 | Swyx prompts Patel to define hardware utilization metrics. Patel gives an in-depth lecture on arithmetic intensity, contrasting training MFU with inference MBU, while calling existing open source inference implementations extremely inefficient. | |
| Interconnect Scaling and Google's Partnership with Broadcom | 5 | 6 | 2 | 2 | Swyx identifies networking as the slowest scaling dimension in hardware systems. Patel explains Broadcom's critical role in Google TPU co-design and why interconnect engineering makes custom silicon difficult for startups. | |
| Strategic Opportunities for GPU-Poor Developers and Startups | 5 | 6 | 2 | 1 | Alessio asks how under-resourced startups can compete in a compute-heavy ecosystem. Patel highlights edge inference, speculative decoding like Medusa, and asynchronous training architectures as practical avenues for the GPU-poor. | |
| Model Depreciation, Latency Profiling, and Noam Shazeer's Vision | 6 | 6 | 4 | 4 | Alessio and Swyx examine rapid model depreciation cycles. Patel argues fine-tuning smaller models is largely a waste of time and explains how to infer proprietary model architectures using Azure inference latency bounds. | |
| The AI Silicon Landscape and Nvidia's Formidable Moat | 5 | 8 | 5 | 2 | Alessio asks about alternative AI silicon startups and who is viable. Patel explains why first-wave startups bet wrongly on on-chip SRAM over off-chip DRAM bandwidth and breaks down why Nvidia's gross margin and release cadence create an insurmountable moat. | |
| Frontier Lab Partnerships and Apple's GenAI Conundrum | 5 | 5 | 3 | 2 | Patel predicts frontier labs will eventually drift apart from cloud partners due to scale ambitions. Swyx asks about Apple's AI positioning, and Patel explains how Apple's brand risk tolerance conflicts with shipping imperfect LLMs. | |
| AI Safety Debates and Semiconductor Supply Chain Fragility | 4 | 7 | 4 | 1 | Patel challenges security-by-obscurity safety stances before Swyx asks about rebuilding the semiconductor supply chain in the US. Patel delivers an unequivocal rejection, citing obscure geographic monopolies across Austrian tools and Japanese specialty chemicals. | |
| Lightning Round, Research Workflow, and the Genie Question | 5 | 4 | 1 | 1 | In the wrap-up lightning round, Patel recommends foundational reading and describes SemiAnalysis's rigorous on-the-ground supply chain research. He answers the genie question by highlighting multi-datacenter distributed training as AI's ultimate unsolved bottleneck. |