High Bandwidth Memory

topic on 9 shows · 25 statements across 17 episodes · said 1 times in 1 episodes since 2026

All-In 1 BG2 Pod Cheeky Pint Latent Space No Priors Invest Like the Best Catalyst the a16z Podcast 20VC

Mentions by year, every show

tap a year for its mentions
0011112026episodesmentions
0112026episodes it came up in
000.50.5112026episodesmentions per episode

All-In 1

2026 1 mention in 1 episode

every mention on every show, scene by scene, with the transcript →

25 statements about High Bandwidth Memory, every show

INVEST LIKE THE BEST Assertion Partly supported
Nvidia Blackwell Has 288GB of HBM but Only 500MB of SRAM
“Blackwell has 288 gigabytes of HBM capacity around the logic die, and the logic die itself maybe only has like 500 megabytes of SRAM. So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus SRAM.”
Neil Movva Aug 25, 2026 ▶ 26:04 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Movva: Sail's Custom Chip Strategy Sidesteps HBM by Offloading to Flash
“When I talk about building custom chips, and they ask me, oh, so what's different? Basically, it's about sidestepping the HBM shortage and focusing on More extreme offload to other forms of memory, such as flash.”
Neil Movva Aug 25, 2026 ▶ 1:14:23 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Movva: Compute startups should primarily attack Nvidia's reliance on HBM
“You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead, which, you know, as an aside, I do think is probably the thing to attack most.”
Neil Movva Aug 25, 2026 ▶ 1:19:42 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Biderman: Personalized AI adapters require hot-swapping millions of endpoints at inference
“And if you truly believe that we can get to the level where we have those kinds of parameter efficient adapters for every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places that need to…”
Dan Biderman Jul 13, 2026 ▶ 43:43 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
CATALYST Prediction Not checkable as stated
Sheldon-Coulson predicts AI hardware will shift away from failure-prone high-bandwidth memory
“There's lots of new accelerators that don't use as much high bandwidth memory and other components that are particularly failure prone, and that's where a lot of the industry will be going.”
Garth Sheldon-Coulson May 28, 2026 ▶ 37:32 Building inference data centers on the high seas
20VC Disclosure
Feldman claims HBM memory shortages limit traditional GPUs but not Cerebras
“That is a limitation for all GPUs, but not us. We don't use it.”
Andrew Feldman May 26, 2026 ▶ 8:29 Cerebras CEO on the Future of Data Centres, Token Costs & Memory | Should US Companies Sell to China · 20VC with Harry Stebbings
CHEEKY PINT Disclosure
Pope: MatX combines SRAM and HBM for cheap, low-latency AI chips
“It is actually possible to do both in the same chip. You, it's kind of an obvious thing. You say you take the HPM, you take the SRAM, put them together in the same chip, you put the weights in SRAM, and you put the all of the inference data in HPM. That is wha…”
Reiner Pope Feb 26, 2026 ▶ 16:12 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
CHEEKY PINT Assertion Supported
Pope: SRAM is an order of magnitude faster per token than HBM
“There's, yeah, there's just some simple math of, like, how long does it take you to read through all of HBM? It takes about 20 milliseconds, and so that's the amount of time per token it runs. Yes. Whereas the amount of time to read through all of SRAM is much…”
Reiner Pope Feb 26, 2026 ▶ 16:57 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
LATENT SPACE Assertion Supported
O'Laughlin: Manufacturing one bit of HBM takes 4x the DRAM wafer capacity
“Each bit of HBM is essentially a four X multiplier onto DRAM.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:43:34 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
ALL-IN Prediction Open · timeframe Jul 2027
Feldman: Memory market will take 18 months to digest surge, keeping prices high
“And it will take us about 18 months to digest. The prices will stay high.”
Andrew Feldman Jan 23, 2026 ▶ 1:04:41 Coinbase CEO's Top 3 Crypto Trends for 2026 + More from Davos!
20VC Assertion Partly supported
Ross: NVIDIA effectively holds a monopsony on High Bandwidth Memory
“The thing is, Nvidia effectively has a monopsony on HBM.”
Jonathan Ross Sep 29, 2025 ▶ 14:26 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
20VC Assertion Not checkable as stated
Ross: Securing AI Memory Allocations Requires Paying Two Years in Advance
“The problem is you have to write that check more than two years in advance.”
Jonathan Ross Sep 29, 2025 ▶ 17:29 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
20VC Insight
Ross: Memory Suppliers Restrict HBM Supply to Protect High Profit Margins
“There's also this situation where the margin on HBM is so high, That no one wants to actually increase the supply, because then the margin goes down.”
Jonathan Ross Sep 29, 2025 ▶ 18:01 Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
a16z Assertion Not checkable as stated
Patel: HBM production capacity remains a bottleneck for Huawei
“I think production capacity wise, it is still absolutely a bottleneck. They certain types of equipment required for making HBM need to be imported. They're working on domestic solutions, but as far as we know, they have not imported enough equipment for this.”
Dylan Patel Sep 22, 2025 ▶ 15:21 Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China
a16z Assertion Partly supported
Patel: HBM makes up over half of GPU cost
“HBM is more than half the cost of the GPU.”
Dylan Patel Sep 22, 2025 ▶ 1:34:28 Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China
NO PRIORS Assertion Not checkable as stated
Krishnan: David Sacks Regularly Explains Deep AI Technical Concepts in White House
“Like you'd be shocked at how often I've seen David Sachs in a meeting explain how inferencing works, you know, what high bandwidth memory is, you know, how the world has shifted from a, you know, pre-training context to a post-training context.”
Sriram Krishnan Jul 31, 2025 ▶ 28:52 No Priors Ep. 125 | With Senior White House Policy Advisor on AI Sriram Krishnan
20VC Assertion Supported
Feldman: GPU High Bandwidth Memory Offers High Capacity but Is Slow
“They use memory, a memory called HBM, the type of DRAM, and it is phenomenal memory. But it is slow. It is slow and high capacity.”
Andrew Feldman Mar 24, 2025 ▶ 6:25 Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance · 20VC with Harry Stebbings
20VC Assertion Contradicted
Ross: NVIDIA holds a cornered resource as HBM monopsony buyer
“You don't normally think of tech companies as having a cornered resource, but Nvidia has a cornered resource. They're a monopsony, the opposite of a monopoly, a single buyer for HBM, and the Interposer, the COOS.”
Jonathan Ross Feb 17, 2025 ▶ 16:58 Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260 · 20VC with Harry Stebbings
BG2 Prediction Not checkable as stated
Haas: TSMC leading-edge capacity and HBM will constrain AI buildout
“Fab capacity does become an issue. No, no doubt. Because when you start thinking about three nanometer and two nanometer you know, TSMC is, is the leader, is the, Only game on town on some level. So I think that is a potential limitation or at least constraint…”
Rene Haas Jan 23, 2025 ▶ 42:57 Stargate, Executive Orders, TikTok, DOGE, Public Valuations | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
BG2 Prediction Not checkable as stated
Gerstner: High-bandwidth memory will face shortages for foreseeable future
“And we think that we have a memory shortage, which is a key part of the NVIDIA supercomputer ecosystem for as far as we can see.”
Brad Gerstner Jan 11, 2025 ▶ 38:06 Market Predictions, Rates & Inflation, DOGE, CES, AI Compute | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
ALL-IN Assertion Not checkable as stated
Baker notes high bandwidth memory costs Nvidia more than TSMC chips
“High bandwidth memory is a bigger part of Nvidia's cogs on GPUs than Taiwan Semi is.”
Gavin Baker Jan 4, 2025 ▶ 1:14:23 2025 Predictions: Tech, Business, Media, Politics!
BG2 Assertion Not checkable as stated
Patel: SK Hynix memory is NVIDIA's largest COGS item, surpassing TSMC
“When you look at the cost of goods sold of NVIDIA their highest cost of goods sold is not TSMC, which is a Thing that people don't realize. It's actually HBM memory primarily SK Hynix.”
Dylan Patel Dec 23, 2024 ▶ 1:02:27 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
BG2 Assertion Partly supported
Patel: Samsung holds almost zero share in HBM memory, especially at NVIDIA
“In HBM, Samsung has almost no share, right? Especially at NVIDIA”
Dylan Patel Dec 23, 2024 ▶ 1:03:16 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
BG2 Assertion Not checkable as stated
Patel: Standard high-end server memory yields higher gross margins than HBM
“The gross margins on HBM have not been fantastic. They've been good, but they haven't been fantastic. Actually, regular memory, high-end, like, server memory that is not HBM is actually higher gross margin than HBM.”
Dylan Patel Dec 23, 2024 ▶ 1:05:08 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
LATENT SPACE Prediction Not checkable as stated
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”
Tri Dao Aug 3, 2023 ▶ 16:07 FlashAttention-2: Making Transformers 800% faster AND exact

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.