Nvidia Blackwell Has 288GB of HBM but Only 500MB of SRAM
“Blackwell has 288 gigabytes of HBM capacity around the logic die, and the logic die itself maybe only has like 500 megabytes of SRAM. So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus SRAM.”
Movva: Sail's Custom Chip Strategy Sidesteps HBM by Offloading to Flash
“When I talk about building custom chips, and they ask me, oh, so what's different? Basically, it's about sidestepping the HBM shortage and focusing on More extreme offload to other forms of memory, such as flash.”
Movva: Compute startups should primarily attack Nvidia's reliance on HBM
“You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead, which, you know, as an aside, I do think is probably the thing to attack most.”
Biderman: Personalized AI adapters require hot-swapping millions of endpoints at inference
“And if you truly believe that we can get to the level where we have those kinds of parameter efficient adapters for every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places that need to…”
Sheldon-Coulson predicts AI hardware will shift away from failure-prone high-bandwidth memory
“There's lots of new accelerators that don't use as much high bandwidth memory and other components that are particularly failure prone, and that's where a lot of the industry will be going.”
Feldman claims HBM memory shortages limit traditional GPUs but not Cerebras
“That is a limitation for all GPUs, but not us. We don't use it.”
Pope: MatX combines SRAM and HBM for cheap, low-latency AI chips
“It is actually possible to do both in the same chip. You, it's kind of an obvious thing. You say you take the HPM, you take the SRAM, put them together in the same chip, you put the weights in SRAM, and you put the all of the inference data in HPM. That is wha…”
Pope: SRAM is an order of magnitude faster per token than HBM
“There's, yeah, there's just some simple math of, like, how long does it take you to read through all of HBM? It takes about 20 milliseconds, and so that's the amount of time per token it runs. Yes. Whereas the amount of time to read through all of SRAM is much…”
O'Laughlin: Manufacturing one bit of HBM takes 4x the DRAM wafer capacity
“Each bit of HBM is essentially a four X multiplier onto DRAM.”
Feldman: Memory market will take 18 months to digest surge, keeping prices high
“And it will take us about 18 months to digest. The prices will stay high.”
Ross: NVIDIA effectively holds a monopsony on High Bandwidth Memory
“The thing is, Nvidia effectively has a monopsony on HBM.”
Ross: Securing AI Memory Allocations Requires Paying Two Years in Advance
“The problem is you have to write that check more than two years in advance.”
Ross: Memory Suppliers Restrict HBM Supply to Protect High Profit Margins
“There's also this situation where the margin on HBM is so high, That no one wants to actually increase the supply, because then the margin goes down.”
Patel: HBM production capacity remains a bottleneck for Huawei
“I think production capacity wise, it is still absolutely a bottleneck. They certain types of equipment required for making HBM need to be imported. They're working on domestic solutions, but as far as we know, they have not imported enough equipment for this.”
Patel: HBM makes up over half of GPU cost
“HBM is more than half the cost of the GPU.”
Krishnan: David Sacks Regularly Explains Deep AI Technical Concepts in White House
“Like you'd be shocked at how often I've seen David Sachs in a meeting explain how inferencing works, you know, what high bandwidth memory is, you know, how the world has shifted from a, you know, pre-training context to a post-training context.”
Feldman: GPU High Bandwidth Memory Offers High Capacity but Is Slow
“They use memory, a memory called HBM, the type of DRAM, and it is phenomenal memory. But it is slow. It is slow and high capacity.”
Ross: NVIDIA holds a cornered resource as HBM monopsony buyer
“You don't normally think of tech companies as having a cornered resource, but Nvidia has a cornered resource. They're a monopsony, the opposite of a monopoly, a single buyer for HBM, and the Interposer, the COOS.”
Haas: TSMC leading-edge capacity and HBM will constrain AI buildout
“Fab capacity does become an issue. No, no doubt. Because when you start thinking about three nanometer and two nanometer you know, TSMC is, is the leader, is the, Only game on town on some level. So I think that is a potential limitation or at least constraint…”
Gerstner: High-bandwidth memory will face shortages for foreseeable future
“And we think that we have a memory shortage, which is a key part of the NVIDIA supercomputer ecosystem for as far as we can see.”
Baker notes high bandwidth memory costs Nvidia more than TSMC chips
“High bandwidth memory is a bigger part of Nvidia's cogs on GPUs than Taiwan Semi is.”
Patel: SK Hynix memory is NVIDIA's largest COGS item, surpassing TSMC
“When you look at the cost of goods sold of NVIDIA their highest cost of goods sold is not TSMC, which is a Thing that people don't realize. It's actually HBM memory primarily SK Hynix.”
Patel: Samsung holds almost zero share in HBM memory, especially at NVIDIA
“In HBM, Samsung has almost no share, right? Especially at NVIDIA”
Patel: Standard high-end server memory yields higher gross margins than HBM
“The gross margins on HBM have not been fantastic. They've been good, but they haven't been fantastic. Actually, regular memory, high-end, like, server memory that is not HBM is actually higher gross margin than HBM.”
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”