SRAM
topic on 6 shows · 13 statements across 10 episodes
Cheeky Pint
Latent Space
No Priors
Invest Like the Best
All-In
20VC
13 statements about SRAM, every show
Movva: Cerebras and Groq Bet on Maximizing On-Chip SRAM Over Traditional GPUs
“Cerebris, Grok and a couple others that are coming out of stealth now, I think have made a very interesting bet on not just building another GPU, but actually building a different kind of accelerator that focuses on a different memory hierarchy. They want to m…”
Movva: An 800mm² Die Made Entirely of SRAM Would Hold Only Single-Digit Gigabytes
“The problem with SRAM is it takes a lot of area on the silicon die so if you want to build a large die, like let's say the NVIDIA Blackwell at 800 millimeters square, if you made that whole die SRAM, It would be in the maybe like single digit gigabytes. It's n…”
Nvidia Blackwell Has 288GB of HBM but Only 500MB of SRAM
“Blackwell has 288 gigabytes of HBM capacity around the logic die, and the logic die itself maybe only has like 500 megabytes of SRAM. So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus SRAM.”
Gavin Baker: SRAM Is Unbeatable for AI Feed-Forward Networks
“You just can't beat SRAM in particular for that feed forward network.”
Feldman notes SRAM costs remain stable with no supply shortages
“We use SRAM, and there's no shortage of SRAM. The cost of SRAM hasn't changed.”
Pope: MatX combines SRAM and HBM for cheap, low-latency AI chips
“It is actually possible to do both in the same chip. You, it's kind of an obvious thing. You say you take the HPM, you take the SRAM, put them together in the same chip, you put the weights in SRAM, and you put the all of the inference data in HPM. That is wha…”
Pope: SRAM is an order of magnitude faster per token than HBM
“There's, yeah, there's just some simple math of, like, how long does it take you to read through all of HBM? It takes about 20 milliseconds, and so that's the amount of time per token it runs. Yes. Whereas the amount of time to read through all of SRAM is much…”
Dean: Accelerator batching is driven by 1000x SRAM data movement energy costs
“And so, all of a sudden, this is why your accelerators require batching, because if you move, like, say, the parameter of a model from SRAM on the chip into the multiplier unit, that's gonna cost you a thousand PicoTools, so you better make use of that, that t…”
Chamath: Next-gen AI silicon will shift to memory-centric SRAM architectures
“The next generation of silicon will not take The same compute-heavy approach, and will probably rely on a much more memory-centric architecture that uses a lot of SRAM.”
Morin: Scaling SRAM is a dead end for AI hardware
“No, SRAM, this will not deliver. It's a dead end in terms of scaling SRAM means scaling the surface mean you get, you know, depreciating problems.”
Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has it…”
Feldman: Cerebras eliminates memory bandwidth bottlenecks using on-wafer SRAM
“We keep a huge amount of SRAM on the wafer. All right, and so there are no memory bandwidth problems ever. That also allows us to harvest sparsity, which is something that others really struggle with.”
SRAM capacity will stagnate, making memory-aware algorithms vital
“And so, yeah, I think in the future SRAM probably won't get that much larger because you don't have that much area. HRAM will get larger and faster, and so I think it becomes more important to design algorithms that take advantage of this memory asymmetry.”