Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
Sarah Chieng · [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) · Dec 7, 2024 · at 9:51
Sarah Chieng of Cerebras explains why the Wafer-Scale Engine architecture achieves higher inference speed compared to GPUs like the NVIDIA H100.
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has its own SRAM. So each core has direct access To each of the values that it needs to do its computations. And this design eliminates the need to repeatedly fetch weights from a central memory, which then significantly reduces the memory bandwidth requirements and power consumption of the wafer scale engine.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →