“And so the wafer scale engine three, as I mentioned, 900,000 cores, 44 gigabytes of SRAM, four trillion transistors, and I do add a note here that the paper focuses on wafer scale engine two, and so the wafer scale engine three is, you know, just an upgraded version of that”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Sarah Chieng
AssertionSupported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah ChiengDec 7, 2024▶ 3:06[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Disclosure
Cerebras avoids model parallelism in production due to communication overhead
“There's a lot of communication overhead with model parallelism. You have to share activation tensors, and that is why in this paper and, you know, in production, Cerebra's focus on data parallelism. So all of this is mentioned in the paper as well, but model p…”
Sarah ChiengDec 7, 2024▶ 21:28[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
AssertionSupported
GPUs cannot handle unstructured sparsity as efficiently as Cerebras hardware
“So both cerebris and GPUs can handle structured sparsity But GPUs are not designed to handle unstructured sparsity, whereas what I've just mentioned before is able to handle this unstructured sparsity.”
Sarah ChiengDec 7, 2024▶ 38:26[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
AssertionContradicted
No competing AI framework disaggregates model storage from compute like Cerebras
“Like basically not, no one is doing anything close to where you're disaggregating. Model storage from compute. And none of these examples above do that either.”
Sarah ChiengDec 7, 2024▶ 42:57[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
AssertionSupported
Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has it…”
Sarah ChiengDec 7, 2024▶ 9:51[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
AssertionSupported
Cerebras MemoryX scales to 2.4 petabytes to support 120-trillion-parameter AI models
“And you know, it scales from four terabytes to 2.4 petabytes, Supports models with up to 120 trillion parameters and then it utilizes DRAM and flash storage.”
Sarah ChiengDec 7, 2024▶ 26:22[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.