Wafer-Scale Engine 3
also referred to as: wafer scale engine 3
part of Wafer Scale Engine
1 statements across 1 episodes · 2 bullish · 0 bearish · 1 people on the record · first statement Dec 7, 2024 by Sarah Chieng · across every show →
Everything said about Wafer-Scale Engine 3, oldest first
Dec 7, 2024 positive
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Dec 7, 2024 bullish
Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has it…”