People, every show

Sarah Chieng

Head of Developer Experience, Cerebras Systems. On 1 show, 1 appearance. The Shows tab opens the full record on each.

engineeroperatorfounderinvestor@sarahchieng ↗LinkedIn ↗cerebras.ai ↗

She leads developer experience at Cerebras Systems, focusing on developer workflows, tooling, and technical content surrounding ultra-fast AI inference. An MIT graduate, she previously co-founded the marketplace startup Thrifthouse, was the first hire at Exa AI, and founded the AI builder community Cafe Compute.

1shows
1appearances
10statements
9resolved
8supported
1contradicted
89%fully supported

Everything Sarah Chieng said on any show that made the record, most notable first. Each card names its show and opens the statement there.

LATENT SPACE Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Disclosure
Cerebras avoids model parallelism in production due to communication overhead
“There's a lot of communication overhead with model parallelism. You have to share activation tensors, and that is why in this paper and, you know, in production, Cerebra's focus on data parallelism. So all of this is mentioned in the paper as well, but model p…”
Sarah Chieng Dec 7, 2024 ▶ 21:28 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
GPUs cannot handle unstructured sparsity as efficiently as Cerebras hardware
“So both cerebris and GPUs can handle structured sparsity But GPUs are not designed to handle unstructured sparsity, whereas what I've just mentioned before is able to handle this unstructured sparsity.”
Sarah Chieng Dec 7, 2024 ▶ 38:26 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Contradicted
No competing AI framework disaggregates model storage from compute like Cerebras
“Like basically not, no one is doing anything close to where you're disaggregating. Model storage from compute. And none of these examples above do that either.”
Sarah Chieng Dec 7, 2024 ▶ 42:57 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has it…”
Sarah Chieng Dec 7, 2024 ▶ 9:51 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras WSE-3 features 900,000 cores, 44GB SRAM, and 4 trillion transistors
“And so the wafer scale engine three, as I mentioned, 900,000 cores, 44 gigabytes of SRAM, four trillion transistors, and I do add a note here that the paper focuses on wafer scale engine two, and so the wafer scale engine three is, you know, just an upgraded v…”
Sarah Chieng Dec 7, 2024 ▶ 26:13 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras MemoryX scales to 2.4 petabytes to support 120-trillion-parameter AI models
“And you know, it scales from four terabytes to 2.4 petabytes, Supports models with up to 120 trillion parameters and then it utilizes DRAM and flash storage.”
Sarah Chieng Dec 7, 2024 ▶ 26:22 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras weight streaming prunes up to 90% of data without accuracy loss
“And so memory, and so as MemoryX streams weights through SwarmX, it eliminates zero and near zero values, and so this reduces bandwidth requirements significantly, pruning up to 90% of the data while maintaining accuracy.”
Sarah Chieng Dec 7, 2024 ▶ 36:28 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras streams weights from MemoryX and computes updates externally
“And instead of storing all the weights that the compute units need on [737] Sarah Chieng: On the compute unit, it's storing it externally in an external memory service. In this case, it's called memory X. And during training, these weights are streamed from me…”
Sarah Chieng Dec 7, 2024 ▶ 12:11 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
LATENT SPACE Assertion Supported
Cerebras weight streaming is exclusively for training, while inference runs on SRAM
“The memory X and swarm X, this whole waste streaming system is just used for is just used for training. So for inference, you just using the SRAM, you know, at 44 gigabytes on the chip, and then you can network multiple chips together to support larger models.”
Sarah Chieng Dec 7, 2024 ▶ 31:38 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)

One line per show, most statements first. The link opens Sarah's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Head of Developer Experience, Cerebras Systems 1 10 89% 8/9 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.