Memory Bandwidth

topic on 4 shows · 9 statements across 7 episodes

the Y Combinator Startup Podcast Latent Space Invest Like the Best 20VC

9 statements about Memory Bandwidth, every show

Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Smulyanski: GPU throughput drops sharply at low concurrency due to kernel overheads
“The moment that you start basically going to lower concurrency because you want better interactivity and better latency, Right? The performance the throughput drops. And it drops very sharply because all of a sudden you have a lot of, like, smaller kernels, yo…”
Misha Smulyanski Jul 29, 2026 ▶ 59:40 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Baker: LLM pre-fill is capacity-bound while decode is bandwidth-constrained
“And that is fundamentally a memory capacity bound problem. Decode is the process of generating new tokens, and that is memory bandwidth constraint.”
Gavin Baker May 20, 2026 ▶ 45:23 Watts, Wafers, and the Future of AI Infra | Gavin Baker · Invest Like The Best
20VC Insight
Feldman: Memory bandwidth, not compute speed, limits GPU AI inference
“And that includes memory, which has, for inference, is the fundamental limiter for the GPU architecture. And so it doesn't matter how much faster the chip goes. It matters how much faster the memory bandwidth is.”
Andrew Feldman Oct 6, 2025 ▶ 19:30 Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
Feldman: Memory bandwidth is the primary bottleneck in AI inference performance
“Inference performance comes from memory bandwidth and the memory bandwidth is the limiting factor. In inference performance. Remember, in order to generate a token, to generate a word, all the weights have to move from memory to compute. If you're constrained …”
Andrew Feldman Oct 1, 2025 ▶ 5:56 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Sohmers: AI hardware over-indexes on raw FLOPS instead of memory bandwidth
“Everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, real bottlenecks.”
Thomas Sohmers Aug 18, 2025 ▶ 6:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
LATENT SPACE Assertion Open · timeframe Aug 2026
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Thomas Sohmers Aug 18, 2025 ▶ 15:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
LATENT SPACE Assertion Supported
Patel: Running LLaMA-70B at reading speed requires 2.1 TB/s memory bandwidth
“Hey, to run Llama's seventy billion requires two terabytes a second of memory bandwidth, 2.1, at reading, human reading speed.”
Dylan Patel Dec 5, 2023 ▶ 26:06 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
20VC Assertion Supported
Seibert: Top AI models are bottlenecked by memory bandwidth, not just compute
“It's very clear the top models are memory bound as well as CPU bound. So you can't just make the CPUs, the GPUs faster. You need to increase memory bandwidth in sort of on par with that.”
Jeff Seibert Nov 22, 2023 ▶ 35:36 Jeff Seibert: Why OpenAI Will Become an Infrastructure Play | E1085 · 20VC with Harry Stebbings

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.