Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes

Misha Smulyanski · Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator · Jul 29, 2026 · at 51:45

Misha Smulyanski (co-founder of Marlowe.ai) explains the arithmetic intensity differences between prefill and decode stages in LLM inference.

0:00 / 1:04exact quote · 64.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once and you operate on all of them, right? The MLP is even more intensive because you can actually batch things, right? And so it's pretty compute bound. Like, the code is different. Like, the code, you actually work on, you know, one token. At a time, it's out there. Regressive. So attention is really horrible. It's like you're fetching, you know, all the ways for every tokens that you want to process, right? So it's not very, you know, it's really inefficient, right? The MLP is a little bit better because you're actually, you can actually batch it, but in reality batches are not that big. You know, the machines are like modern accelerator have a lot of compute, so you end up being bandwidth bound for a, You know, pretty large batches as well.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Misha Smulyanski

Insight
Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Misha Smulyanski Jul 29, 2026 ▶ 47:35 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Insight
Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Insight
Smulyanski: GPU throughput drops sharply at low concurrency due to kernel overheads
“The moment that you start basically going to lower concurrency because you want better interactivity and better latency, Right? The performance the throughput drops. And it drops very sharply because all of a sudden you have a lot of, like, smaller kernels, yo…”
Misha Smulyanski Jul 29, 2026 ▶ 59:40 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.