Decode

topic on 3 shows · 5 statements across 3 episodes

the Y Combinator Startup Podcast Latent Space Invest Like the Best

5 statements about Decode, every show

Y COMBINATOR Assertion Supported
Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once an…”
Misha Smulyanski Jul 29, 2026 ▶ 51:45 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Baker: LLM pre-fill is capacity-bound while decode is bandwidth-constrained
“And that is fundamentally a memory capacity bound problem. Decode is the process of generating new tokens, and that is memory bandwidth constraint.”
Gavin Baker May 20, 2026 ▶ 45:23 Watts, Wafers, and the Future of AI Infra | Gavin Baker · Invest Like The Best
LLM prefill remains compute-bound while decoding phases are strictly memory-bound
“So prefill typically, and this changes as model architecture changes, prefill is right now compute bound. Most of the time. If the sequence is sufficiently long, it's compute bound on the decode side because you're doing a full pass over all the weights and th…”
Kyle Kranen Mar 8, 2026 ▶ 41:14 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Local document pre-fill with global sequence decode solves transformer quadratic scaling
“If pre-fill becomes local and decode is, is still global, you solve that pre-fill quadratic scaling problem because you have a bunch of like small chunks that you pre-fill independently.”
Kyle Kranen Mar 8, 2026 ▶ 53:49 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.