decode

2 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Mar 8, 2026 by Kyle Kranen · across every show →

Everything said about decode, oldest first

Mar 8, 2026 neutral
Insight
LLM prefill remains compute-bound while decoding phases are strictly memory-bound
“So prefill typically, and this changes as model architecture changes, prefill is right now compute bound. Most of the time. If the sequence is sufficiently long, it's compute bound on the decode side because you're doing a full pass over all the weights and th…”
Kyle Kranen Mar 8, 2026 ▶ 41:14 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 positive
Insight
Local document pre-fill with global sequence decode solves transformer quadratic scaling
“If pre-fill becomes local and decode is, is still global, you solve that pre-fill quadratic scaling problem because you have a bunch of like small chunks that you pre-fill independently.”
Kyle Kranen Mar 8, 2026 ▶ 53:49 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.