KV cache

also referred to as: kvcache

13 statements across 9 episodes · 7 bullish · 1 bearish · 9 people on the record · first statement Jun 25, 2024 by Jonathan Frankle · across every show →

Everything said about KV cache, oldest first

Jun 25, 2024 negative
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan Frankle Jun 25, 2024 ▶ 1:17:50 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Sep 19, 2024
Assertion Supported
Writing in the Margins requires inference engine modifications but zero fine-tuning
“We are not doing any change to the model architecture, so you don't have to fine-tune anything, you don't have to change anything, like the can you use this stuff with, like, a Lank pane? No, because it requires a modification on how the inference engine is us…”
Umar Jamil Sep 19, 2024 ▶ 26:20 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Sep 19, 2024
Insight
Masking out prior KV cache tokens pushes autoregressive transformers out of distribution
“The token number two in the KVCache is a contextualized version of the token zero, one, and two. So if you tell the model to only look at the last tokens you are creating an autoregressive model that is generating the logits of a P of let's say X, but only loo…”
Umar Jamil Sep 19, 2024 ▶ 33:09 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Sep 19, 2024 bullish
Insight
Jamil: Writing in the Margins avoids re-prefilling tokens, halving compute cost
“Again, to the language model to generate the answer, and it would cost you another million, because the model has to reprocess this prefilling again of one million tokens, so it would cost you two million tokens, but with writing in the margins, it would cost …”
Umar Jamil Sep 19, 2024 ▶ 16:43 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Jan 26, 2025 positive
Assertion Supported
Beauchamp: DeepSeek slashes inference costs by shrinking KV cache
“What DeepSeek have achieved that's quite special is they've got this amazing inference engine. They've been able to reduce the size of the KV cache significantly. And then by being able to do that, they're able to significantly reduce their inference costs.”
William Beauchamp Jan 26, 2025 ▶ 23:30 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Jun 13, 2025
Assertion Supported
CPU latency in KV cache management bottlenecks GPU utilization
“Like your eviction policy runs on a CPU. Like that radix hashing algorithm and block hashing and all that stuff happens like primarily CPU. That's really important for performance because if you have latency in these steps, like you're not keeping your GPU uti…”
Chris Lattner Jun 13, 2025 ▶ 1:00:09 The Shape of Compute (Chris Lattner of Modular)
Sep 23, 2025 positive
Disclosure
Bachman: Manifest AI is releasing Power Retention architecture with fixed-size memory
“Power retention is the specific variant that we're about to release. And it basically works by instead of the memory constantly growing, this constantly Growing KVCache. You have a memory that is a fixed size and each new token simply gets compressed into this…”
Diego Bachman Sep 23, 2025 ▶ 4:17 ⚡️ Beyond Transformers with Power Retention
Sep 23, 2025 bullish
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46 ⚡️ Beyond Transformers with Power Retention
Oct 11, 2025 bullish
Assertion Not checkable as stated
Lenz: Local smartphone AI requires hybrid models due to KV cache limits
“So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or without doing drastically changes because the model plus KVCache won't fit.”
Barak Lenz Oct 11, 2025 ▶ 13:09 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Mar 8, 2026 neutral
Insight
LLM prefill remains compute-bound while decoding phases are strictly memory-bound
“So prefill typically, and this changes as model architecture changes, prefill is right now compute bound. Most of the time. If the sequence is sufficiently long, it's compute bound on the decode side because you're doing a full pass over all the weights and th…”
Kyle Kranen Mar 8, 2026 ▶ 41:14 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Jul 8, 2026
Insight
Bubna: Transferring RL weights is fundamentally an OS memory problem
“Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is, there's a lot of degrees of freedom, and it is basically a systems problem of Moving me…”
Akshat Bubna Jul 8, 2026 ▶ 31:40 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Aug 3, 2026 bullish
Opinion
KV cache compaction is the viable path to solve LLM continual learning
“If you're able to sort of make your KV almost infinite, and you're able to compact in such a way that you don't lose any of the knowledge. In that case, you can actually do a continual learning as you can actually solve continual learning. And this, as a, it's…”
Ali Taha Aug 3, 2026 ▶ 1:40:41 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 bullish
Insight
Ultra-fast network interface cards could deliver 100x speedups in disaggregated AI inference
“If you were to somehow be able to, in like this theoretical dreamland, have extremely fast NICs, you could, in theory, spare that HBM and you could just transfer KVCache trans, like directly from one node to another. This would give you like almost a hundred X…”
Ali Taha Aug 3, 2026 ▶ 1:38:00 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.