KV Cache Transfers
topic on 1 show · 1 statements across 1 episodes
1 statements about KV Cache Transfers, every show
Network interface card speed is the primary bottleneck in large-scale AI serving
“I think the answer is just faster next, like faster network chip communications. It seems to me that like more and more memory is the bottom, like you want to have larger models. Right now, when you're doing serving at large, you have to transfer KV cache from…”