FlashInfer
product on 1 show · 1 statements across 1 episodes · said 2 times in 2 episodes since 2025
Mentions by year, every show
tap a year for its mentions
Latent Space 2
2025 2 mentions in 2 episodes 1 per episode
every mention on every show, scene by scene, with the transcript →
1 statements about FlashInfer, every show
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”