FlashInfer, every mention
2 scenes · ← back to FlashInfer
tap a year for its mentions
every year anyone Yining Zhang 1Diego Bachman 1
Verbatim, from the transcripts: the passages where FlashInfer comes up
⚡️ Beyond Transformers with Power Retention
- ▶ 7:47 Diego Bachman We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power…
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 44:54 Yining Zhang And we also co-host some meetups, something like the first meetup we co-host with the MLCLM and FlashInfer.