FlashInfer, every mention

2 scenes · ← back to FlashInfer

tap a year for its mentions
0011222025episodesmentions
0122025episodes it came up in
000.51122025episodesmentions per episode

every year anyone Yining Zhang 1Diego Bachman 1

Verbatim, from the transcripts: the passages where FlashInfer comes up

loading…

⚡️ Beyond Transformers with Power Retention Sep 23, 2025 · 1 mention

  • ▶ 7:47 Diego Bachman We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power…

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 1 mention

  • ▶ 44:54 Yining Zhang And we also co-host some meetups, something like the first meetup we co-host with the MLCLM and FlashInfer.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.