GPU Kernels

topic on 2 shows · 3 statements across 3 episodes

Latent Space No Priors

3 statements about GPU Kernels, every show

Bubna: Speculative decoding accept length delivers multiplicative speedups over kernel tuning
“People talk a lot about, we made these kernels faster and whatnot, but improving kernel only give you like a few percentage points of improvement and increasing except length literally is a multiplicative decrease.”
Akshat Bubna Jul 8, 2026 ▶ 18:48 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Writing custom GPU kernels is a last resort in model optimization
“Well, kernel is a last resort. So first I go, oh, another kernel. And I try and find some way to go around it.”
Quentin Anthony Nov 3, 2025 ▶ 15:50 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
NO PRIORS Insight
Srivastava: Inference lacks abstractions, forcing engineers to rewrite low-level GPU kernels
“A lot of the optimization you're doing is pretty low level and there's no real abstraction. So you either have to learn how to use open source very well or rewrite some of these kernels by yourself.”
Tuhin Srivastava Mar 21, 2024 ▶ 10:38 No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.