Inference Optimization

topic on 3 shows · 3 statements across 3 episodes

Latent Space No Priors TBPN

3 statements about Inference Optimization, every show

LATENT SPACE Assertion Not checkable as stated
Software optimizations yield 2x to 4x inference speedups on identical hardware
“Yeah, then you're looking at, like, a two to four X improvement, depending on the inference optimizations.”
Philip Kiely Aug 3, 2026 ▶ 39:00 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
TBPN Insight
Kilpatrick: Large High-Demand AI Models Require Massive Inference Investments
“Especially with larger models and especially with models that have lots of demand, like there's no world where you can get away with not putting a large order of magnitude of investment into inference.”
Logan Kilpatrick Apr 25, 2025 ▶ 14:44 Google's AI Comeback in Their Own Words - Logan Kilpatrick
NO PRIORS Insight
Srivastava: Inference lacks abstractions, forcing engineers to rewrite low-level GPU kernels
“A lot of the optimization you're doing is pretty low level and there's no real abstraction. So you either have to learn how to use open source very well or rewrite some of these kernels by yourself.”
Tuhin Srivastava Mar 21, 2024 ▶ 10:38 No Priors Ep 56 | With Baseten CEO and Co-Founder Tuhin Srivastava

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.