LLM Inference

topic on 5 shows · 7 statements across 7 episodes

In Depth Latent Space the MAD Podcast TBPN 20VC

7 statements about LLM Inference, every show

MAD Disclosure
Pedregal: Audio transcription is Granola's largest cost, exceeding LLM inference
“So the most expensive thing about our business is actually transcription. And historically it's been actually transcription and high quality transcription versus LLM inference.”
Chris Pedregal Aug 21, 2025 ▶ 45:42 How to Build a Beloved AI Product - Granola CEO Chris Pedregal
TBPN Assertion Supported
Russ d'Sa: LLM inference is now faster than text-to-speech generation
“Now like LLM inference can actually be done in less time than generating speech with TTS.”
Russell D'Sa Apr 26, 2025 ▶ 25:54 Building the World’s Smallest Fusion Reactor | Robin Langtry on TBPN
TBPN Assertion Supported
D'Sa: LLM inference is now faster than TTS speech generation
“And now like LLM inference can actually be done in less time than generating speech with TTS.”
Russell D'Sa Apr 26, 2025 ▶ 5:38 Why Hallucination Is Good For AI Voice Agents | Russell D'Sa on TBPN
20VC Opinion
Lazarte: Now is the best time in over a decade to start a startup
“It's quite clear. It's the best time in over a decade to start a company. Because we have this insane phenomenon of adoption, like adoption of LLM inference.”
Victor Lazarte Apr 14, 2025 ▶ 5:36 Benchmark GP, Victor Lazarte: The 3 Traits All the Best Founders Have · 20VC with Harry Stebbings
MAD Insight
LLM prompt processing is compute-bound; next-token generation is memory-bound
“Prompt processing is bottlenecked by computation, and generating next, predicting next token is bottlenecked by memory bandwidth.”
Lin Qiao Mar 27, 2025 ▶ 30:42 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
LATENT SPACE Prediction Didn’t hold up
Patel: AI inference will deploy more GPUs than training by 2024
“LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact at least in terms of GPUs deployed,”
Dylan Patel Dec 5, 2023 ▶ 14:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
IN DEPTH Prediction Held up
LLM inference costs will decrease significantly in relatively short order
“I think you're gonna see the cost of inference tend to go down. Inference is the cost to actually serve and run the model. I think you're gonna see those things go down in relatively short order.”
Jack Krawczyk Nov 30, 2023 ▶ 1:11:37 The Bard blueprint | Creating value, shipping fast, and advancing AI | Jack Krawczyk (Google)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.