Quantization

topic on 2 shows · 6 statements across 5 episodes

Latent Space the MAD Podcast

6 statements about Quantization, every show

Most inference optimizations like KV caching are lossless, unlike quantization
“Most inference optimizations are lossless. KV caching, for example, you are just recomputing or preventing recomputing the same values. Speculation, of course, If a draft token is wrong, it gets rejected. The main lossy optimization is quantization”
Philip Kiely Aug 3, 2026 ▶ 26:55 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Quantizing more layers can actually improve model fidelity via error cancellation
“It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in is that you can predict which …”
Ali Taha Aug 3, 2026 ▶ 31:54 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
MAD Assertion Not checkable as stated
Dettmers: 4-bit precision is the end of quantization
“Four bit precision is the end of quantization.”
Tim Dettmers Jan 22, 2026 ▶ 16:03 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
MAD Assertion Not checkable as stated
Chip Huyen: Quantization works universally well across tasks and models
“Quantizations, which is like very universally very working really well. For a lot of tasks across model.”
Chip Huyen Jan 16, 2025 ▶ 23:41 What You MUST Know About AI Engineering | Chip Huyen, Author of “AI Engineering”
Cheah: Over-quantized LLMs rapidly enter repetition loops at long contexts
“When you over-quantize, right, at longer context length, right, it starts going into repetition rapidly.”
Eugene Cheah Jul 29, 2024 ▶ 1:03:20 [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
LATENT SPACE Assertion Supported
Royzen: INT8 quantization offers storage optimization without guaranteed inference speedups
“But with int eight, there's not necessarily a Speed increase. It's just the storage optimization.”
Michael Royzen Nov 3, 2023 ▶ 1:05:46 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.