inference

7 statements across 6 episodes · 1 bullish · 1 bearish · 7 people on the record · first statement Nov 25, 2024 by Lin Qiao · across every show →

Everything said about inference, oldest first

Nov 25, 2024 bullish
Insight
AI inference matters more than training because it scales with global population
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
Lin Qiao Nov 25, 2024 ▶ 8:59 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Jul 16, 2025 negative
Insight
Rizwan: AI coding tools should not monetize through inference markups
“Our thesis is inference is not the business. We, yeah, we want to give the end user total transparency into price, into, which I think is like incredibly important to, you know, even get comfortable with the idea of spending as much money as you do. I think th…”
Saoud (Saud) Rizwan Jul 16, 2025 ▶ 55:44 Cline: The Collaborative AI Coder
Aug 29, 2025 neutral
Insight
Morcos: AI total cost of ownership is dominated by inference
“When you think about the total cost of ownership of these models, it's gonna be very, very heavily weighted towards inference. It's all inference.”
Ari Morcos Aug 29, 2025 ▶ 58:04 Better Data is All You Need — Ari Morcos, Datology
Dec 30, 2025 neutral
Assertion Not checkable as stated
Catanzaro: Continual Learning Requires Making Stateless Inference Stateful
“If you must update weights, then, like, you know, weights become stateful, and today, like, inference is not stateful.”
Sarah Catanzaro Dec 30, 2025 ▶ 23:00 [State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify
Dec 31, 2025 neutral
Disclosure
Angelopoulos: LMArena's top expenses are free-tier inference, hiring, and SF office
“Primarily inference that funds the free usage of the platform and then also hiring, of course, headcount. We have an office, you know. That's an SF.”
Anastasios Angelopoulos Dec 31, 2025 ▶ 9:57 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Aug 3, 2026
Insight
Optimal inference parallelism cannot be mathematically calculated; it must be auto-tuned
“And with training, it's more of like a math, like you can run the math and see the flops and maximize it. With inference, it's more of like an auto-tuning, like GPU kernel auto-tuning... You shadow the same traffic, like real traffic, and you just see which co…”
Ali Taha Aug 3, 2026 ▶ 56:06 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026
Insight
Inference optimization is only solved once researchers report mere 1% speedups
“Like, you'll, you'll know that influence is pretty much solved when researchers start publishing about how they got one percent faster at something.”
Philip Kiely Aug 3, 2026 ▶ 34:37 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.