Insight certainty 4/5 debate potential 2/5

Pipeline parallelism is only necessary for multi-node AI model inference

Philip Kiely · Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten · Aug 3, 2026 · at 54:23

Baseten engineer Philip Kiely explains distributed parallelism strategies across GPU clusters for large language models.

0:00 / 0:15exact quote · 15.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The only reason you would have to do pipeline parallelism which is where you separate, like, different layers, and you put, like, half the layers on one hardware and half on another, is if you are forced to do multi-node inference because a model is bigger than you have.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Philip Kiely

Insight
Quantization, speculative decoding, and prefill-decode disaggregation each yield 2x speedups
“Going from BF-sixteen to NVFP four is, it's not quite a two X, right? It's like, I think it's about like 30, 30 to 40% from 16 to eight, and then another 30 to 40% multiplied from eight to four. So that doesn't quite get you a two X, but like roughly a two X. …”
Philip Kiely Aug 3, 2026 ▶ 40:03 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
Leading AI agents will utilize continuous inference-to-training learning loops within two years
“And I think within a few months to a couple of years, like a lot of, Leading agent builders are going to have these loops, like, really up and running in production, where you are doing inference, learning from the inference.”
Philip Kiely Aug 3, 2026 ▶ 1:31:15 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Insight
AI customers migrate from pay-per-token APIs to dedicated GPUs upon finding fit
“I think that we've increasingly seen a lot of demand for the sort of pay per token APIs just because everyone wants to try open models. And then once they find a use case that's really sticky then they move over to Dedicated.”
Philip Kiely Aug 3, 2026 ▶ 5:03 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Supported
Open-source inference engines often receive model weights before official launches
“Getting to the point of I can make a token out of this model is not that hard because generally the open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the people making the model merg…”
Philip Kiely Aug 3, 2026 ▶ 13:51 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Opinion
DeepSeek models are the hardest to support due to architectural novelties
“I think that, like, obviously the DeepSeq models tend to be the most challenging as they have, like, the most novel architectural stuff going on model over model.”
Philip Kiely Aug 3, 2026 ▶ 16:01 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Insight
Modifying base LLM weights for vision degrades original text performance
“You don't want to mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision.”
Philip Kiely Aug 3, 2026 ▶ 17:13 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.