Assertion certainty 4/5 debate potential 2/5

DeepSeek runs each single model replica across more than 300 GPUs

Lin Qiao · Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI · Mar 27, 2025 · at 36:10

Lin Qiao, CEO of Fireworks AI and former PyTorch engineering leader at Meta, discusses the extreme compute scale and deployment complexity required for DeepSeek's model architecture.

0:00 / 0:16exact quote · 16.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“DeepSeq actually that company itself was running and still running this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Lin Qiao

Prediction Not checkable as stated
Lin Qiao predicts a 10x AI cost reduction yields 100x more applications
“If this bar can be lowered by 10 times, you can imagine there's so many more, it will be hundred times more applications enter the, this arena to create a brand new experience to end consumers and prosumers. And by that, we'll see a much bigger consumption acr…”
Lin Qiao Mar 27, 2025 ▶ 51:56 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Insight
For AI applications, moats lie in curated data rather than user experience
“Their mode is probably not the user experience, but because it's very easy to copy. Anyone can study the product and copy. Their mode is data.”
Lin Qiao Mar 27, 2025 ▶ 54:20 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Prediction Not checkable as stated
Lin Qiao: The future of AI modeling belongs to open-source models
“The future of the future of modeling sits on open model side. And I believe that side is gonna be much more active in creating those hundreds or maybe thousands of expert models that is specialized delivering much better quality in certain domain.”
Lin Qiao Mar 27, 2025 ▶ 57:29 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Insight
Lin Qiao: AI frameworks must reconcile researcher flexibility with strict production cost and latency constraints
“For researchers, you want the flexibility. You want ease of use. You want them to just think about what's possible, right? And for production, it's a constraint problem solving. As in, you have latency budget, you have cost budget you want to scale, you want t…”
Lin Qiao Mar 27, 2025 ▶ 4:44 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Insight
Lin Qiao: PyTorch's primary success lesson is that simplicity scales
“I think one of the biggest success we saw from the PyTorch experience is simplicity scales.”
Lin Qiao Mar 27, 2025 ▶ 16:26 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Assertion Not checkable as stated
Fireworks AI improved speculative execution hit rates from 30% to 90%
“We have seen cases improving the prediction hit from 30% to 90%, and that's huge speed.”
Lin Qiao Mar 27, 2025 ▶ 29:25 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.