Nov 25, 2024 · 55m · latent-space

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI

Lin Qiao · 38m spoken Shawn Wang · 9m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode, Fireworks AI CEO Lin Qiao sits down with Alessio Fanelli and Swyx to discuss the technical architecture of distributed inference, the shift toward Compound AI and declarative reasoning, and why open-source AI infrastructure is poised to outperform closed-source alternatives.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 22.7% of the talking time here. How this is scored →

The hosts as informed peer 4.9 Guest teaching 5.7 Guest disagreement 1.5 The hosts pushing back 2.2
05100:0015:0030:0045:002:08–6:58 · The hosts as informed peer 4/10 Lin Qiao's Background at Meta and PyTorch Evolution Swyx sets up the prehistory of PyTorch and Meta's early GenAI efforts, referencing past guest Soumith Chintala. Lin delivers an extensive masterclass on scaling PyTorch from research to production across Meta's ubiquitous systems.6:58–9:49 · The hosts as informed peer 4/10 Pivoting from Horizontal PyTorch Cloud to Generative AI Inference Swyx probes Fireworks' initial pivot from a horizontal PyTorch cloud to inference hosting. Lin explains the market dynamics of 2022, detailing why inference scales with the global population whereas training only scales with researchers.9:49–13:52 · The hosts as informed peer 5/10 Focusing on App Developers and the Emergence of Llama Stack Swyx pushes back on Meta's Llama Stack, expressing skepticism that developer adoption will follow simply because Llama is open source. Lin acknowledges it is very early and emphasizes Fireworks' role in delivering direct community feedback.13:52–20:26 · The hosts as informed peer 4/10 The Concept of Compound AI and Multi-Modal Optimization Swyx asks why Fireworks leaned heavily into Compound AI post-Series B. Lin articulates the architectural shift from one-size-fits-all distributed inference to multi-modal compound systems and three-dimensional optimization across quality, latency, and cost.20:26–24:13 · The hosts as informed peer 5/10 Distributed Inference Architecture and Custom Kernel Acceleration Alessio and Swyx press Lin on what distributed inference actually entails versus raw GPU rental moats. Lin details custom CUDA kernels, disaggregated execution, regional routing, and hardware specialization across heterogeneous architectures.24:14–30:40 · The hosts as informed peer 5/10 Declarative AI Architecture and Fireworks' New Reasoning Model Lin uses database SQL analogies to explain declarative AI system architectures over complex DAG pipelines. She reveals Fireworks' upcoming reasoning model trained to approach o1 reasoning quality.30:41–38:38 · The hosts as informed peer 7/10 Inference Scaling Laws, Model Specialization, and the Bitter Lesson Swyx directly challenges compound specialist models using Rich Sutton's Bitter Lesson, arguing larger general models eventually subsume narrow experts. Lin counters with human societal specialization and the shift toward test-time inference scaling laws.38:38–44:23 · The hosts as informed peer 6/10 Fireworks Engineering Culture and the Cursor Partnership Case Study Swyx and Lin discuss Cursor's fast-apply speculative decoding pipeline. Lin details how co-engineering high-throughput inference stacks with aggressive developer teams validated their Fire Optimizer product.44:24–54:13 · The hosts as informed peer 7/10 Quantization Nuances, Open-Source Economics, and Multi-LoRA Serving Swyx brings up public benchmark callouts and breaks down OpenAI's P&L compute amortization. Lin criticizes unfair competitor benchmarking and explains Fireworks' multi-LoRA architecture sharing base model memory across hundreds of adapters.54:13–55:37 · The hosts as informed peer 2/10 Community Feedback, Discord Office Hours, and Hiring Expansion Standard wrap-up segment with Lin inviting developer feedback on Discord and announcing engineering hiring across multiple disciplines.2:08–6:58 · Guest teaching 7/10 Lin Qiao's Background at Meta and PyTorch Evolution Swyx sets up the prehistory of PyTorch and Meta's early GenAI efforts, referencing past guest Soumith Chintala. Lin delivers an extensive masterclass on scaling PyTorch from research to production across Meta's ubiquitous systems.6:58–9:49 · Guest teaching 6/10 Pivoting from Horizontal PyTorch Cloud to Generative AI Inference Swyx probes Fireworks' initial pivot from a horizontal PyTorch cloud to inference hosting. Lin explains the market dynamics of 2022, detailing why inference scales with the global population whereas training only scales with researchers.9:49–13:52 · Guest teaching 5/10 Focusing on App Developers and the Emergence of Llama Stack Swyx pushes back on Meta's Llama Stack, expressing skepticism that developer adoption will follow simply because Llama is open source. Lin acknowledges it is very early and emphasizes Fireworks' role in delivering direct community feedback.13:52–20:26 · Guest teaching 7/10 The Concept of Compound AI and Multi-Modal Optimization Swyx asks why Fireworks leaned heavily into Compound AI post-Series B. Lin articulates the architectural shift from one-size-fits-all distributed inference to multi-modal compound systems and three-dimensional optimization across quality, latency, and cost.20:26–24:13 · Guest teaching 6/10 Distributed Inference Architecture and Custom Kernel Acceleration Alessio and Swyx press Lin on what distributed inference actually entails versus raw GPU rental moats. Lin details custom CUDA kernels, disaggregated execution, regional routing, and hardware specialization across heterogeneous architectures.24:14–30:40 · Guest teaching 7/10 Declarative AI Architecture and Fireworks' New Reasoning Model Lin uses database SQL analogies to explain declarative AI system architectures over complex DAG pipelines. She reveals Fireworks' upcoming reasoning model trained to approach o1 reasoning quality.30:41–38:38 · Guest teaching 6/10 Inference Scaling Laws, Model Specialization, and the Bitter Lesson Swyx directly challenges compound specialist models using Rich Sutton's Bitter Lesson, arguing larger general models eventually subsume narrow experts. Lin counters with human societal specialization and the shift toward test-time inference scaling laws.38:38–44:23 · Guest teaching 5/10 Fireworks Engineering Culture and the Cursor Partnership Case Study Swyx and Lin discuss Cursor's fast-apply speculative decoding pipeline. Lin details how co-engineering high-throughput inference stacks with aggressive developer teams validated their Fire Optimizer product.44:24–54:13 · Guest teaching 6/10 Quantization Nuances, Open-Source Economics, and Multi-LoRA Serving Swyx brings up public benchmark callouts and breaks down OpenAI's P&L compute amortization. Lin criticizes unfair competitor benchmarking and explains Fireworks' multi-LoRA architecture sharing base model memory across hundreds of adapters.54:13–55:37 · Guest teaching 2/10 Community Feedback, Discord Office Hours, and Hiring Expansion Standard wrap-up segment with Lin inviting developer feedback on Discord and announcing engineering hiring across multiple disciplines.2:08–6:58 · Guest disagreement 1/10 Lin Qiao's Background at Meta and PyTorch Evolution Swyx sets up the prehistory of PyTorch and Meta's early GenAI efforts, referencing past guest Soumith Chintala. Lin delivers an extensive masterclass on scaling PyTorch from research to production across Meta's ubiquitous systems.6:58–9:49 · Guest disagreement 1/10 Pivoting from Horizontal PyTorch Cloud to Generative AI Inference Swyx probes Fireworks' initial pivot from a horizontal PyTorch cloud to inference hosting. Lin explains the market dynamics of 2022, detailing why inference scales with the global population whereas training only scales with researchers.9:49–13:52 · Guest disagreement 2/10 Focusing on App Developers and the Emergence of Llama Stack Swyx pushes back on Meta's Llama Stack, expressing skepticism that developer adoption will follow simply because Llama is open source. Lin acknowledges it is very early and emphasizes Fireworks' role in delivering direct community feedback.13:52–20:26 · Guest disagreement 1/10 The Concept of Compound AI and Multi-Modal Optimization Swyx asks why Fireworks leaned heavily into Compound AI post-Series B. Lin articulates the architectural shift from one-size-fits-all distributed inference to multi-modal compound systems and three-dimensional optimization across quality, latency, and cost.20:26–24:13 · Guest disagreement 2/10 Distributed Inference Architecture and Custom Kernel Acceleration Alessio and Swyx press Lin on what distributed inference actually entails versus raw GPU rental moats. Lin details custom CUDA kernels, disaggregated execution, regional routing, and hardware specialization across heterogeneous architectures.24:14–30:40 · Guest disagreement 1/10 Declarative AI Architecture and Fireworks' New Reasoning Model Lin uses database SQL analogies to explain declarative AI system architectures over complex DAG pipelines. She reveals Fireworks' upcoming reasoning model trained to approach o1 reasoning quality.30:41–38:38 · Guest disagreement 3/10 Inference Scaling Laws, Model Specialization, and the Bitter Lesson Swyx directly challenges compound specialist models using Rich Sutton's Bitter Lesson, arguing larger general models eventually subsume narrow experts. Lin counters with human societal specialization and the shift toward test-time inference scaling laws.38:38–44:23 · Guest disagreement 1/10 Fireworks Engineering Culture and the Cursor Partnership Case Study Swyx and Lin discuss Cursor's fast-apply speculative decoding pipeline. Lin details how co-engineering high-throughput inference stacks with aggressive developer teams validated their Fire Optimizer product.44:24–54:13 · Guest disagreement 3/10 Quantization Nuances, Open-Source Economics, and Multi-LoRA Serving Swyx brings up public benchmark callouts and breaks down OpenAI's P&L compute amortization. Lin criticizes unfair competitor benchmarking and explains Fireworks' multi-LoRA architecture sharing base model memory across hundreds of adapters.54:13–55:37 · Guest disagreement 0/10 Community Feedback, Discord Office Hours, and Hiring Expansion Standard wrap-up segment with Lin inviting developer feedback on Discord and announcing engineering hiring across multiple disciplines.2:08–6:58 · The hosts pushing back 1/10 Lin Qiao's Background at Meta and PyTorch Evolution Swyx sets up the prehistory of PyTorch and Meta's early GenAI efforts, referencing past guest Soumith Chintala. Lin delivers an extensive masterclass on scaling PyTorch from research to production across Meta's ubiquitous systems.6:58–9:49 · The hosts pushing back 1/10 Pivoting from Horizontal PyTorch Cloud to Generative AI Inference Swyx probes Fireworks' initial pivot from a horizontal PyTorch cloud to inference hosting. Lin explains the market dynamics of 2022, detailing why inference scales with the global population whereas training only scales with researchers.9:49–13:52 · The hosts pushing back 4/10 Focusing on App Developers and the Emergence of Llama Stack Swyx pushes back on Meta's Llama Stack, expressing skepticism that developer adoption will follow simply because Llama is open source. Lin acknowledges it is very early and emphasizes Fireworks' role in delivering direct community feedback.13:52–20:26 · The hosts pushing back 1/10 The Concept of Compound AI and Multi-Modal Optimization Swyx asks why Fireworks leaned heavily into Compound AI post-Series B. Lin articulates the architectural shift from one-size-fits-all distributed inference to multi-modal compound systems and three-dimensional optimization across quality, latency, and cost.20:26–24:13 · The hosts pushing back 3/10 Distributed Inference Architecture and Custom Kernel Acceleration Alessio and Swyx press Lin on what distributed inference actually entails versus raw GPU rental moats. Lin details custom CUDA kernels, disaggregated execution, regional routing, and hardware specialization across heterogeneous architectures.24:14–30:40 · The hosts pushing back 1/10 Declarative AI Architecture and Fireworks' New Reasoning Model Lin uses database SQL analogies to explain declarative AI system architectures over complex DAG pipelines. She reveals Fireworks' upcoming reasoning model trained to approach o1 reasoning quality.30:41–38:38 · The hosts pushing back 6/10 Inference Scaling Laws, Model Specialization, and the Bitter Lesson Swyx directly challenges compound specialist models using Rich Sutton's Bitter Lesson, arguing larger general models eventually subsume narrow experts. Lin counters with human societal specialization and the shift toward test-time inference scaling laws.38:38–44:23 · The hosts pushing back 2/10 Fireworks Engineering Culture and the Cursor Partnership Case Study Swyx and Lin discuss Cursor's fast-apply speculative decoding pipeline. Lin details how co-engineering high-throughput inference stacks with aggressive developer teams validated their Fire Optimizer product.44:24–54:13 · The hosts pushing back 3/10 Quantization Nuances, Open-Source Economics, and Multi-LoRA Serving Swyx brings up public benchmark callouts and breaks down OpenAI's P&L compute amortization. Lin criticizes unfair competitor benchmarking and explains Fireworks' multi-LoRA architecture sharing base model memory across hundreds of adapters.54:13–55:37 · The hosts pushing back 0/10 Community Feedback, Discord Office Hours, and Hiring Expansion Standard wrap-up segment with Lin inviting developer feedback on Discord and announcing engineering hiring across multiple disciplines.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 42.6% · guest 57.4%0:00 · the hosts 42.6% · guest 57.4%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 9.3% · guest 90.7%6:00 · the hosts 9.3% · guest 90.7%9:00 · the hosts 17% · guest 83%9:00 · the hosts 17% · guest 83%12:00 · the hosts 34.7% · guest 65.3%12:00 · the hosts 34.7% · guest 65.3%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 10.9% · guest 89.1%18:00 · the hosts 10.9% · guest 89.1%21:00 · the hosts 36.5% · guest 63.5%21:00 · the hosts 36.5% · guest 63.5%24:00 · the hosts 7.6% · guest 92.4%24:00 · the hosts 7.6% · guest 92.4%27:00 · the hosts 5.1% · guest 94.9%27:00 · the hosts 5.1% · guest 94.9%30:00 · the hosts 28% · guest 72%30:00 · the hosts 28% · guest 72%33:00 · the hosts 34.4% · guest 65.6%33:00 · the hosts 34.4% · guest 65.6%36:00 · the hosts 58.6% · guest 41.4%36:00 · the hosts 58.6% · guest 41.4%39:00 · the hosts 24.9% · guest 75.1%39:00 · the hosts 24.9% · guest 75.1%42:00 · the hosts 42.4% · guest 57.6%42:00 · the hosts 42.4% · guest 57.6%45:00 · the hosts 14.3% · guest 85.7%45:00 · the hosts 14.3% · guest 85.7%48:00 · the hosts 28.7% · guest 71.3%48:00 · the hosts 28.7% · guest 71.3%51:00 · the hosts 21.2% · guest 78.8%51:00 · the hosts 21.2% · guest 78.8%54:00 · the hosts 12.4% · guest 87.6%54:00 · the hosts 12.4% · guest 87.6%
Sharpest disagreement ▶ 46:54 Lin rejects competitor benchmarking tactics

Lin strongly criticizes a competitor publicly calling out Fireworks by name with biased benchmarks, insisting evaluations must be independently validated.

Hardest push from the hosts ▶ 34:09 Swyx invokes the Bitter Lesson against compound AI

Swyx refuses the premise that domain-specialized models have lasting moats, arguing a 10x larger model trained on generalized data will invalidate narrow expert models.

Biggest teaching moment ▶ 27:20 Lin contrasts declarative and imperative AI design

Lin systematically educates the hosts using relational database SQL optimization analogies to explain why declarative LLM orchestration will win developer adoption over brittle DAG pipelines.

The host holds their own ▶ 50:45 Swyx itemizes OpenAI's multi-billion dollar cost breakdown

Swyx demonstrates deep domain financial knowledge by rattling off the exact line items of OpenAI's training, inference compute, and research amortization expenses.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Lin Qiao's Background at Meta and PyTorch Evolution 4711 Swyx sets up the prehistory of PyTorch and Meta's early GenAI efforts, referencing past guest Soumith Chintala. Lin delivers an extensive masterclass on scaling PyTorch from research to production across Meta's ubiquitous systems.
Pivoting from Horizontal PyTorch Cloud to Generative AI Inference 4611 Swyx probes Fireworks' initial pivot from a horizontal PyTorch cloud to inference hosting. Lin explains the market dynamics of 2022, detailing why inference scales with the global population whereas training only scales with researchers.
Focusing on App Developers and the Emergence of Llama Stack 5524 Swyx pushes back on Meta's Llama Stack, expressing skepticism that developer adoption will follow simply because Llama is open source. Lin acknowledges it is very early and emphasizes Fireworks' role in delivering direct community feedback.
The Concept of Compound AI and Multi-Modal Optimization 4711 Swyx asks why Fireworks leaned heavily into Compound AI post-Series B. Lin articulates the architectural shift from one-size-fits-all distributed inference to multi-modal compound systems and three-dimensional optimization across quality, latency, and cost.
Distributed Inference Architecture and Custom Kernel Acceleration 5623 Alessio and Swyx press Lin on what distributed inference actually entails versus raw GPU rental moats. Lin details custom CUDA kernels, disaggregated execution, regional routing, and hardware specialization across heterogeneous architectures.
Declarative AI Architecture and Fireworks' New Reasoning Model 5711 Lin uses database SQL analogies to explain declarative AI system architectures over complex DAG pipelines. She reveals Fireworks' upcoming reasoning model trained to approach o1 reasoning quality.
Inference Scaling Laws, Model Specialization, and the Bitter Lesson 7636 Swyx directly challenges compound specialist models using Rich Sutton's Bitter Lesson, arguing larger general models eventually subsume narrow experts. Lin counters with human societal specialization and the shift toward test-time inference scaling laws.
Fireworks Engineering Culture and the Cursor Partnership Case Study 6512 Swyx and Lin discuss Cursor's fast-apply speculative decoding pipeline. Lin details how co-engineering high-throughput inference stacks with aggressive developer teams validated their Fire Optimizer product.
Quantization Nuances, Open-Source Economics, and Multi-LoRA Serving 7633 Swyx brings up public benchmark callouts and breaks down OpenAI's P&L compute amortization. Lin criticizes unfair competitor benchmarking and explains Fireworks' multi-LoRA architecture sharing base model memory across hundreds of adapters.
Community Feedback, Discord Office Hours, and Hiring Expansion 2200 Standard wrap-up segment with Lin inviting developer feedback on Discord and announcing engineering hiring across multiple disciplines.

Statements from this episode (13)

Assertion Supported
PyTorch was originally built for researchers without considering production requirements
“PyTorch actually started as the framework for researchers. Don't care about production at all.”
Lin Qiao Nov 25, 2024 ▶ 4:22
Insight
AI inference matters more than training because it scales with global population
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
Lin Qiao Nov 25, 2024 ▶ 8:59
Opinion
Shawn Wang doubts Meta's Llama Stack will win broad developer adoption
“I've been a little bit more doubtful on Lama stack. I think you've been more positive. Basically, it's just like the meta version of whatever HuggingFace offers, you know, or TensorRT, or BLM, or whatever the open source opportunity is. But like, to me, it's n…”
Shawn Wang Nov 25, 2024 ▶ 13:02
Insight
Customer inference workloads rarely align with foundation model training distributions
“The data distribution in their inference workload doesn't align with the data distribution in the training data for the model, right? It's a given, actually. If you think about this, because researchers have to guesstimate what is important, what's not importa…”
Lin Qiao Nov 25, 2024 ▶ 15:29
Insight
Compelling GenAI applications require compound AI systems spanning multiple modalities
“In order to really build a compelling application on top of JNI, we need a compound AI system. Compact AI system basically is going to have multiple models across modalities along with APIs, whether it's public APIs, internal proprietary APIs, storage systems,…”
Lin Qiao Nov 25, 2024 ▶ 19:41
Assertion Not checkable as stated
Fireworks AI runs custom acceleration kernels for almost all served models
“For almost for all models, for all large language models, all your models.”
Lin Qiao Nov 25, 2024 ▶ 23:24
Disclosure
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI. You should see that by the time we announce this o…”
Lin Qiao Nov 25, 2024 ▶ 29:26
Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao Nov 25, 2024 ▶ 33:45
Opinion
Pre-training on human data is hitting limits; synthetic data is required
“So I think on the data side, we're approaching the limit and the only data to increase that is synthetic generated data.”
Lin Qiao Nov 25, 2024 ▶ 35:05
Assertion Supported
Fireworks AI operates with a team of only forty people
“No, but only 40 people.”
Lin Qiao Nov 25, 2024 ▶ 39:12
Disclosure
Fireworks AI deployed a custom workload optimization stack specifically for Cursor
“We have a unique automation stack that is one size fits one. We actually deploy to cursor early on. Basically optimize for their specific workload, and that's a lot of juice to extract out of there, and we see success in, in that product is actually can be wid…”
Lin Qiao Nov 25, 2024 ▶ 43:26
Assertion Partly supported
Fireworks AI serves fine-tuned LoRA adapters at base model pricing
“We wrote multi LoRa last year, actually, and we actually have this function for a long time and many people have been using it, but it's not well known that, oh, if you find your model, you don't need to use on demand. If you find your model is LoRa. You can u…”
Lin Qiao Nov 25, 2024 ▶ 52:11
Assertion Supported
Fireworks AI's multi-LoRA system serves up to 1,000 adapters per base model
“One base model can sustain a hundred to a thousand LoRa adapters. And then basically all these different LoRa adapters can share the same, like direct the same traffic to the same base model where base model is dominating the cost.”
Lin Qiao Nov 25, 2024 ▶ 53:45
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.