Mar 27, 2025 · 59m · mad

Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI

Lin Qiao · 45m spoken Matt Turck · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, Fireworks AI CEO Lin Qiao discusses her journey from leading PyTorch development at Meta to building a fast, cost-effective AI inference and model customization platform. She shares deep insights on AI infrastructure economics, agentic workflows, open-source model ecosystems, and enterprise AI adoption.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 13.6% of the talking time here. How this is scored →

Matt as informed peer 3.5 Guest teaching 4.4 Guest disagreement 0.3 Matt pushing back 0.8
05100:0015:0030:0045:001:17–5:38 · Matt as informed peer 3/10 The Elevator Pitch and Mission of Fireworks AI Matt opens with a clear overview of Fireworks AI and invites Lin to share the elevator pitch and company origins. Lin outlines the platform focus on optimizing quality, speed, and cost for developers while sharing her early work on PyTorch at Meta.5:38–7:58 · Matt as informed peer 2/10 Overcoming Framework Unification Challenges and PyTorch Adoption Lin explains the technical challenge of unifying research flexibility and production optimization under PyTorch at Meta, noting that initial zipper attempts failed before rebuilding the backend over five years.7:58–14:54 · Matt as informed peer 5/10 The Industry Shift from Pre-GenAI to Post-GenAI Era Lin explains the shift from pre-GenAI model training to post-GenAI foundation model deployment. Matt demonstrates strong domain knowledge by synthesizing and playing back how foundation models shift technical complexity from model creation to deployment infrastructure.14:54–20:41 · Matt as informed peer 4/10 Meta's Open Source Legacy and PyTorch Governance Transition Matt cites PyTorch governance transitions and Contextual AI discussions, while Lin offers a minor correction that Meta retains a representative presence in governance rather than a total exit.20:41–27:24 · Matt as informed peer 3/10 Enterprise Adoption and Customer Use Cases: Cursor, Uber, and DoorDash Lin discusses customer adoption across Cursor, DoorDash, Uber, and Samsung, introducing the multi-dimensional search space (quality, speed, cost) targeted by Fire Optimizer.27:24–34:44 · Matt as informed peer 3/10 Advanced Optimization Techniques: Speculative Execution and Partitioning Lin delivers an in-depth technical explanation of speculative execution, quantization options, and disaggregated model partitioning. Matt asks concise clarifying questions on quantization terms.34:44–38:56 · Matt as informed peer 3/10 Agentic Development, Human-in-the-Loop, and Verifiable Rewards Lin covers human-in-the-loop vs automated optimization modes, detailing RL with verifiable rewards across programmatic domains like coding and subjective domains like design.38:56–44:23 · Matt as informed peer 3/10 Constrained Generation, Grammar Mode, and Multimodal Orchestration Lin details structured outputs, grammar mode (BNF format), and multimodal orchestration for domain-specific expert models like legal medical claim review.44:23–49:51 · Matt as informed peer 5/10 Low-Level Hardware Kernel Optimization and Multi-Cloud Infrastructure Matt challenges the industry narrative around enterprise demand for VPC and on-prem deployments versus public APIs. Lin reframes the issue, explaining that rapid model and hardware depreciation drives enterprise clients toward hosted single-tenant APIs.49:51–52:46 · Matt as informed peer 6/10 Inference Economics and Building a Sustainable Business Matt pushes Lin on how an infrastructure business stays sustainable in the face of rapidly falling unit inference costs. Lin firmly rejects the premise that falling costs threaten their business, arguing price drops expand total market volume 100x.52:46–55:04 · Matt as informed peer 2/10 Fireworks AI Roadmap and Model Customization Engine Lin shares forward roadmap priorities around Fire Optimizer, emphasizing quality customization that leverages customer proprietary data as a competitive moat.55:04–58:53 · Matt as informed peer 3/10 Industry Predictions: AI Agents, Open Models, and Ecosystem Complexity Lin predicts 2025 as the year of agents and open models, describing the intermediate integration challenges as a 'spaghetti layer' that Fireworks aims to simplify.1:17–5:38 · Guest teaching 3/10 The Elevator Pitch and Mission of Fireworks AI Matt opens with a clear overview of Fireworks AI and invites Lin to share the elevator pitch and company origins. Lin outlines the platform focus on optimizing quality, speed, and cost for developers while sharing her early work on PyTorch at Meta.5:38–7:58 · Guest teaching 5/10 Overcoming Framework Unification Challenges and PyTorch Adoption Lin explains the technical challenge of unifying research flexibility and production optimization under PyTorch at Meta, noting that initial zipper attempts failed before rebuilding the backend over five years.7:58–14:54 · Guest teaching 5/10 The Industry Shift from Pre-GenAI to Post-GenAI Era Lin explains the shift from pre-GenAI model training to post-GenAI foundation model deployment. Matt demonstrates strong domain knowledge by synthesizing and playing back how foundation models shift technical complexity from model creation to deployment infrastructure.14:54–20:41 · Guest teaching 3/10 Meta's Open Source Legacy and PyTorch Governance Transition Matt cites PyTorch governance transitions and Contextual AI discussions, while Lin offers a minor correction that Meta retains a representative presence in governance rather than a total exit.20:41–27:24 · Guest teaching 4/10 Enterprise Adoption and Customer Use Cases: Cursor, Uber, and DoorDash Lin discusses customer adoption across Cursor, DoorDash, Uber, and Samsung, introducing the multi-dimensional search space (quality, speed, cost) targeted by Fire Optimizer.27:24–34:44 · Guest teaching 6/10 Advanced Optimization Techniques: Speculative Execution and Partitioning Lin delivers an in-depth technical explanation of speculative execution, quantization options, and disaggregated model partitioning. Matt asks concise clarifying questions on quantization terms.34:44–38:56 · Guest teaching 5/10 Agentic Development, Human-in-the-Loop, and Verifiable Rewards Lin covers human-in-the-loop vs automated optimization modes, detailing RL with verifiable rewards across programmatic domains like coding and subjective domains like design.38:56–44:23 · Guest teaching 5/10 Constrained Generation, Grammar Mode, and Multimodal Orchestration Lin details structured outputs, grammar mode (BNF format), and multimodal orchestration for domain-specific expert models like legal medical claim review.44:23–49:51 · Guest teaching 4/10 Low-Level Hardware Kernel Optimization and Multi-Cloud Infrastructure Matt challenges the industry narrative around enterprise demand for VPC and on-prem deployments versus public APIs. Lin reframes the issue, explaining that rapid model and hardware depreciation drives enterprise clients toward hosted single-tenant APIs.49:51–52:46 · Guest teaching 5/10 Inference Economics and Building a Sustainable Business Matt pushes Lin on how an infrastructure business stays sustainable in the face of rapidly falling unit inference costs. Lin firmly rejects the premise that falling costs threaten their business, arguing price drops expand total market volume 100x.52:46–55:04 · Guest teaching 4/10 Fireworks AI Roadmap and Model Customization Engine Lin shares forward roadmap priorities around Fire Optimizer, emphasizing quality customization that leverages customer proprietary data as a competitive moat.55:04–58:53 · Guest teaching 4/10 Industry Predictions: AI Agents, Open Models, and Ecosystem Complexity Lin predicts 2025 as the year of agents and open models, describing the intermediate integration challenges as a 'spaghetti layer' that Fireworks aims to simplify.1:17–5:38 · Guest disagreement 0/10 The Elevator Pitch and Mission of Fireworks AI Matt opens with a clear overview of Fireworks AI and invites Lin to share the elevator pitch and company origins. Lin outlines the platform focus on optimizing quality, speed, and cost for developers while sharing her early work on PyTorch at Meta.5:38–7:58 · Guest disagreement 0/10 Overcoming Framework Unification Challenges and PyTorch Adoption Lin explains the technical challenge of unifying research flexibility and production optimization under PyTorch at Meta, noting that initial zipper attempts failed before rebuilding the backend over five years.7:58–14:54 · Guest disagreement 0/10 The Industry Shift from Pre-GenAI to Post-GenAI Era Lin explains the shift from pre-GenAI model training to post-GenAI foundation model deployment. Matt demonstrates strong domain knowledge by synthesizing and playing back how foundation models shift technical complexity from model creation to deployment infrastructure.14:54–20:41 · Guest disagreement 1/10 Meta's Open Source Legacy and PyTorch Governance Transition Matt cites PyTorch governance transitions and Contextual AI discussions, while Lin offers a minor correction that Meta retains a representative presence in governance rather than a total exit.20:41–27:24 · Guest disagreement 0/10 Enterprise Adoption and Customer Use Cases: Cursor, Uber, and DoorDash Lin discusses customer adoption across Cursor, DoorDash, Uber, and Samsung, introducing the multi-dimensional search space (quality, speed, cost) targeted by Fire Optimizer.27:24–34:44 · Guest disagreement 0/10 Advanced Optimization Techniques: Speculative Execution and Partitioning Lin delivers an in-depth technical explanation of speculative execution, quantization options, and disaggregated model partitioning. Matt asks concise clarifying questions on quantization terms.34:44–38:56 · Guest disagreement 0/10 Agentic Development, Human-in-the-Loop, and Verifiable Rewards Lin covers human-in-the-loop vs automated optimization modes, detailing RL with verifiable rewards across programmatic domains like coding and subjective domains like design.38:56–44:23 · Guest disagreement 0/10 Constrained Generation, Grammar Mode, and Multimodal Orchestration Lin details structured outputs, grammar mode (BNF format), and multimodal orchestration for domain-specific expert models like legal medical claim review.44:23–49:51 · Guest disagreement 1/10 Low-Level Hardware Kernel Optimization and Multi-Cloud Infrastructure Matt challenges the industry narrative around enterprise demand for VPC and on-prem deployments versus public APIs. Lin reframes the issue, explaining that rapid model and hardware depreciation drives enterprise clients toward hosted single-tenant APIs.49:51–52:46 · Guest disagreement 1/10 Inference Economics and Building a Sustainable Business Matt pushes Lin on how an infrastructure business stays sustainable in the face of rapidly falling unit inference costs. Lin firmly rejects the premise that falling costs threaten their business, arguing price drops expand total market volume 100x.52:46–55:04 · Guest disagreement 0/10 Fireworks AI Roadmap and Model Customization Engine Lin shares forward roadmap priorities around Fire Optimizer, emphasizing quality customization that leverages customer proprietary data as a competitive moat.55:04–58:53 · Guest disagreement 0/10 Industry Predictions: AI Agents, Open Models, and Ecosystem Complexity Lin predicts 2025 as the year of agents and open models, describing the intermediate integration challenges as a 'spaghetti layer' that Fireworks aims to simplify.1:17–5:38 · Matt pushing back 0/10 The Elevator Pitch and Mission of Fireworks AI Matt opens with a clear overview of Fireworks AI and invites Lin to share the elevator pitch and company origins. Lin outlines the platform focus on optimizing quality, speed, and cost for developers while sharing her early work on PyTorch at Meta.5:38–7:58 · Matt pushing back 0/10 Overcoming Framework Unification Challenges and PyTorch Adoption Lin explains the technical challenge of unifying research flexibility and production optimization under PyTorch at Meta, noting that initial zipper attempts failed before rebuilding the backend over five years.7:58–14:54 · Matt pushing back 1/10 The Industry Shift from Pre-GenAI to Post-GenAI Era Lin explains the shift from pre-GenAI model training to post-GenAI foundation model deployment. Matt demonstrates strong domain knowledge by synthesizing and playing back how foundation models shift technical complexity from model creation to deployment infrastructure.14:54–20:41 · Matt pushing back 2/10 Meta's Open Source Legacy and PyTorch Governance Transition Matt cites PyTorch governance transitions and Contextual AI discussions, while Lin offers a minor correction that Meta retains a representative presence in governance rather than a total exit.20:41–27:24 · Matt pushing back 0/10 Enterprise Adoption and Customer Use Cases: Cursor, Uber, and DoorDash Lin discusses customer adoption across Cursor, DoorDash, Uber, and Samsung, introducing the multi-dimensional search space (quality, speed, cost) targeted by Fire Optimizer.27:24–34:44 · Matt pushing back 0/10 Advanced Optimization Techniques: Speculative Execution and Partitioning Lin delivers an in-depth technical explanation of speculative execution, quantization options, and disaggregated model partitioning. Matt asks concise clarifying questions on quantization terms.34:44–38:56 · Matt pushing back 0/10 Agentic Development, Human-in-the-Loop, and Verifiable Rewards Lin covers human-in-the-loop vs automated optimization modes, detailing RL with verifiable rewards across programmatic domains like coding and subjective domains like design.38:56–44:23 · Matt pushing back 0/10 Constrained Generation, Grammar Mode, and Multimodal Orchestration Lin details structured outputs, grammar mode (BNF format), and multimodal orchestration for domain-specific expert models like legal medical claim review.44:23–49:51 · Matt pushing back 3/10 Low-Level Hardware Kernel Optimization and Multi-Cloud Infrastructure Matt challenges the industry narrative around enterprise demand for VPC and on-prem deployments versus public APIs. Lin reframes the issue, explaining that rapid model and hardware depreciation drives enterprise clients toward hosted single-tenant APIs.49:51–52:46 · Matt pushing back 4/10 Inference Economics and Building a Sustainable Business Matt pushes Lin on how an infrastructure business stays sustainable in the face of rapidly falling unit inference costs. Lin firmly rejects the premise that falling costs threaten their business, arguing price drops expand total market volume 100x.52:46–55:04 · Matt pushing back 0/10 Fireworks AI Roadmap and Model Customization Engine Lin shares forward roadmap priorities around Fire Optimizer, emphasizing quality customization that leverages customer proprietary data as a competitive moat.55:04–58:53 · Matt pushing back 0/10 Industry Predictions: AI Agents, Open Models, and Ecosystem Complexity Lin predicts 2025 as the year of agents and open models, describing the intermediate integration challenges as a 'spaghetti layer' that Fireworks aims to simplify.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 42.7% · guest 57.3%0:00 · Matt 42.7% · guest 57.3%3:00 · Matt 4.6% · guest 95.4%3:00 · Matt 4.6% · guest 95.4%6:00 · Matt 14.4% · guest 85.6%6:00 · Matt 14.4% · guest 85.6%9:00 · Matt 3.1% · guest 96.9%9:00 · Matt 3.1% · guest 96.9%12:00 · Matt 25% · guest 75%12:00 · Matt 25% · guest 75%15:00 · Matt 32.7% · guest 67.3%15:00 · Matt 32.7% · guest 67.3%18:00 · Matt 18.6% · guest 81.4%18:00 · Matt 18.6% · guest 81.4%21:00 · Matt 10.1% · guest 89.9%21:00 · Matt 10.1% · guest 89.9%24:00 · Matt 8.6% · guest 91.4%24:00 · Matt 8.6% · guest 91.4%27:00 · Matt 3.7% · guest 96.3%27:00 · Matt 3.7% · guest 96.3%30:00 · Matt 3.6% · guest 96.4%30:00 · Matt 3.6% · guest 96.4%33:00 · Matt 7.7% · guest 92.3%33:00 · Matt 7.7% · guest 92.3%36:00 · Matt 5.5% · guest 94.5%36:00 · Matt 5.5% · guest 94.5%39:00 · Matt 3.2% · guest 96.8%39:00 · Matt 3.2% · guest 96.8%42:00 · Matt 7.8% · guest 92.2%42:00 · Matt 7.8% · guest 92.2%45:00 · Matt 24.1% · guest 75.9%45:00 · Matt 24.1% · guest 75.9%48:00 · Matt 15.8% · guest 84.2%48:00 · Matt 15.8% · guest 84.2%51:00 · Matt 13.3% · guest 86.7%51:00 · Matt 13.3% · guest 86.7%54:00 · Matt 8.1% · guest 91.9%54:00 · Matt 8.1% · guest 91.9%57:00 · Matt 20.4% · guest 79.6%57:00 · Matt 20.4% · guest 79.6%
Sharpest disagreement ▶ 50:15 Lin rejects premise on declining inference prices

Lin forcefully pushes back against the premise that falling inference prices threaten their revenue, asserting that infrastructure must drop by orders of magnitude to unlock sustainable application ROIs.

Hardest push from Matt ▶ 49:50 Matt challenges inference provider unit economics

Matt presses Lin directly on business sustainability, challenging how an infrastructure vendor maintains a durable business when unit costs continually collapse.

Biggest teaching moment ▶ 27:40 Lin breakdown of speculative execution mechanics

Lin provides a comprehensive technical masterclass explaining how pairing draft and target models yields large-model quality at small-model speeds.

Matt holds his own ▶ 12:50 Matt synthesizes the fundamental pre vs post GenAI paradigm

Matt crisply synthesizes Lin's narrative, demonstrating strong domain mastery by highlighting how foundation models shifted AI complexity from model training to post-training deployment.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Elevator Pitch and Mission of Fireworks AI 3300 Matt opens with a clear overview of Fireworks AI and invites Lin to share the elevator pitch and company origins. Lin outlines the platform focus on optimizing quality, speed, and cost for developers while sharing her early work on PyTorch at Meta.
Overcoming Framework Unification Challenges and PyTorch Adoption 2500 Lin explains the technical challenge of unifying research flexibility and production optimization under PyTorch at Meta, noting that initial zipper attempts failed before rebuilding the backend over five years.
The Industry Shift from Pre-GenAI to Post-GenAI Era 5501 Lin explains the shift from pre-GenAI model training to post-GenAI foundation model deployment. Matt demonstrates strong domain knowledge by synthesizing and playing back how foundation models shift technical complexity from model creation to deployment infrastructure.
Meta's Open Source Legacy and PyTorch Governance Transition 4312 Matt cites PyTorch governance transitions and Contextual AI discussions, while Lin offers a minor correction that Meta retains a representative presence in governance rather than a total exit.
Enterprise Adoption and Customer Use Cases: Cursor, Uber, and DoorDash 3400 Lin discusses customer adoption across Cursor, DoorDash, Uber, and Samsung, introducing the multi-dimensional search space (quality, speed, cost) targeted by Fire Optimizer.
Advanced Optimization Techniques: Speculative Execution and Partitioning 3600 Lin delivers an in-depth technical explanation of speculative execution, quantization options, and disaggregated model partitioning. Matt asks concise clarifying questions on quantization terms.
Agentic Development, Human-in-the-Loop, and Verifiable Rewards 3500 Lin covers human-in-the-loop vs automated optimization modes, detailing RL with verifiable rewards across programmatic domains like coding and subjective domains like design.
Constrained Generation, Grammar Mode, and Multimodal Orchestration 3500 Lin details structured outputs, grammar mode (BNF format), and multimodal orchestration for domain-specific expert models like legal medical claim review.
Low-Level Hardware Kernel Optimization and Multi-Cloud Infrastructure 5413 Matt challenges the industry narrative around enterprise demand for VPC and on-prem deployments versus public APIs. Lin reframes the issue, explaining that rapid model and hardware depreciation drives enterprise clients toward hosted single-tenant APIs.
Inference Economics and Building a Sustainable Business 6514 Matt pushes Lin on how an infrastructure business stays sustainable in the face of rapidly falling unit inference costs. Lin firmly rejects the premise that falling costs threaten their business, arguing price drops expand total market volume 100x.
Fireworks AI Roadmap and Model Customization Engine 2400 Lin shares forward roadmap priorities around Fire Optimizer, emphasizing quality customization that leverages customer proprietary data as a competitive moat.
Industry Predictions: AI Agents, Open Models, and Ecosystem Complexity 3400 Lin predicts 2025 as the year of agents and open models, describing the intermediate integration challenges as a 'spaghetti layer' that Fireworks aims to simplify.

Statements from this episode (15)

Assertion Supported
Meta historically maintained three separate AI frameworks for mobile, research, and production
“Even within Mata, there are three different flavors. One for mobile, one for research, one for production.”
Lin Qiao Mar 27, 2025 ▶ 4:14
Insight
Lin Qiao: AI frameworks must reconcile researcher flexibility with strict production cost and latency constraints
“For researchers, you want the flexibility. You want ease of use. You want them to just think about what's possible, right? And for production, it's a constraint problem solving. As in, you have latency budget, you have cost budget you want to scale, you want t…”
Lin Qiao Mar 27, 2025 ▶ 4:44
Assertion Not checkable as stated
Meta spent five years rebuilding PyTorch's backend for internal scale
“It took us five years. Took us five years to get the stage supporting almost all internal needs using deep learning and mass and massive scale.”
Lin Qiao Mar 27, 2025 ▶ 6:48
Assertion Supported
Lin Qiao: OpenAI switched completely from TensorFlow to PyTorch
“OpenAI switched to use PyTorch fully.”
Lin Qiao Mar 27, 2025 ▶ 8:00
Insight
Lin Qiao: PyTorch's primary success lesson is that simplicity scales
“I think one of the biggest success we saw from the PyTorch experience is simplicity scales.”
Lin Qiao Mar 27, 2025 ▶ 16:26
Assertion Not checkable as stated
Lin Qiao: Meta had hundreds of engineers building PyTorch and its infrastructure
“We have hundreds of engineers building PyTorch and infrastructure around PyTorch, but at the same time, I believe PyTorch within Meta probably has thousands of users.”
Lin Qiao Mar 27, 2025 ▶ 19:23
Assertion Not checkable as stated
Fireworks AI improved speculative execution hit rates from 30% to 90%
“We have seen cases improving the prediction hit from 30% to 90%, and that's huge speed.”
Lin Qiao Mar 27, 2025 ▶ 29:25
Insight
LLM prompt processing is compute-bound; next-token generation is memory-bound
“Prompt processing is bottlenecked by computation, and generating next, predicting next token is bottlenecked by memory bandwidth.”
Lin Qiao Mar 27, 2025 ▶ 30:42
Assertion Not checkable as stated
Lin Qiao: Fireworks AI's optimization space has over 80,000 options
“And all these different options add up together, it can lead into more than 80,000 possible, possible way to optimize.”
Lin Qiao Mar 27, 2025 ▶ 32:38
Assertion Not checkable as stated
DeepSeek runs each single model replica across more than 300 GPUs
“DeepSeq actually that company itself was running and still running this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas.”
Lin Qiao Mar 27, 2025 ▶ 36:10
Assertion Contradicted
Fireworks AI was first to enable function calling for DeepSeek models
“We have been working on function for calling for a long time, and we are the first one to enable function calling for deep seek models.”
Lin Qiao Mar 27, 2025 ▶ 43:58
Prediction Not checkable as stated
Lin Qiao predicts a 10x AI cost reduction yields 100x more applications
“If this bar can be lowered by 10 times, you can imagine there's so many more, it will be hundred times more applications enter the, this arena to create a brand new experience to end consumers and prosumers. And by that, we'll see a much bigger consumption acr…”
Lin Qiao Mar 27, 2025 ▶ 51:56
Insight
For AI applications, moats lie in curated data rather than user experience
“Their mode is probably not the user experience, but because it's very easy to copy. Anyone can study the product and copy. Their mode is data.”
Lin Qiao Mar 27, 2025 ▶ 54:20
Assertion Supported
Over 500 DeepSeek model variants hit Hugging Face within a month
“DeepSeq for example, just within one month of releasing their new models, There are, despite DeepSeq model, extremely hard to tune and optimize, extremely hard. There are 500, more than 500 variants published on Hugging Face, optimizing for local device, optim…”
Lin Qiao Mar 27, 2025 ▶ 56:29
Prediction Not checkable as stated
Lin Qiao: The future of AI modeling belongs to open-source models
“The future of the future of modeling sits on open model side. And I believe that side is gonna be much more active in creating those hundreds or maybe thousands of expert models that is specialized delivering much better quality in certain domain.”
Lin Qiao Mar 27, 2025 ▶ 57:29
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.