Jul 20, 2026 · 1h 28m · news

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Lin Qiao · 1h 2m spoken Harry Stebbings · 14m spoken

Clips from this episode (6)

Short vertical cuts produced from the tape, captions burned in. Where a cut lands on a statement from the ledger, its card says so.

0:00 / 1:08
the statementLin Qiao: Corporate Uniqueness and User Data Intelligence Cannot Be Replicated ExternallyLin Qiaoread it →
0:00 / 0:36
the statementLin Qiao: Product-market fit no longer guarantees business durability in AI SaaSLin Qiaoread it →
0:00 / 0:42
the statementLin Qiao: A single-company monopoly on AI makes no senseLin Qiaoread it →
0:00 / 0:54
the statementQiao: Fireworks daily token count could grow 20x-100x by next yearLin Qiaoread it →
0:00 / 0:56
the statementLin Qiao: token costs will fall drastically and cheaper infrastructure will invite far more usageLin Qiaoread it →
0:00 / 0:38
the statementQiao: Fireworks AI expects to at least double ARR by year-endOpenLin Qiaoread it →
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of 20VC, host Harry Stebbings interviews Lin Qiao, co-founder and CEO of Fireworks AI, exploring the economics of generative AI, the strategic imperative of open-source models over monolithic AGI, physical infrastructure constraints, and insights on scaling high-growth AI startups.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Harry holds 17.8% of the talking time here. How this is scored →

Harry as informed peer 5.7 Guest teaching 5.3 Guest disagreement 2.2 Harry pushing back 4.6
05100:0020:0040:001:00:001:20:000:57–3:46 · Harry as informed peer 4/10 Lin Qiao's Career Journey and Founding Fireworks AI Harry sets up the interview citing background diligence with Eric Vishria and Dima, asking Lin about starting a company at 48. Lin shares her technical background in distributed systems and seven years at Meta learning people management.3:46–9:14 · Harry as informed peer 6/10 Positioning Fireworks AI: Inference, Specialized Intelligence, and Enterprise Data Harry questions why inference and specialized intelligence are not commodities and pushes Anthropic's enterprise moat as a counterexample. Lin counters with Jensen Huang's insight that every business is specialized because proprietary data is private.9:14–14:57 · Harry as informed peer 7/10 Open-Source Model Quality and the Viability of Frontier Labs Harry asks if frontier model companies are dramatically overvalued if open-source models solve 90% of enterprise workflows at 15x lower cost. Lin details how open models allow enterprises to retain weight control and hill-climb on private data.14:57–19:45 · Harry as informed peer 6/10 AI Economics: Scaling to Bankruptcy, and One-Size-Fits-One Optimization Lin outlines the concept of startups and incumbents 'scaling into bankruptcy' with high inference COGS. Harry pushes back asking if OpenAI's aggressively price-reduced frontier models eliminate the open-weight advantage.19:45–23:04 · Harry as informed peer 6/10 Geopolitical AI Risk, Guardrails, and Vertical Model Adoption Harry raises national security concerns regarding Chinese dominance on open LLM leaderboards and potential export restrictions. Lin argues that every company needs custom guardrails regardless of origin and open ecosystems naturally decentralize.23:04–27:14 · Harry as informed peer 6/10 Disruption of SaaS Application Moats and Rapid Model Iteration Harry challenges the idea that SaaS moats have collapsed by citing long enterprise legal sales cycles at portfolio company Leya vs Harvey. Lin points out that workflow harnesses must be co-trained with underlying specialized models to prevent errors.27:14–30:23 · Harry as informed peer 5/10 Base Model Step Functions, Government AI Infrastructure, and Monopoly Risks Harry asks if foundation AI should be treated as national infrastructure or partially government-owned. Lin warns against single-company monopolies, comparing base intelligence to utilities while defending specialized intelligence.30:23–33:10 · Harry as informed peer 7/10 The AI Routing Layer and Autonomous Self-Evolving Systems Harry directly probes whether routing layers like OpenRouter or Requesty will be disintermediated if companies build autonomous routing internally. Lin concedes that standalone routers might be redundant once self-evolving automated routing matures.37:33–40:16 · Harry as informed peer 6/10 Transitioning Beyond Coding to Enterprise Co-Work and Consumer AI Harry asks bluntly about customer concentration risk if Cursor churns. Lin details the diversification from coding to enterprise co-work and consumer recommendation engines across multiple verticals.40:16–43:22 · Harry as informed peer 6/10 Token Volume Explosions and AI Supply Chain Bottlenecks Harry reacts incredulously to Lin predicting a 20-100x explosion in token count, noting that Capex bubble fears would be groundless under that demand. Lin explains that physical supply chain bottlenecks in energy and chips throttle delivery.43:22–47:19 · Harry as informed peer 7/10 Startup Agility, Platform Specialization, and Enterprise AI Expenditure Harry presses on whether AI infrastructure players must go full-stack and why Nvidia's Jensen Huang does not expand into Fireworks' platform layer. Lin clarifies that Nvidia enters models purely to clear supply bottlenecks without building cloud inference.47:19–50:41 · Harry as informed peer 6/10 Long-Term Token Economics and Infrastructure Cost Deflation Harry references Marc Benioff's developer spend data and asks whether token prices will truly fall when they have remained stubbornly high. Lin explains how task-level token efficiency and hardware supply easing will lead to a 10x cost reduction driving 100x usage.50:41–53:30 · Harry as informed peer 6/10 Quality-First Inference, Zero KLD, and Bit-Wise Model Equivalence Harry asks if Fireworks is viewed as more expensive than low-cost providers like Together AI. Lin reframes the pricing dynamic, introducing the concept of zero KLD and exact bit-wise equivalence between training and inference.53:30–55:59 · Harry as informed peer 6/10 Gross Margins in AI Infrastructure and Growth-Stage Priorities Harry asks if 30-40% gross margins represent the permanent new normal for AI infrastructure. Lin explains that low initial margins represent an intentional constraint choice to prioritize hypergrowth and geographic expansion over premature optimization.55:59–58:58 · Harry as informed peer 6/10 Custom Data Center Architecture and Heterogeneous Hardware Clusters Lin outlines hardware disaggregation between FLOPs-intensive pre-fill and SRAM-intensive token generation using Groq ASICs. Harry connects this to his background diligence with Groq CEO Jonathan Ross.58:58–1:01:04 · Harry as informed peer 6/10 Physical Data Center Execution Constraints and Global Competition Harry asks whether rapid physical infrastructure buildouts give China a structural advantage over US permitting hurdles. Lin agrees on civil engineering velocity but notes specialized operations talent is still scarce globally.1:01:04–1:04:57 · Harry as informed peer 6/10 Hardware and Model Depreciation Cycles and Long-Term Margin Strategy Lin breaks down how hardware depreciation cycles have compressed from six years to annual multi-SKU releases, mismatching rapid model obsolescence. Harry synthesizes the implication that model speed outstrips hardware depreciation schedules.1:04:57–1:07:11 · Harry as informed peer 5/10 Market Dynamics and Owning Intelligence vs. Renting Harry asks whether inference markets are winner-take-all like ridesharing or multi-player like hyperscale cloud. Lin argues that enterprises inevitably transition from renting general intelligence to owning and optimizing proprietary models.1:07:11–1:12:26 · Harry as informed peer 6/10 Sovereign Models and National AI Independence Harry asks why startups avoid chip manufacturing if Meta, OpenAI, Anthropic, and DeepSeek build their own custom silicon. Lin educates him on silicon workload stabilization: taping out chips requires unchanging model architectures that premature startups cannot freeze.1:12:26–1:15:05 · Harry as informed peer 6/10 System Level Infrastructure and Architecture Bottlenecks Harry cites Groq's Jonathan Ross claiming HBM memory bandwidth is the primary bottleneck. Lin counters that the broader architectural bottleneck is system-level co-design across 10-trillion parameter clusters, celebrating Fireworks hitting $800M run-rate potential.1:15:05–1:18:44 · Harry as informed peer 5/10 Strategic Executive Hiring and AI Leadership Traits Harry inquires into the hiring of former Salesforce president George Hu. Lin explains how maintaining high curiosity alongside deep executive seniority allowed them to accelerate go-to-market without bureaucratic ossification.1:18:44–1:21:17 · Harry as informed peer 4/10 Quick Fire: Scaling Velocity and Culture of Extreme Ownership In the quick fire round, Harry shares 20VC hiring profiles while Lin articulates the culture of extreme ownership and hiring talent that rejects artificial organizational boundaries.1:21:17–1:25:47 · Harry as informed peer 4/10 Quick Fire: Leadership Lessons from Jensen Huang Lin reflects on leadership lessons learned from Jensen Huang, highlighting instant responsiveness and context-driven decision making. She shares her past mistake of underinvesting in marketing.1:25:47–1:28:07 · Harry as informed peer 4/10 Quick Fire: Untapped Enterprise Segments Lin outlines untapped opportunities in traditional enterprise segments like insurance and banking, concluding with her core thesis that every enterprise must eventually own its software intelligence stack.0:57–3:46 · Guest teaching 2/10 Lin Qiao's Career Journey and Founding Fireworks AI Harry sets up the interview citing background diligence with Eric Vishria and Dima, asking Lin about starting a company at 48. Lin shares her technical background in distributed systems and seven years at Meta learning people management.3:46–9:14 · Guest teaching 6/10 Positioning Fireworks AI: Inference, Specialized Intelligence, and Enterprise Data Harry questions why inference and specialized intelligence are not commodities and pushes Anthropic's enterprise moat as a counterexample. Lin counters with Jensen Huang's insight that every business is specialized because proprietary data is private.9:14–14:57 · Guest teaching 6/10 Open-Source Model Quality and the Viability of Frontier Labs Harry asks if frontier model companies are dramatically overvalued if open-source models solve 90% of enterprise workflows at 15x lower cost. Lin details how open models allow enterprises to retain weight control and hill-climb on private data.14:57–19:45 · Guest teaching 6/10 AI Economics: Scaling to Bankruptcy, and One-Size-Fits-One Optimization Lin outlines the concept of startups and incumbents 'scaling into bankruptcy' with high inference COGS. Harry pushes back asking if OpenAI's aggressively price-reduced frontier models eliminate the open-weight advantage.19:45–23:04 · Guest teaching 5/10 Geopolitical AI Risk, Guardrails, and Vertical Model Adoption Harry raises national security concerns regarding Chinese dominance on open LLM leaderboards and potential export restrictions. Lin argues that every company needs custom guardrails regardless of origin and open ecosystems naturally decentralize.23:04–27:14 · Guest teaching 6/10 Disruption of SaaS Application Moats and Rapid Model Iteration Harry challenges the idea that SaaS moats have collapsed by citing long enterprise legal sales cycles at portfolio company Leya vs Harvey. Lin points out that workflow harnesses must be co-trained with underlying specialized models to prevent errors.27:14–30:23 · Guest teaching 4/10 Base Model Step Functions, Government AI Infrastructure, and Monopoly Risks Harry asks if foundation AI should be treated as national infrastructure or partially government-owned. Lin warns against single-company monopolies, comparing base intelligence to utilities while defending specialized intelligence.30:23–33:10 · Guest teaching 6/10 The AI Routing Layer and Autonomous Self-Evolving Systems Harry directly probes whether routing layers like OpenRouter or Requesty will be disintermediated if companies build autonomous routing internally. Lin concedes that standalone routers might be redundant once self-evolving automated routing matures.37:33–40:16 · Guest teaching 5/10 Transitioning Beyond Coding to Enterprise Co-Work and Consumer AI Harry asks bluntly about customer concentration risk if Cursor churns. Lin details the diversification from coding to enterprise co-work and consumer recommendation engines across multiple verticals.40:16–43:22 · Guest teaching 6/10 Token Volume Explosions and AI Supply Chain Bottlenecks Harry reacts incredulously to Lin predicting a 20-100x explosion in token count, noting that Capex bubble fears would be groundless under that demand. Lin explains that physical supply chain bottlenecks in energy and chips throttle delivery.43:22–47:19 · Guest teaching 5/10 Startup Agility, Platform Specialization, and Enterprise AI Expenditure Harry presses on whether AI infrastructure players must go full-stack and why Nvidia's Jensen Huang does not expand into Fireworks' platform layer. Lin clarifies that Nvidia enters models purely to clear supply bottlenecks without building cloud inference.47:19–50:41 · Guest teaching 6/10 Long-Term Token Economics and Infrastructure Cost Deflation Harry references Marc Benioff's developer spend data and asks whether token prices will truly fall when they have remained stubbornly high. Lin explains how task-level token efficiency and hardware supply easing will lead to a 10x cost reduction driving 100x usage.50:41–53:30 · Guest teaching 7/10 Quality-First Inference, Zero KLD, and Bit-Wise Model Equivalence Harry asks if Fireworks is viewed as more expensive than low-cost providers like Together AI. Lin reframes the pricing dynamic, introducing the concept of zero KLD and exact bit-wise equivalence between training and inference.53:30–55:59 · Guest teaching 6/10 Gross Margins in AI Infrastructure and Growth-Stage Priorities Harry asks if 30-40% gross margins represent the permanent new normal for AI infrastructure. Lin explains that low initial margins represent an intentional constraint choice to prioritize hypergrowth and geographic expansion over premature optimization.55:59–58:58 · Guest teaching 7/10 Custom Data Center Architecture and Heterogeneous Hardware Clusters Lin outlines hardware disaggregation between FLOPs-intensive pre-fill and SRAM-intensive token generation using Groq ASICs. Harry connects this to his background diligence with Groq CEO Jonathan Ross.58:58–1:01:04 · Guest teaching 5/10 Physical Data Center Execution Constraints and Global Competition Harry asks whether rapid physical infrastructure buildouts give China a structural advantage over US permitting hurdles. Lin agrees on civil engineering velocity but notes specialized operations talent is still scarce globally.1:01:04–1:04:57 · Guest teaching 7/10 Hardware and Model Depreciation Cycles and Long-Term Margin Strategy Lin breaks down how hardware depreciation cycles have compressed from six years to annual multi-SKU releases, mismatching rapid model obsolescence. Harry synthesizes the implication that model speed outstrips hardware depreciation schedules.1:04:57–1:07:11 · Guest teaching 5/10 Market Dynamics and Owning Intelligence vs. Renting Harry asks whether inference markets are winner-take-all like ridesharing or multi-player like hyperscale cloud. Lin argues that enterprises inevitably transition from renting general intelligence to owning and optimizing proprietary models.1:07:11–1:12:26 · Guest teaching 7/10 Sovereign Models and National AI Independence Harry asks why startups avoid chip manufacturing if Meta, OpenAI, Anthropic, and DeepSeek build their own custom silicon. Lin educates him on silicon workload stabilization: taping out chips requires unchanging model architectures that premature startups cannot freeze.1:12:26–1:15:05 · Guest teaching 6/10 System Level Infrastructure and Architecture Bottlenecks Harry cites Groq's Jonathan Ross claiming HBM memory bandwidth is the primary bottleneck. Lin counters that the broader architectural bottleneck is system-level co-design across 10-trillion parameter clusters, celebrating Fireworks hitting $800M run-rate potential.1:15:05–1:18:44 · Guest teaching 4/10 Strategic Executive Hiring and AI Leadership Traits Harry inquires into the hiring of former Salesforce president George Hu. Lin explains how maintaining high curiosity alongside deep executive seniority allowed them to accelerate go-to-market without bureaucratic ossification.1:18:44–1:21:17 · Guest teaching 3/10 Quick Fire: Scaling Velocity and Culture of Extreme Ownership In the quick fire round, Harry shares 20VC hiring profiles while Lin articulates the culture of extreme ownership and hiring talent that rejects artificial organizational boundaries.1:21:17–1:25:47 · Guest teaching 4/10 Quick Fire: Leadership Lessons from Jensen Huang Lin reflects on leadership lessons learned from Jensen Huang, highlighting instant responsiveness and context-driven decision making. She shares her past mistake of underinvesting in marketing.1:25:47–1:28:07 · Guest teaching 3/10 Quick Fire: Untapped Enterprise Segments Lin outlines untapped opportunities in traditional enterprise segments like insurance and banking, concluding with her core thesis that every enterprise must eventually own its software intelligence stack.0:57–3:46 · Guest disagreement 1/10 Lin Qiao's Career Journey and Founding Fireworks AI Harry sets up the interview citing background diligence with Eric Vishria and Dima, asking Lin about starting a company at 48. Lin shares her technical background in distributed systems and seven years at Meta learning people management.3:46–9:14 · Guest disagreement 3/10 Positioning Fireworks AI: Inference, Specialized Intelligence, and Enterprise Data Harry questions why inference and specialized intelligence are not commodities and pushes Anthropic's enterprise moat as a counterexample. Lin counters with Jensen Huang's insight that every business is specialized because proprietary data is private.9:14–14:57 · Guest disagreement 2/10 Open-Source Model Quality and the Viability of Frontier Labs Harry asks if frontier model companies are dramatically overvalued if open-source models solve 90% of enterprise workflows at 15x lower cost. Lin details how open models allow enterprises to retain weight control and hill-climb on private data.14:57–19:45 · Guest disagreement 3/10 AI Economics: Scaling to Bankruptcy, and One-Size-Fits-One Optimization Lin outlines the concept of startups and incumbents 'scaling into bankruptcy' with high inference COGS. Harry pushes back asking if OpenAI's aggressively price-reduced frontier models eliminate the open-weight advantage.19:45–23:04 · Guest disagreement 2/10 Geopolitical AI Risk, Guardrails, and Vertical Model Adoption Harry raises national security concerns regarding Chinese dominance on open LLM leaderboards and potential export restrictions. Lin argues that every company needs custom guardrails regardless of origin and open ecosystems naturally decentralize.23:04–27:14 · Guest disagreement 3/10 Disruption of SaaS Application Moats and Rapid Model Iteration Harry challenges the idea that SaaS moats have collapsed by citing long enterprise legal sales cycles at portfolio company Leya vs Harvey. Lin points out that workflow harnesses must be co-trained with underlying specialized models to prevent errors.27:14–30:23 · Guest disagreement 2/10 Base Model Step Functions, Government AI Infrastructure, and Monopoly Risks Harry asks if foundation AI should be treated as national infrastructure or partially government-owned. Lin warns against single-company monopolies, comparing base intelligence to utilities while defending specialized intelligence.30:23–33:10 · Guest disagreement 4/10 The AI Routing Layer and Autonomous Self-Evolving Systems Harry directly probes whether routing layers like OpenRouter or Requesty will be disintermediated if companies build autonomous routing internally. Lin concedes that standalone routers might be redundant once self-evolving automated routing matures.37:33–40:16 · Guest disagreement 2/10 Transitioning Beyond Coding to Enterprise Co-Work and Consumer AI Harry asks bluntly about customer concentration risk if Cursor churns. Lin details the diversification from coding to enterprise co-work and consumer recommendation engines across multiple verticals.40:16–43:22 · Guest disagreement 2/10 Token Volume Explosions and AI Supply Chain Bottlenecks Harry reacts incredulously to Lin predicting a 20-100x explosion in token count, noting that Capex bubble fears would be groundless under that demand. Lin explains that physical supply chain bottlenecks in energy and chips throttle delivery.43:22–47:19 · Guest disagreement 3/10 Startup Agility, Platform Specialization, and Enterprise AI Expenditure Harry presses on whether AI infrastructure players must go full-stack and why Nvidia's Jensen Huang does not expand into Fireworks' platform layer. Lin clarifies that Nvidia enters models purely to clear supply bottlenecks without building cloud inference.47:19–50:41 · Guest disagreement 3/10 Long-Term Token Economics and Infrastructure Cost Deflation Harry references Marc Benioff's developer spend data and asks whether token prices will truly fall when they have remained stubbornly high. Lin explains how task-level token efficiency and hardware supply easing will lead to a 10x cost reduction driving 100x usage.50:41–53:30 · Guest disagreement 3/10 Quality-First Inference, Zero KLD, and Bit-Wise Model Equivalence Harry asks if Fireworks is viewed as more expensive than low-cost providers like Together AI. Lin reframes the pricing dynamic, introducing the concept of zero KLD and exact bit-wise equivalence between training and inference.53:30–55:59 · Guest disagreement 2/10 Gross Margins in AI Infrastructure and Growth-Stage Priorities Harry asks if 30-40% gross margins represent the permanent new normal for AI infrastructure. Lin explains that low initial margins represent an intentional constraint choice to prioritize hypergrowth and geographic expansion over premature optimization.55:59–58:58 · Guest disagreement 2/10 Custom Data Center Architecture and Heterogeneous Hardware Clusters Lin outlines hardware disaggregation between FLOPs-intensive pre-fill and SRAM-intensive token generation using Groq ASICs. Harry connects this to his background diligence with Groq CEO Jonathan Ross.58:58–1:01:04 · Guest disagreement 2/10 Physical Data Center Execution Constraints and Global Competition Harry asks whether rapid physical infrastructure buildouts give China a structural advantage over US permitting hurdles. Lin agrees on civil engineering velocity but notes specialized operations talent is still scarce globally.1:01:04–1:04:57 · Guest disagreement 3/10 Hardware and Model Depreciation Cycles and Long-Term Margin Strategy Lin breaks down how hardware depreciation cycles have compressed from six years to annual multi-SKU releases, mismatching rapid model obsolescence. Harry synthesizes the implication that model speed outstrips hardware depreciation schedules.1:04:57–1:07:11 · Guest disagreement 2/10 Market Dynamics and Owning Intelligence vs. Renting Harry asks whether inference markets are winner-take-all like ridesharing or multi-player like hyperscale cloud. Lin argues that enterprises inevitably transition from renting general intelligence to owning and optimizing proprietary models.1:07:11–1:12:26 · Guest disagreement 3/10 Sovereign Models and National AI Independence Harry asks why startups avoid chip manufacturing if Meta, OpenAI, Anthropic, and DeepSeek build their own custom silicon. Lin educates him on silicon workload stabilization: taping out chips requires unchanging model architectures that premature startups cannot freeze.1:12:26–1:15:05 · Guest disagreement 2/10 System Level Infrastructure and Architecture Bottlenecks Harry cites Groq's Jonathan Ross claiming HBM memory bandwidth is the primary bottleneck. Lin counters that the broader architectural bottleneck is system-level co-design across 10-trillion parameter clusters, celebrating Fireworks hitting $800M run-rate potential.1:15:05–1:18:44 · Guest disagreement 1/10 Strategic Executive Hiring and AI Leadership Traits Harry inquires into the hiring of former Salesforce president George Hu. Lin explains how maintaining high curiosity alongside deep executive seniority allowed them to accelerate go-to-market without bureaucratic ossification.1:18:44–1:21:17 · Guest disagreement 1/10 Quick Fire: Scaling Velocity and Culture of Extreme Ownership In the quick fire round, Harry shares 20VC hiring profiles while Lin articulates the culture of extreme ownership and hiring talent that rejects artificial organizational boundaries.1:21:17–1:25:47 · Guest disagreement 1/10 Quick Fire: Leadership Lessons from Jensen Huang Lin reflects on leadership lessons learned from Jensen Huang, highlighting instant responsiveness and context-driven decision making. She shares her past mistake of underinvesting in marketing.1:25:47–1:28:07 · Guest disagreement 1/10 Quick Fire: Untapped Enterprise Segments Lin outlines untapped opportunities in traditional enterprise segments like insurance and banking, concluding with her core thesis that every enterprise must eventually own its software intelligence stack.0:57–3:46 · Harry pushing back 2/10 Lin Qiao's Career Journey and Founding Fireworks AI Harry sets up the interview citing background diligence with Eric Vishria and Dima, asking Lin about starting a company at 48. Lin shares her technical background in distributed systems and seven years at Meta learning people management.3:46–9:14 · Harry pushing back 5/10 Positioning Fireworks AI: Inference, Specialized Intelligence, and Enterprise Data Harry questions why inference and specialized intelligence are not commodities and pushes Anthropic's enterprise moat as a counterexample. Lin counters with Jensen Huang's insight that every business is specialized because proprietary data is private.9:14–14:57 · Harry pushing back 6/10 Open-Source Model Quality and the Viability of Frontier Labs Harry asks if frontier model companies are dramatically overvalued if open-source models solve 90% of enterprise workflows at 15x lower cost. Lin details how open models allow enterprises to retain weight control and hill-climb on private data.14:57–19:45 · Harry pushing back 5/10 AI Economics: Scaling to Bankruptcy, and One-Size-Fits-One Optimization Lin outlines the concept of startups and incumbents 'scaling into bankruptcy' with high inference COGS. Harry pushes back asking if OpenAI's aggressively price-reduced frontier models eliminate the open-weight advantage.19:45–23:04 · Harry pushing back 5/10 Geopolitical AI Risk, Guardrails, and Vertical Model Adoption Harry raises national security concerns regarding Chinese dominance on open LLM leaderboards and potential export restrictions. Lin argues that every company needs custom guardrails regardless of origin and open ecosystems naturally decentralize.23:04–27:14 · Harry pushing back 6/10 Disruption of SaaS Application Moats and Rapid Model Iteration Harry challenges the idea that SaaS moats have collapsed by citing long enterprise legal sales cycles at portfolio company Leya vs Harvey. Lin points out that workflow harnesses must be co-trained with underlying specialized models to prevent errors.27:14–30:23 · Harry pushing back 4/10 Base Model Step Functions, Government AI Infrastructure, and Monopoly Risks Harry asks if foundation AI should be treated as national infrastructure or partially government-owned. Lin warns against single-company monopolies, comparing base intelligence to utilities while defending specialized intelligence.30:23–33:10 · Harry pushing back 7/10 The AI Routing Layer and Autonomous Self-Evolving Systems Harry directly probes whether routing layers like OpenRouter or Requesty will be disintermediated if companies build autonomous routing internally. Lin concedes that standalone routers might be redundant once self-evolving automated routing matures.37:33–40:16 · Harry pushing back 5/10 Transitioning Beyond Coding to Enterprise Co-Work and Consumer AI Harry asks bluntly about customer concentration risk if Cursor churns. Lin details the diversification from coding to enterprise co-work and consumer recommendation engines across multiple verticals.40:16–43:22 · Harry pushing back 5/10 Token Volume Explosions and AI Supply Chain Bottlenecks Harry reacts incredulously to Lin predicting a 20-100x explosion in token count, noting that Capex bubble fears would be groundless under that demand. Lin explains that physical supply chain bottlenecks in energy and chips throttle delivery.43:22–47:19 · Harry pushing back 6/10 Startup Agility, Platform Specialization, and Enterprise AI Expenditure Harry presses on whether AI infrastructure players must go full-stack and why Nvidia's Jensen Huang does not expand into Fireworks' platform layer. Lin clarifies that Nvidia enters models purely to clear supply bottlenecks without building cloud inference.47:19–50:41 · Harry pushing back 6/10 Long-Term Token Economics and Infrastructure Cost Deflation Harry references Marc Benioff's developer spend data and asks whether token prices will truly fall when they have remained stubbornly high. Lin explains how task-level token efficiency and hardware supply easing will lead to a 10x cost reduction driving 100x usage.50:41–53:30 · Harry pushing back 5/10 Quality-First Inference, Zero KLD, and Bit-Wise Model Equivalence Harry asks if Fireworks is viewed as more expensive than low-cost providers like Together AI. Lin reframes the pricing dynamic, introducing the concept of zero KLD and exact bit-wise equivalence between training and inference.53:30–55:59 · Harry pushing back 5/10 Gross Margins in AI Infrastructure and Growth-Stage Priorities Harry asks if 30-40% gross margins represent the permanent new normal for AI infrastructure. Lin explains that low initial margins represent an intentional constraint choice to prioritize hypergrowth and geographic expansion over premature optimization.55:59–58:58 · Harry pushing back 4/10 Custom Data Center Architecture and Heterogeneous Hardware Clusters Lin outlines hardware disaggregation between FLOPs-intensive pre-fill and SRAM-intensive token generation using Groq ASICs. Harry connects this to his background diligence with Groq CEO Jonathan Ross.58:58–1:01:04 · Harry pushing back 5/10 Physical Data Center Execution Constraints and Global Competition Harry asks whether rapid physical infrastructure buildouts give China a structural advantage over US permitting hurdles. Lin agrees on civil engineering velocity but notes specialized operations talent is still scarce globally.1:01:04–1:04:57 · Harry pushing back 5/10 Hardware and Model Depreciation Cycles and Long-Term Margin Strategy Lin breaks down how hardware depreciation cycles have compressed from six years to annual multi-SKU releases, mismatching rapid model obsolescence. Harry synthesizes the implication that model speed outstrips hardware depreciation schedules.1:04:57–1:07:11 · Harry pushing back 5/10 Market Dynamics and Owning Intelligence vs. Renting Harry asks whether inference markets are winner-take-all like ridesharing or multi-player like hyperscale cloud. Lin argues that enterprises inevitably transition from renting general intelligence to owning and optimizing proprietary models.1:07:11–1:12:26 · Harry pushing back 6/10 Sovereign Models and National AI Independence Harry asks why startups avoid chip manufacturing if Meta, OpenAI, Anthropic, and DeepSeek build their own custom silicon. Lin educates him on silicon workload stabilization: taping out chips requires unchanging model architectures that premature startups cannot freeze.1:12:26–1:15:05 · Harry pushing back 5/10 System Level Infrastructure and Architecture Bottlenecks Harry cites Groq's Jonathan Ross claiming HBM memory bandwidth is the primary bottleneck. Lin counters that the broader architectural bottleneck is system-level co-design across 10-trillion parameter clusters, celebrating Fireworks hitting $800M run-rate potential.1:15:05–1:18:44 · Harry pushing back 3/10 Strategic Executive Hiring and AI Leadership Traits Harry inquires into the hiring of former Salesforce president George Hu. Lin explains how maintaining high curiosity alongside deep executive seniority allowed them to accelerate go-to-market without bureaucratic ossification.1:18:44–1:21:17 · Harry pushing back 2/10 Quick Fire: Scaling Velocity and Culture of Extreme Ownership In the quick fire round, Harry shares 20VC hiring profiles while Lin articulates the culture of extreme ownership and hiring talent that rejects artificial organizational boundaries.1:21:17–1:25:47 · Harry pushing back 2/10 Quick Fire: Leadership Lessons from Jensen Huang Lin reflects on leadership lessons learned from Jensen Huang, highlighting instant responsiveness and context-driven decision making. She shares her past mistake of underinvesting in marketing.1:25:47–1:28:07 · Harry pushing back 2/10 Quick Fire: Untapped Enterprise Segments Lin outlines untapped opportunities in traditional enterprise segments like insurance and banking, concluding with her core thesis that every enterprise must eventually own its software intelligence stack.

speaking balance: gold is Harry, purple is the guest (3 minute bins)

0:00 · Harry 36.8% · guest 63.2%0:00 · Harry 36.8% · guest 63.2%3:00 · Harry 24.4% · guest 75.6%3:00 · Harry 24.4% · guest 75.6%6:00 · Harry 0.7% · guest 99.3%6:00 · Harry 0.7% · guest 99.3%9:00 · Harry 27% · guest 73%9:00 · Harry 27% · guest 73%12:00 · Harry 13.2% · guest 86.8%12:00 · Harry 13.2% · guest 86.8%15:00 · Harry 10.1% · guest 89.9%15:00 · Harry 10.1% · guest 89.9%18:00 · Harry 12.2% · guest 87.8%18:00 · Harry 12.2% · guest 87.8%21:00 · Harry 37.1% · guest 62.9%21:00 · Harry 37.1% · guest 62.9%24:00 · Harry 17.4% · guest 82.6%24:00 · Harry 17.4% · guest 82.6%27:00 · Harry 21.7% · guest 78.3%27:00 · Harry 21.7% · guest 78.3%30:00 · Harry 25.8% · guest 74.2%30:00 · Harry 25.8% · guest 74.2%33:00 · Harry 5.6% · guest 94.4%33:00 · Harry 5.6% · guest 94.4%36:00 · Harry 10.3% · guest 89.7%36:00 · Harry 10.3% · guest 89.7%39:00 · Harry 18.5% · guest 81.5%39:00 · Harry 18.5% · guest 81.5%42:00 · Harry 13.3% · guest 86.7%42:00 · Harry 13.3% · guest 86.7%45:00 · Harry 33% · guest 67%45:00 · Harry 33% · guest 67%48:00 · Harry 14.8% · guest 85.2%48:00 · Harry 14.8% · guest 85.2%51:00 · Harry 16.2% · guest 83.8%51:00 · Harry 16.2% · guest 83.8%54:00 · Harry 13.6% · guest 86.4%54:00 · Harry 13.6% · guest 86.4%57:00 · Harry 19.4% · guest 80.6%57:00 · Harry 19.4% · guest 80.6%1:00:00 · Harry 9.9% · guest 90.1%1:00:00 · Harry 9.9% · guest 90.1%1:03:00 · Harry 21.2% · guest 78.8%1:03:00 · Harry 21.2% · guest 78.8%1:06:00 · Harry 21.7% · guest 78.3%1:06:00 · Harry 21.7% · guest 78.3%1:09:00 · Harry 9.3% · guest 90.7%1:09:00 · Harry 9.3% · guest 90.7%1:12:00 · Harry 30.4% · guest 69.6%1:12:00 · Harry 30.4% · guest 69.6%1:15:00 · Harry 13.8% · guest 86.2%1:15:00 · Harry 13.8% · guest 86.2%1:18:00 · Harry 20% · guest 80%1:18:00 · Harry 20% · guest 80%1:21:00 · Harry 5.2% · guest 94.8%1:21:00 · Harry 5.2% · guest 94.8%1:24:00 · Harry 10.5% · guest 89.5%1:24:00 · Harry 10.5% · guest 89.5%1:27:00 · Harry 26.5% · guest 73.5%1:27:00 · Harry 26.5% · guest 73.5%
Sharpest disagreement ▶ 32:27 Direct confrontation over routing layer defensibility

Lin resists Harry's suggestion that standalone routing is a permanently valuable independent layer, acknowledging automated self-evolving systems will replace manual routers.

Hardest push from Harry ▶ 24:16 Harry challenges SaaS moat collapse with legal enterprise reality

Harry pushes back firmly against Lin's thesis that application moats have dissolved, citing lengthy enterprise legal sales cycles and custom implementation barriers.

Biggest teaching moment ▶ 1:09:00 Lin breaks down ASIC and chip tape-out requirements

Lin educates Harry on hardware economics, demonstrating that custom chip design requires completely frozen, mature software workloads before committing to tape-out.

Harry holds his own ▶ 11:10 Harry calculates frontier lab valuation vs open-source reality

Harry synthesizes open-source efficiency gains, cost differentials, and workflow coverage to challenge the fundamental market valuation of frontier labs.

the scores for every segment, with the reasoning behind each
ChapterTopicHarry as informed peerGuest teachingGuest disagreementHarry pushing backWhy
Lin Qiao's Career Journey and Founding Fireworks AI 4212 Harry sets up the interview citing background diligence with Eric Vishria and Dima, asking Lin about starting a company at 48. Lin shares her technical background in distributed systems and seven years at Meta learning people management.
Positioning Fireworks AI: Inference, Specialized Intelligence, and Enterprise Data 6635 Harry questions why inference and specialized intelligence are not commodities and pushes Anthropic's enterprise moat as a counterexample. Lin counters with Jensen Huang's insight that every business is specialized because proprietary data is private.
Open-Source Model Quality and the Viability of Frontier Labs 7626 Harry asks if frontier model companies are dramatically overvalued if open-source models solve 90% of enterprise workflows at 15x lower cost. Lin details how open models allow enterprises to retain weight control and hill-climb on private data.
AI Economics: Scaling to Bankruptcy, and One-Size-Fits-One Optimization 6635 Lin outlines the concept of startups and incumbents 'scaling into bankruptcy' with high inference COGS. Harry pushes back asking if OpenAI's aggressively price-reduced frontier models eliminate the open-weight advantage.
Geopolitical AI Risk, Guardrails, and Vertical Model Adoption 6525 Harry raises national security concerns regarding Chinese dominance on open LLM leaderboards and potential export restrictions. Lin argues that every company needs custom guardrails regardless of origin and open ecosystems naturally decentralize.
Disruption of SaaS Application Moats and Rapid Model Iteration 6636 Harry challenges the idea that SaaS moats have collapsed by citing long enterprise legal sales cycles at portfolio company Leya vs Harvey. Lin points out that workflow harnesses must be co-trained with underlying specialized models to prevent errors.
Base Model Step Functions, Government AI Infrastructure, and Monopoly Risks 5424 Harry asks if foundation AI should be treated as national infrastructure or partially government-owned. Lin warns against single-company monopolies, comparing base intelligence to utilities while defending specialized intelligence.
The AI Routing Layer and Autonomous Self-Evolving Systems 7647 Harry directly probes whether routing layers like OpenRouter or Requesty will be disintermediated if companies build autonomous routing internally. Lin concedes that standalone routers might be redundant once self-evolving automated routing matures.
Transitioning Beyond Coding to Enterprise Co-Work and Consumer AI 6525 Harry asks bluntly about customer concentration risk if Cursor churns. Lin details the diversification from coding to enterprise co-work and consumer recommendation engines across multiple verticals.
Token Volume Explosions and AI Supply Chain Bottlenecks 6625 Harry reacts incredulously to Lin predicting a 20-100x explosion in token count, noting that Capex bubble fears would be groundless under that demand. Lin explains that physical supply chain bottlenecks in energy and chips throttle delivery.
Startup Agility, Platform Specialization, and Enterprise AI Expenditure 7536 Harry presses on whether AI infrastructure players must go full-stack and why Nvidia's Jensen Huang does not expand into Fireworks' platform layer. Lin clarifies that Nvidia enters models purely to clear supply bottlenecks without building cloud inference.
Long-Term Token Economics and Infrastructure Cost Deflation 6636 Harry references Marc Benioff's developer spend data and asks whether token prices will truly fall when they have remained stubbornly high. Lin explains how task-level token efficiency and hardware supply easing will lead to a 10x cost reduction driving 100x usage.
Quality-First Inference, Zero KLD, and Bit-Wise Model Equivalence 6735 Harry asks if Fireworks is viewed as more expensive than low-cost providers like Together AI. Lin reframes the pricing dynamic, introducing the concept of zero KLD and exact bit-wise equivalence between training and inference.
Gross Margins in AI Infrastructure and Growth-Stage Priorities 6625 Harry asks if 30-40% gross margins represent the permanent new normal for AI infrastructure. Lin explains that low initial margins represent an intentional constraint choice to prioritize hypergrowth and geographic expansion over premature optimization.
Custom Data Center Architecture and Heterogeneous Hardware Clusters 6724 Lin outlines hardware disaggregation between FLOPs-intensive pre-fill and SRAM-intensive token generation using Groq ASICs. Harry connects this to his background diligence with Groq CEO Jonathan Ross.
Physical Data Center Execution Constraints and Global Competition 6525 Harry asks whether rapid physical infrastructure buildouts give China a structural advantage over US permitting hurdles. Lin agrees on civil engineering velocity but notes specialized operations talent is still scarce globally.
Hardware and Model Depreciation Cycles and Long-Term Margin Strategy 6735 Lin breaks down how hardware depreciation cycles have compressed from six years to annual multi-SKU releases, mismatching rapid model obsolescence. Harry synthesizes the implication that model speed outstrips hardware depreciation schedules.
Market Dynamics and Owning Intelligence vs. Renting 5525 Harry asks whether inference markets are winner-take-all like ridesharing or multi-player like hyperscale cloud. Lin argues that enterprises inevitably transition from renting general intelligence to owning and optimizing proprietary models.
Sovereign Models and National AI Independence 6736 Harry asks why startups avoid chip manufacturing if Meta, OpenAI, Anthropic, and DeepSeek build their own custom silicon. Lin educates him on silicon workload stabilization: taping out chips requires unchanging model architectures that premature startups cannot freeze.
System Level Infrastructure and Architecture Bottlenecks 6625 Harry cites Groq's Jonathan Ross claiming HBM memory bandwidth is the primary bottleneck. Lin counters that the broader architectural bottleneck is system-level co-design across 10-trillion parameter clusters, celebrating Fireworks hitting $800M run-rate potential.
Strategic Executive Hiring and AI Leadership Traits 5413 Harry inquires into the hiring of former Salesforce president George Hu. Lin explains how maintaining high curiosity alongside deep executive seniority allowed them to accelerate go-to-market without bureaucratic ossification.
Quick Fire: Scaling Velocity and Culture of Extreme Ownership 4312 In the quick fire round, Harry shares 20VC hiring profiles while Lin articulates the culture of extreme ownership and hiring talent that rejects artificial organizational boundaries.
Quick Fire: Leadership Lessons from Jensen Huang 4412 Lin reflects on leadership lessons learned from Jensen Huang, highlighting instant responsiveness and context-driven decision making. She shares her past mistake of underinvesting in marketing.
Quick Fire: Untapped Enterprise Segments 4312 Lin outlines untapped opportunities in traditional enterprise segments like insurance and banking, concluding with her core thesis that every enterprise must eventually own its software intelligence stack.

Statements from this episode (43)

Opinion
Lin Qiao: A single-company monopoly on AI makes no sense
“What I don't want to see is there's only one company owns intelligence. I think that doesn't make sense to me.”
Lin Qiao Jul 20, 2026 ▶ 29:40
Opinion
Lin Qiao: Collaborative AI tools define the current market shift
“I think last year is the year of coding, and this year is the year of co-work.”
Lin Qiao Jul 20, 2026 ▶ 0:06
Disclosure
Stebbings wrote a $10M check for Fireworks after a 15-minute meeting
“This is 20 VC with me, Harry Stebbings, and in the hot seat today, a founder who I wrote a ten million dollar check for after just a 15 minute meeting, Lin Kuo, founder of Fireworks.”
Harry Stebbings Jul 20, 2026 ▶ 0:38
Prediction Open · timeframe Jul 2029
Lin Qiao: AI token costs will drop 10x in three years, driving 100x usage
“I do think the cost of token will go down drastically. 10 X cost reduction in the next three years, and this 10 X cost reduction will drive a hundred X usage.”
Lin Qiao Jul 20, 2026 ▶ 0:23
Prediction Open · timeframe Jul 2029
Lin Qiao: Fireworks AI will not enter the application layer
“We absolutely are not going to move into application layer. Very clear to us.”
Lin Qiao Jul 20, 2026 ▶ 0:34
Assertion Supported
Fireworks hired former Salesforce president George Hu
“Lin just hired George Hu. He was the former president of Salesforce.”
Harry Stebbings Jul 20, 2026 ▶ 1:01
Prediction Not checkable as stated
Stebbings: Fireworks can become a $500 billion company
“If so, fireworks, honestly, It can be a five hundred billion dollar company.”
Harry Stebbings Jul 20, 2026 ▶ 1:08
Assertion Supported
Lin Qiao: Most world data is private enterprise data, not public internet
“If you think intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model. The training data is coming from public internet and the label data. Public internet is very small. Corpus of data comp…”
Lin Qiao Jul 20, 2026 ▶ 4:31
Prediction Not checkable as stated
Lin Qiao: Enterprise data will never be shared with third parties
“It will never get shared with anyone else because this is company's proprietary IP.”
Lin Qiao Jul 20, 2026 ▶ 5:02
Prediction Not checkable as stated
Lin Qiao: The future of AI is private, specialized intelligence
“We believe the future of the frontier of the intelligence, actually private intelligence, our specialized intelligence.”
Lin Qiao Jul 20, 2026 ▶ 5:22
Opinion
Qiao: Anthropic relies on a single-AGI belief over specialization
“I view Anthropic as a company fully believing AGI.”
Lin Qiao Jul 20, 2026 ▶ 5:57
Insight
Lin Qiao: Corporate Uniqueness and User Data Intelligence Cannot Be Replicated Externally
“Every single company is doing something unique that justify their existence, and this something unique is deeply baked into their product design, is deeply baked into their software design and system building, and that's deeply baked into the data and the inte…”
Lin Qiao Jul 20, 2026 ▶ 8:58
Disclosure
Fireworks AI uses open-source models internally for hiring, finance, and coding
“So within, ah, within Firewalls, obviously, we dock for our own product. We use OpenModel to drive our recruiting process, candidate sourcing, and the feedback collection. We use OpenModel to even drive some internal finance processes. Obviously, for coding, w…”
Lin Qiao Jul 20, 2026 ▶ 13:08
Insight
Lin Qiao: Fine-tuned open models outperform general models on specialized enterprise tasks
“And the model intelligence has password threshold is much easier to steer, especially with small amount of data. A small amount of unique data a particular company has, and then we can hill climb towards your eval, and often time, the end result of hill climbi…”
Lin Qiao Jul 20, 2026 ▶ 14:05
Insight
Lin Qiao: Product-market fit no longer guarantees business durability in AI SaaS
“During SaaS time, product market fit and the durable business almost are equivalent to each other. The hardest thing is find product market fit, and then once you find it just scales as fast as you can. Because CPU is a commodity. The infrastructure you build …”
Lin Qiao Jul 20, 2026 ▶ 15:08
Assertion Not checkable as stated
Lin Qiao: Fireworks has achieved 5x to 10x inference cost reductions
“What we have seen in the past is five times to 10 times cost reduction.”
Lin Qiao Jul 20, 2026 ▶ 19:19
Prediction Not checkable as stated
Lin Qiao: AI future will feature millions of specialized models, not AGI
“I really believe the future will not be a few small number of AGI models dominating the world. I really believe the future will be, it may be scary, but I think that's true, it will be millions of specialized models, one per application, per use case.”
Lin Qiao Jul 20, 2026 ▶ 20:59
Prediction Not checkable as stated
Lin Qiao: US can build its own open-source AI ecosystem if China restricts access
“I do believe, in terms of talent density and resources, I do believe U.S. Will be able to build that open system by ourselves, and we should.”
Lin Qiao Jul 20, 2026 ▶ 21:55
Insight
Lin Qiao: AI coding models have eliminated software implementation as a competitive barrier
“In the past, it requires tens of very strong product engineers and PMs to convert from idea to implementation to production scale. Multiple quarters of years of investment. That's a deep mode. And today, one person, a few weeks, can possibly launch their ideas…”
Lin Qiao Jul 20, 2026 ▶ 23:33
Assertion Supported
Lin Qiao: Almost all AI coding startups now fine-tune custom models
“In coding space, Cursor probably is one of the pioneers starting to tune their model, and now almost all coding companies tune their own models.”
Lin Qiao Jul 20, 2026 ▶ 26:38
Prediction Not checkable as stated
Lin Qiao: Major AI base IQ leaps will occur every 9-12 months
“I see those as every year or every three quarters, there's a major leap.”
Lin Qiao Jul 20, 2026 ▶ 27:56
Insight
Qiao: Enterprise AI frontier is routing architecture, not single models
“So again, my thinking of what is the frontier is not just this one model. The frontier could be your special routing mechanism for your business, and you decompose that based on, hey, in order to fulfill this task, and you need a highly intelligent layer, mayb…”
Lin Qiao Jul 20, 2026 ▶ 30:48
Prediction Not checkable as stated
Qiao: AI model routing will evolve into fully automated self-learning systems
“We also think there's a space to build a automatic routing system that can learn by itself, and that compound with automatic tuning system eventually. We think it should all be automated, and then you can see a self-evolving system based on what flow through y…”
Lin Qiao Jul 20, 2026 ▶ 31:41
Opinion
Qiao: Standalone model router services like OpenRouter won't be needed long-term
“You probably don't. We're not there yet, but I do think this is, this could be area of innovation.”
Lin Qiao Jul 20, 2026 ▶ 32:38
Assertion Not checkable as stated
Lin Qiao: Fireworks runs distributed RL across 5-6 global data center regions
“We've designed a fully distributed system. We run across five, six data center regions globally, and tap into scattered GPUs, and they are able to run massive jobs, our jobs.”
Lin Qiao Jul 20, 2026 ▶ 36:27
Assertion Not checkable as stated
Lin Qiao: All major AI coding companies use Fireworks infrastructure
“I think all major coding companies are on us.”
Lin Qiao Jul 20, 2026 ▶ 38:40
Assertion Not checkable as stated
Lin Qiao: Fireworks processes over 40 trillion tokens daily
“Today we process more than 40 trillion tokens a day. So majority of those tokens are coming from a customized model, not from off-the-shelf models, are coming from a customized model.”
Lin Qiao Jul 20, 2026 ▶ 40:59
Prediction Not checkable as stated
Qiao: Fireworks daily token count could grow 20x-100x by next year
“Anywhere ranging from 20 to a hundred X could be possible.”
Lin Qiao Jul 20, 2026 ▶ 41:22
Disclosure
Qiao: Cursor was a single-digit million-dollar company before 1000x growth
“I would say Curso is the first company they have decided to work with us. Early on. I still remember when they worked with us, they were single digit million dollar. Very small. That's only two years ago. They grew by a hundred, a thousand X over two years.”
Lin Qiao Jul 20, 2026 ▶ 44:29
Prediction Not checkable as stated
Lin Qiao: token costs will fall drastically and cheaper infrastructure will invite far more usage
“I do think the cost of token will go down drastically. High price will invite a lot of people coming in to solve the problem, and it will invite competition. Competition will bring down the cost, and then eventually will lead into a very economical solution, r…”
Lin Qiao Jul 20, 2026 ▶ 47:19
Assertion Partly supported
Fireworks achieves exact bit equivalence between AI training and inference
“Between the training system and the inference system, when models move over, We have bit equivalence. So, as in, the numerics are fully the same. We do not lose a bit of accuracy.”
Lin Qiao Jul 20, 2026 ▶ 52:06
Assertion Not checkable as stated
Qiao: Fireworks could optimize gross margins now but chooses rapid expansion
“If our focus is only optimized growth margin, we absolutely can do that, but we are sacrificing the speed of growth because we want to go everywhere.”
Lin Qiao Jul 20, 2026 ▶ 55:24
Assertion Supported
Lin Qiao: AI hardware vendors now release three SKUs per year
“And now within a year from one vendor alone, we have three SKUs.”
Lin Qiao Jul 20, 2026 ▶ 1:01:24
Assertion Not checkable as stated
Lin Qiao: Legal AI market consolidated from many startups to two
“Take legal, for example. I was on a dinner table, and interesting, it seems like there were a lot of those companies around two years ago, but now it's pretty much two.”
Lin Qiao Jul 20, 2026 ▶ 1:05:17
Prediction Not checkable as stated
Lin Qiao: Nations will build sovereign AI models like power grids
“I definitely see that possibility. I also see, if we think about the general intelligence model as the electricity layer, as a power line, every country should own their own power line, right?”
Lin Qiao Jul 20, 2026 ▶ 1:07:46
Assertion Supported
Meta Has Been Building Custom Silicon Chips for Over Five Years
“I think Meta has been building their chips for more than five years, way more than five years.”
Lin Qiao Jul 20, 2026 ▶ 1:09:20
Insight
Qiao: Generative AI workloads are too dynamic to justify custom silicon
“We're still in the early stage of workload maturity for now to warrant a chip that will be durable.”
Lin Qiao Jul 20, 2026 ▶ 1:11:22
Assertion Not checkable as stated
The AI Industry Lacks Systems Designed for 10-Trillion Parameter Models
“Great system designed for 10 trillion parameter models today.”
Lin Qiao Jul 20, 2026 ▶ 1:13:15
Assertion Supported
Qiao: Fireworks AI currently has 200 employees
“So today we're at 200 people.”
Lin Qiao Jul 20, 2026 ▶ 1:15:35
Opinion
Stebbings: 20VC hires mostly immigrants because British people lack strong work ethic
“Pretty much only hire immigrants. British people don't work very hard.”
Harry Stebbings Jul 20, 2026 ▶ 1:19:57
Prediction Not checkable as stated
Qiao: AI market will shift from token maxing to ROI maxing
“In the next couple of years, as AI is getting more and more into production, there will be a lot of focus. In getting that clarity and getting that discipline out. The token maxing is just a thing in time, but we're quickly moving to ROI maxing, which is about…”
Lin Qiao Jul 20, 2026 ▶ 1:25:23
Disclosure
Fireworks AI counts Geico and Capital One as enterprise customers
“We have customers like Geico, like Capital One, all these companies.”
Lin Qiao Jul 20, 2026 ▶ 1:26:16
Prediction Not checkable as stated
Qiao: Every company will own its custom AI intelligence within three years
“I really see people will own their, every single company will own their own intelligence as a must-have. It's not optional.”
Lin Qiao Jul 20, 2026 ▶ 1:27:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.