Jul 17, 2026 · 59m · sourcery

Inference 101: SambaNova CEO Rodrigo Liang · Sourcery with Molly O'Shea

Rodrigo Liang · 45m spoken Molly O'Shea · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Sourcery recorded at the RAISE Summit in Paris, host Molly O'Shea interviews SambaNova Systems CEO Rodrigo Liang following the company's $1 billion fundraising round at an $11 billion valuation. Liang details SambaNova's hardware architecture optimized for AI inference, distributed data center economics, agentic latency requirements, and enterprise AI sovereignty.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Molly holds 17.6% of the talking time here. How this is scored →

Molly as informed peer 4.3 Guest teaching 4.6 Guest disagreement 1.1 Molly pushing back 1.6
05100:0015:0030:0045:000:59–7:18 · Molly as informed peer 4/10 SambaNova's $1 Billion Fundraising Announcement Molly prompts Rodrigo on SambaNova's latest $1B funding announcement and how the company evolved from training to inference. Rodrigo gives a detailed technical breakdown comparing SN-40's 10kW air-cooled rack footprint against 140kW NVIDIA GPU clusters.7:18–11:00 · Molly as informed peer 3/10 Rack Composition and Data Center Deployment Efficiency Molly asks how rack composition has shifted between training and inference paradigms. Rodrigo educates on the differences in clustering requirements, noting how training requires complex synchronization and high-cost networking whereas inference can scale out cleanly in single-rack increments.11:00–14:17 · Molly as informed peer 4/10 Distributed Data Centers and Latency in Agentic AI Molly challenges the narrative around building $50B to $100B gigawatt data centers. Rodrigo nuances the premise, explaining that while mega-clusters exist, agentic workflows require distributed, ultra-low-latency metropolitan data centers where multi-agent compounding latencies cannot tolerate distant compute.14:17–17:29 · Molly as informed peer 4/10 Defining Premium Inference: Accuracy and Speed Molly asks Rodrigo to define premium inference in practical terms. Rodrigo explains that premium inference is defined by two vectors: model size for accuracy (running full precision without quantizing) and token generation speed.17:29–20:43 · Molly as informed peer 4/10 Sponsor Segment: Brex Following a sponsor read, Molly asks whether market pricing will bifurcate between consumer and enterprise speed tiers. Rodrigo draws an analogy to telecommunications (5G vs 2G), arguing that speed becomes baseline expectation as delivery costs decline.20:43–24:27 · Molly as informed peer 6/10 Edge Computing and Modular Data Centers Molly demonstrates industry expertise by referencing Starlink and naming edge modular data center startup Armada. Rodrigo validates her reference, confirming SambaNova partners directly with Armada to deploy 10kW racks inside rugged shipping containers for remote industrial edge use cases.24:27–27:48 · Molly as informed peer 4/10 Coopetition and Data Center Economics Molly asks how SambaNova navigates coopetition when co-existing alongside established chipmakers like NVIDIA in shared data centers. Rodrigo breaks down the unit economics, showing how offloading inference traffic to SambaNova lifts operator margins and frees up GPU racks for training and HPC.27:48–31:02 · Molly as informed peer 4/10 Model Routing and Revenue Per Rack Metrics Molly inquires about the mechanics of API-level model routing and hardware metric tracking. Rodrigo explains that modern API standards simplify routing across heterogeneous chips and defines the core business KPI as monthly revenue generated per token per rack.31:02–35:12 · Molly as informed peer 6/10 Industry Bottlenecks and the Global AI Land Grab Molly contextualizes industry scaling bottlenecks by citing Dylan Field on AI taxonomy and Theresa Carlson on AWS government cloud adoption. Rodrigo frames the market as a global land grab where pure NVIDIA compute is commoditized and custom inference provides necessary differentiation.35:12–37:53 · Molly as informed peer 4/10 Capital Efficiency and Heterogeneous Infrastructure Molly questions the sustainability of massive fundraising rounds across rival chip and infrastructure startups. Rodrigo defends high valuations for durable winners but argues data centers will only support three or four chip architectures due to operational overhead.37:53–40:19 · Molly as informed peer 3/10 Cloud Strategy and NeoCloud Partnerships Molly asks about the necessity of building an internal first-party cloud product. Rodrigo outlines SambaNova's strategic choice to sell racks and partner with NeoClouds like Vista Equity's Vector Core Compute rather than competing directly with hyperscalers.40:19–43:48 · Molly as informed peer 5/10 Sponsor Segment: AssemblyAI After reading sponsor spots, Molly references Alex Karp's remarks on sovereign data infrastructure to ask whether Europe must build indigenous chips. Rodrigo explains that sovereign computing is driven primarily by protecting proprietary corporate and national data from global foundation model ingestion.43:48–47:55 · Molly as informed peer 6/10 On-Premises Repatriation and Proprietary Enterprise Models Molly highlights the ongoing on-prem repatriation trend, referencing Harvey AI's defensive posture against frontier labs. Rodrigo emphasizes that enterprises relying entirely on public frontier models risk margin erosion and must train proprietary models on their own private data.47:55–51:15 · Molly as informed peer 5/10 Economic AI Spending and Business Value Creation Molly notes that enterprise guidance is shifting from indiscriminate AI spending towards capital discipline. Rodrigo agrees, contrasting exploratory 'AI-first' experimentation with structured business transformation aimed at expanding service offerings.51:15–56:06 · Molly as informed peer 4/10 Unasked Questions: Service Differentiation and Scaling Molly asks what fundamental questions the tech sector is overlooking today. Rodrigo warns that companies must focus on true service differentiation and scaling infrastructure early, invoking Netflix's disruption of Blockbuster and Circuit City as a cautionary tale.56:06–58:00 · Molly as informed peer 3/10 Career Reflections: Resilience and Team Building Molly asks Rodrigo to reflect on mentorship and career resilience across his 32 years in semiconductor engineering. Rodrigo shares lessons on perseverance and presents SambaNova's latest SN-50 silicon chip before closing the show.0:59–7:18 · Guest teaching 5/10 SambaNova's $1 Billion Fundraising Announcement Molly prompts Rodrigo on SambaNova's latest $1B funding announcement and how the company evolved from training to inference. Rodrigo gives a detailed technical breakdown comparing SN-40's 10kW air-cooled rack footprint against 140kW NVIDIA GPU clusters.7:18–11:00 · Guest teaching 6/10 Rack Composition and Data Center Deployment Efficiency Molly asks how rack composition has shifted between training and inference paradigms. Rodrigo educates on the differences in clustering requirements, noting how training requires complex synchronization and high-cost networking whereas inference can scale out cleanly in single-rack increments.11:00–14:17 · Guest teaching 6/10 Distributed Data Centers and Latency in Agentic AI Molly challenges the narrative around building $50B to $100B gigawatt data centers. Rodrigo nuances the premise, explaining that while mega-clusters exist, agentic workflows require distributed, ultra-low-latency metropolitan data centers where multi-agent compounding latencies cannot tolerate distant compute.14:17–17:29 · Guest teaching 5/10 Defining Premium Inference: Accuracy and Speed Molly asks Rodrigo to define premium inference in practical terms. Rodrigo explains that premium inference is defined by two vectors: model size for accuracy (running full precision without quantizing) and token generation speed.17:29–20:43 · Guest teaching 4/10 Sponsor Segment: Brex Following a sponsor read, Molly asks whether market pricing will bifurcate between consumer and enterprise speed tiers. Rodrigo draws an analogy to telecommunications (5G vs 2G), arguing that speed becomes baseline expectation as delivery costs decline.20:43–24:27 · Guest teaching 3/10 Edge Computing and Modular Data Centers Molly demonstrates industry expertise by referencing Starlink and naming edge modular data center startup Armada. Rodrigo validates her reference, confirming SambaNova partners directly with Armada to deploy 10kW racks inside rugged shipping containers for remote industrial edge use cases.24:27–27:48 · Guest teaching 5/10 Coopetition and Data Center Economics Molly asks how SambaNova navigates coopetition when co-existing alongside established chipmakers like NVIDIA in shared data centers. Rodrigo breaks down the unit economics, showing how offloading inference traffic to SambaNova lifts operator margins and frees up GPU racks for training and HPC.27:48–31:02 · Guest teaching 5/10 Model Routing and Revenue Per Rack Metrics Molly inquires about the mechanics of API-level model routing and hardware metric tracking. Rodrigo explains that modern API standards simplify routing across heterogeneous chips and defines the core business KPI as monthly revenue generated per token per rack.31:02–35:12 · Guest teaching 4/10 Industry Bottlenecks and the Global AI Land Grab Molly contextualizes industry scaling bottlenecks by citing Dylan Field on AI taxonomy and Theresa Carlson on AWS government cloud adoption. Rodrigo frames the market as a global land grab where pure NVIDIA compute is commoditized and custom inference provides necessary differentiation.35:12–37:53 · Guest teaching 5/10 Capital Efficiency and Heterogeneous Infrastructure Molly questions the sustainability of massive fundraising rounds across rival chip and infrastructure startups. Rodrigo defends high valuations for durable winners but argues data centers will only support three or four chip architectures due to operational overhead.37:53–40:19 · Guest teaching 5/10 Cloud Strategy and NeoCloud Partnerships Molly asks about the necessity of building an internal first-party cloud product. Rodrigo outlines SambaNova's strategic choice to sell racks and partner with NeoClouds like Vista Equity's Vector Core Compute rather than competing directly with hyperscalers.40:19–43:48 · Guest teaching 5/10 Sponsor Segment: AssemblyAI After reading sponsor spots, Molly references Alex Karp's remarks on sovereign data infrastructure to ask whether Europe must build indigenous chips. Rodrigo explains that sovereign computing is driven primarily by protecting proprietary corporate and national data from global foundation model ingestion.43:48–47:55 · Guest teaching 4/10 On-Premises Repatriation and Proprietary Enterprise Models Molly highlights the ongoing on-prem repatriation trend, referencing Harvey AI's defensive posture against frontier labs. Rodrigo emphasizes that enterprises relying entirely on public frontier models risk margin erosion and must train proprietary models on their own private data.47:55–51:15 · Guest teaching 4/10 Economic AI Spending and Business Value Creation Molly notes that enterprise guidance is shifting from indiscriminate AI spending towards capital discipline. Rodrigo agrees, contrasting exploratory 'AI-first' experimentation with structured business transformation aimed at expanding service offerings.51:15–56:06 · Guest teaching 5/10 Unasked Questions: Service Differentiation and Scaling Molly asks what fundamental questions the tech sector is overlooking today. Rodrigo warns that companies must focus on true service differentiation and scaling infrastructure early, invoking Netflix's disruption of Blockbuster and Circuit City as a cautionary tale.56:06–58:00 · Guest teaching 3/10 Career Reflections: Resilience and Team Building Molly asks Rodrigo to reflect on mentorship and career resilience across his 32 years in semiconductor engineering. Rodrigo shares lessons on perseverance and presents SambaNova's latest SN-50 silicon chip before closing the show.0:59–7:18 · Guest disagreement 1/10 SambaNova's $1 Billion Fundraising Announcement Molly prompts Rodrigo on SambaNova's latest $1B funding announcement and how the company evolved from training to inference. Rodrigo gives a detailed technical breakdown comparing SN-40's 10kW air-cooled rack footprint against 140kW NVIDIA GPU clusters.7:18–11:00 · Guest disagreement 1/10 Rack Composition and Data Center Deployment Efficiency Molly asks how rack composition has shifted between training and inference paradigms. Rodrigo educates on the differences in clustering requirements, noting how training requires complex synchronization and high-cost networking whereas inference can scale out cleanly in single-rack increments.11:00–14:17 · Guest disagreement 2/10 Distributed Data Centers and Latency in Agentic AI Molly challenges the narrative around building $50B to $100B gigawatt data centers. Rodrigo nuances the premise, explaining that while mega-clusters exist, agentic workflows require distributed, ultra-low-latency metropolitan data centers where multi-agent compounding latencies cannot tolerate distant compute.14:17–17:29 · Guest disagreement 1/10 Defining Premium Inference: Accuracy and Speed Molly asks Rodrigo to define premium inference in practical terms. Rodrigo explains that premium inference is defined by two vectors: model size for accuracy (running full precision without quantizing) and token generation speed.17:29–20:43 · Guest disagreement 1/10 Sponsor Segment: Brex Following a sponsor read, Molly asks whether market pricing will bifurcate between consumer and enterprise speed tiers. Rodrigo draws an analogy to telecommunications (5G vs 2G), arguing that speed becomes baseline expectation as delivery costs decline.20:43–24:27 · Guest disagreement 1/10 Edge Computing and Modular Data Centers Molly demonstrates industry expertise by referencing Starlink and naming edge modular data center startup Armada. Rodrigo validates her reference, confirming SambaNova partners directly with Armada to deploy 10kW racks inside rugged shipping containers for remote industrial edge use cases.24:27–27:48 · Guest disagreement 1/10 Coopetition and Data Center Economics Molly asks how SambaNova navigates coopetition when co-existing alongside established chipmakers like NVIDIA in shared data centers. Rodrigo breaks down the unit economics, showing how offloading inference traffic to SambaNova lifts operator margins and frees up GPU racks for training and HPC.27:48–31:02 · Guest disagreement 1/10 Model Routing and Revenue Per Rack Metrics Molly inquires about the mechanics of API-level model routing and hardware metric tracking. Rodrigo explains that modern API standards simplify routing across heterogeneous chips and defines the core business KPI as monthly revenue generated per token per rack.31:02–35:12 · Guest disagreement 2/10 Industry Bottlenecks and the Global AI Land Grab Molly contextualizes industry scaling bottlenecks by citing Dylan Field on AI taxonomy and Theresa Carlson on AWS government cloud adoption. Rodrigo frames the market as a global land grab where pure NVIDIA compute is commoditized and custom inference provides necessary differentiation.35:12–37:53 · Guest disagreement 2/10 Capital Efficiency and Heterogeneous Infrastructure Molly questions the sustainability of massive fundraising rounds across rival chip and infrastructure startups. Rodrigo defends high valuations for durable winners but argues data centers will only support three or four chip architectures due to operational overhead.37:53–40:19 · Guest disagreement 1/10 Cloud Strategy and NeoCloud Partnerships Molly asks about the necessity of building an internal first-party cloud product. Rodrigo outlines SambaNova's strategic choice to sell racks and partner with NeoClouds like Vista Equity's Vector Core Compute rather than competing directly with hyperscalers.40:19–43:48 · Guest disagreement 1/10 Sponsor Segment: AssemblyAI After reading sponsor spots, Molly references Alex Karp's remarks on sovereign data infrastructure to ask whether Europe must build indigenous chips. Rodrigo explains that sovereign computing is driven primarily by protecting proprietary corporate and national data from global foundation model ingestion.43:48–47:55 · Guest disagreement 1/10 On-Premises Repatriation and Proprietary Enterprise Models Molly highlights the ongoing on-prem repatriation trend, referencing Harvey AI's defensive posture against frontier labs. Rodrigo emphasizes that enterprises relying entirely on public frontier models risk margin erosion and must train proprietary models on their own private data.47:55–51:15 · Guest disagreement 1/10 Economic AI Spending and Business Value Creation Molly notes that enterprise guidance is shifting from indiscriminate AI spending towards capital discipline. Rodrigo agrees, contrasting exploratory 'AI-first' experimentation with structured business transformation aimed at expanding service offerings.51:15–56:06 · Guest disagreement 1/10 Unasked Questions: Service Differentiation and Scaling Molly asks what fundamental questions the tech sector is overlooking today. Rodrigo warns that companies must focus on true service differentiation and scaling infrastructure early, invoking Netflix's disruption of Blockbuster and Circuit City as a cautionary tale.56:06–58:00 · Guest disagreement 0/10 Career Reflections: Resilience and Team Building Molly asks Rodrigo to reflect on mentorship and career resilience across his 32 years in semiconductor engineering. Rodrigo shares lessons on perseverance and presents SambaNova's latest SN-50 silicon chip before closing the show.0:59–7:18 · Molly pushing back 1/10 SambaNova's $1 Billion Fundraising Announcement Molly prompts Rodrigo on SambaNova's latest $1B funding announcement and how the company evolved from training to inference. Rodrigo gives a detailed technical breakdown comparing SN-40's 10kW air-cooled rack footprint against 140kW NVIDIA GPU clusters.7:18–11:00 · Molly pushing back 1/10 Rack Composition and Data Center Deployment Efficiency Molly asks how rack composition has shifted between training and inference paradigms. Rodrigo educates on the differences in clustering requirements, noting how training requires complex synchronization and high-cost networking whereas inference can scale out cleanly in single-rack increments.11:00–14:17 · Molly pushing back 3/10 Distributed Data Centers and Latency in Agentic AI Molly challenges the narrative around building $50B to $100B gigawatt data centers. Rodrigo nuances the premise, explaining that while mega-clusters exist, agentic workflows require distributed, ultra-low-latency metropolitan data centers where multi-agent compounding latencies cannot tolerate distant compute.14:17–17:29 · Molly pushing back 1/10 Defining Premium Inference: Accuracy and Speed Molly asks Rodrigo to define premium inference in practical terms. Rodrigo explains that premium inference is defined by two vectors: model size for accuracy (running full precision without quantizing) and token generation speed.17:29–20:43 · Molly pushing back 2/10 Sponsor Segment: Brex Following a sponsor read, Molly asks whether market pricing will bifurcate between consumer and enterprise speed tiers. Rodrigo draws an analogy to telecommunications (5G vs 2G), arguing that speed becomes baseline expectation as delivery costs decline.20:43–24:27 · Molly pushing back 1/10 Edge Computing and Modular Data Centers Molly demonstrates industry expertise by referencing Starlink and naming edge modular data center startup Armada. Rodrigo validates her reference, confirming SambaNova partners directly with Armada to deploy 10kW racks inside rugged shipping containers for remote industrial edge use cases.24:27–27:48 · Molly pushing back 2/10 Coopetition and Data Center Economics Molly asks how SambaNova navigates coopetition when co-existing alongside established chipmakers like NVIDIA in shared data centers. Rodrigo breaks down the unit economics, showing how offloading inference traffic to SambaNova lifts operator margins and frees up GPU racks for training and HPC.27:48–31:02 · Molly pushing back 2/10 Model Routing and Revenue Per Rack Metrics Molly inquires about the mechanics of API-level model routing and hardware metric tracking. Rodrigo explains that modern API standards simplify routing across heterogeneous chips and defines the core business KPI as monthly revenue generated per token per rack.31:02–35:12 · Molly pushing back 2/10 Industry Bottlenecks and the Global AI Land Grab Molly contextualizes industry scaling bottlenecks by citing Dylan Field on AI taxonomy and Theresa Carlson on AWS government cloud adoption. Rodrigo frames the market as a global land grab where pure NVIDIA compute is commoditized and custom inference provides necessary differentiation.35:12–37:53 · Molly pushing back 3/10 Capital Efficiency and Heterogeneous Infrastructure Molly questions the sustainability of massive fundraising rounds across rival chip and infrastructure startups. Rodrigo defends high valuations for durable winners but argues data centers will only support three or four chip architectures due to operational overhead.37:53–40:19 · Molly pushing back 1/10 Cloud Strategy and NeoCloud Partnerships Molly asks about the necessity of building an internal first-party cloud product. Rodrigo outlines SambaNova's strategic choice to sell racks and partner with NeoClouds like Vista Equity's Vector Core Compute rather than competing directly with hyperscalers.40:19–43:48 · Molly pushing back 2/10 Sponsor Segment: AssemblyAI After reading sponsor spots, Molly references Alex Karp's remarks on sovereign data infrastructure to ask whether Europe must build indigenous chips. Rodrigo explains that sovereign computing is driven primarily by protecting proprietary corporate and national data from global foundation model ingestion.43:48–47:55 · Molly pushing back 2/10 On-Premises Repatriation and Proprietary Enterprise Models Molly highlights the ongoing on-prem repatriation trend, referencing Harvey AI's defensive posture against frontier labs. Rodrigo emphasizes that enterprises relying entirely on public frontier models risk margin erosion and must train proprietary models on their own private data.47:55–51:15 · Molly pushing back 2/10 Economic AI Spending and Business Value Creation Molly notes that enterprise guidance is shifting from indiscriminate AI spending towards capital discipline. Rodrigo agrees, contrasting exploratory 'AI-first' experimentation with structured business transformation aimed at expanding service offerings.51:15–56:06 · Molly pushing back 1/10 Unasked Questions: Service Differentiation and Scaling Molly asks what fundamental questions the tech sector is overlooking today. Rodrigo warns that companies must focus on true service differentiation and scaling infrastructure early, invoking Netflix's disruption of Blockbuster and Circuit City as a cautionary tale.56:06–58:00 · Molly pushing back 0/10 Career Reflections: Resilience and Team Building Molly asks Rodrigo to reflect on mentorship and career resilience across his 32 years in semiconductor engineering. Rodrigo shares lessons on perseverance and presents SambaNova's latest SN-50 silicon chip before closing the show.

speaking balance: gold is Molly, purple is the guest (3 minute bins)

0:00 · Molly 20.5% · guest 79.5%0:00 · Molly 20.5% · guest 79.5%3:00 · Molly 10.4% · guest 89.6%3:00 · Molly 10.4% · guest 89.6%6:00 · Molly 1.4% · guest 98.6%6:00 · Molly 1.4% · guest 98.6%9:00 · Molly 3.9% · guest 96.1%9:00 · Molly 3.9% · guest 96.1%12:00 · Molly 5.7% · guest 94.3%12:00 · Molly 5.7% · guest 94.3%15:00 · Molly 16.9% · guest 83.1%15:00 · Molly 16.9% · guest 83.1%18:00 · Molly 34.4% · guest 65.6%18:00 · Molly 34.4% · guest 65.6%21:00 · Molly 23.1% · guest 76.9%21:00 · Molly 23.1% · guest 76.9%24:00 · Molly 7.9% · guest 92.1%24:00 · Molly 7.9% · guest 92.1%27:00 · Molly 5.2% · guest 94.8%27:00 · Molly 5.2% · guest 94.8%30:00 · Molly 45% · guest 55%30:00 · Molly 45% · guest 55%33:00 · Molly 5% · guest 95%33:00 · Molly 5% · guest 95%36:00 · Molly 1.6% · guest 98.4%36:00 · Molly 1.6% · guest 98.4%39:00 · Molly 61.1% · guest 38.9%39:00 · Molly 61.1% · guest 38.9%42:00 · Molly 32.5% · guest 67.5%42:00 · Molly 32.5% · guest 67.5%45:00 · Molly 2.5% · guest 97.5%45:00 · Molly 2.5% · guest 97.5%48:00 · Molly 15.3% · guest 84.7%48:00 · Molly 15.3% · guest 84.7%51:00 · Molly 6% · guest 94%51:00 · Molly 6% · guest 94%54:00 · Molly 18.1% · guest 81.9%54:00 · Molly 18.1% · guest 81.9%57:00 · Molly 37.1% · guest 62.9%57:00 · Molly 37.1% · guest 62.9%
Sharpest disagreement ▶ 33:45 Rodrigo labeling NVIDIA hardware as a commodity

Rodrigo forcefully dismisses conventional reverence for incumbent hardware, asserting that NVIDIA GPUs are fundamentally commoditized assets offering negligible margin differentiation.

Hardest push from Molly ▶ 11:00 Molly challenging the necessity of mega data centers

Molly questions the viability of pervasive mega-scale infrastructure, directly asking whether headlines hyping $50B to $100B data centers reflect realistic market demands.

Biggest teaching moment ▶ 11:50 Rodrigo detailing compounding multi-agent latency constraints

Rodrigo educates Molly on why distributed metro deployments are mathematically mandatory for agentic AI, walking through how 20 interacting sub-agents compound single-model response delays into unacceptable 40-second user latencies.

Molly holds their own ▶ 22:50 Molly connecting modular edge compute directly to Armada

Molly demonstrates deep domain insight by pinpointing edge infrastructure startup Armada and industrial mission-critical use cases before Rodrigo confirms their existing technical partnership.

the scores for every segment, with the reasoning behind each
ChapterTopicMolly as informed peerGuest teachingGuest disagreementMolly pushing backWhy
SambaNova's $1 Billion Fundraising Announcement 4511 Molly prompts Rodrigo on SambaNova's latest $1B funding announcement and how the company evolved from training to inference. Rodrigo gives a detailed technical breakdown comparing SN-40's 10kW air-cooled rack footprint against 140kW NVIDIA GPU clusters.
Rack Composition and Data Center Deployment Efficiency 3611 Molly asks how rack composition has shifted between training and inference paradigms. Rodrigo educates on the differences in clustering requirements, noting how training requires complex synchronization and high-cost networking whereas inference can scale out cleanly in single-rack increments.
Distributed Data Centers and Latency in Agentic AI 4623 Molly challenges the narrative around building $50B to $100B gigawatt data centers. Rodrigo nuances the premise, explaining that while mega-clusters exist, agentic workflows require distributed, ultra-low-latency metropolitan data centers where multi-agent compounding latencies cannot tolerate distant compute.
Defining Premium Inference: Accuracy and Speed 4511 Molly asks Rodrigo to define premium inference in practical terms. Rodrigo explains that premium inference is defined by two vectors: model size for accuracy (running full precision without quantizing) and token generation speed.
Sponsor Segment: Brex 4412 Following a sponsor read, Molly asks whether market pricing will bifurcate between consumer and enterprise speed tiers. Rodrigo draws an analogy to telecommunications (5G vs 2G), arguing that speed becomes baseline expectation as delivery costs decline.
Edge Computing and Modular Data Centers 6311 Molly demonstrates industry expertise by referencing Starlink and naming edge modular data center startup Armada. Rodrigo validates her reference, confirming SambaNova partners directly with Armada to deploy 10kW racks inside rugged shipping containers for remote industrial edge use cases.
Coopetition and Data Center Economics 4512 Molly asks how SambaNova navigates coopetition when co-existing alongside established chipmakers like NVIDIA in shared data centers. Rodrigo breaks down the unit economics, showing how offloading inference traffic to SambaNova lifts operator margins and frees up GPU racks for training and HPC.
Model Routing and Revenue Per Rack Metrics 4512 Molly inquires about the mechanics of API-level model routing and hardware metric tracking. Rodrigo explains that modern API standards simplify routing across heterogeneous chips and defines the core business KPI as monthly revenue generated per token per rack.
Industry Bottlenecks and the Global AI Land Grab 6422 Molly contextualizes industry scaling bottlenecks by citing Dylan Field on AI taxonomy and Theresa Carlson on AWS government cloud adoption. Rodrigo frames the market as a global land grab where pure NVIDIA compute is commoditized and custom inference provides necessary differentiation.
Capital Efficiency and Heterogeneous Infrastructure 4523 Molly questions the sustainability of massive fundraising rounds across rival chip and infrastructure startups. Rodrigo defends high valuations for durable winners but argues data centers will only support three or four chip architectures due to operational overhead.
Cloud Strategy and NeoCloud Partnerships 3511 Molly asks about the necessity of building an internal first-party cloud product. Rodrigo outlines SambaNova's strategic choice to sell racks and partner with NeoClouds like Vista Equity's Vector Core Compute rather than competing directly with hyperscalers.
Sponsor Segment: AssemblyAI 5512 After reading sponsor spots, Molly references Alex Karp's remarks on sovereign data infrastructure to ask whether Europe must build indigenous chips. Rodrigo explains that sovereign computing is driven primarily by protecting proprietary corporate and national data from global foundation model ingestion.
On-Premises Repatriation and Proprietary Enterprise Models 6412 Molly highlights the ongoing on-prem repatriation trend, referencing Harvey AI's defensive posture against frontier labs. Rodrigo emphasizes that enterprises relying entirely on public frontier models risk margin erosion and must train proprietary models on their own private data.
Economic AI Spending and Business Value Creation 5412 Molly notes that enterprise guidance is shifting from indiscriminate AI spending towards capital discipline. Rodrigo agrees, contrasting exploratory 'AI-first' experimentation with structured business transformation aimed at expanding service offerings.
Unasked Questions: Service Differentiation and Scaling 4511 Molly asks what fundamental questions the tech sector is overlooking today. Rodrigo warns that companies must focus on true service differentiation and scaling infrastructure early, invoking Netflix's disruption of Blockbuster and Circuit City as a cautionary tale.
Career Reflections: Resilience and Team Building 3300 Molly asks Rodrigo to reflect on mentorship and career resilience across his 32 years in semiconductor engineering. Rodrigo shares lessons on perseverance and presents SambaNova's latest SN-50 silicon chip before closing the show.

Statements from this episode (25)

Disclosure
Rodrigo Liang: SambaNova closes $1B round at $11B valuation
“We just did the first close of a billion dollar fund raise at an eleven billion valuation.”
Rodrigo Liang Jul 17, 2026 ▶ 1:34
Prediction Open · timeframe Jul 2031
Liang: AI inference chip deployments will dwarf training by orders of magnitude
“Because at scale, the number of chips deployed for inferencing will be orders of magnitude greater than whatever you're doing for training.”
Rodrigo Liang Jul 17, 2026 ▶ 4:37
Disclosure
Liang: SambaNova taped out six chips in seven years, tape-out seven next year
“We've taped out six chips in the last seven years, and we'll tape out seven for the next year.”
Rodrigo Liang Jul 17, 2026 ▶ 4:54
Assertion Contradicted
Liang: SambaNova 10kW SN40 rack outperformed 140kW Nvidia GPU racks
“And so by the time we released SN-Forty a couple years ago, it became incredibly popular, because instead of a 130, a 140 kilowatt rack of NVIDIA GPU, we were outperforming them with a 10 kilowatt SN-Forty rack.”
Rodrigo Liang Jul 17, 2026 ▶ 5:52
Assertion Supported
Liang: SambaNova serves 1.5T parameter models on one rack versus 10-20
“And so with sum it over, that minimum quantum is down to one rack. Right, where if you have other, other service providers, you just run, say, a DeepSeq model, which is now one and a half trillion parameters, just to run that, the minimum for some of the other…”
Rodrigo Liang Jul 17, 2026 ▶ 8:22
Assertion Supported
Liang: SambaNova inference uses standard air-cooled racks and Ethernet
“Standard 19 inch rack, standard air cooling, no complicated liquid cooling retrofit in the data center. We're using standard Kubernetes, standard Red Hat Linux, standard Ethernet at the top for networking. We don't have to use all this kind of really expensive…”
Rodrigo Liang Jul 17, 2026 ▶ 9:31
Prediction Not checkable as stated
Liang: Agentic AI will drive a wave of mid-sized distributed data centers
“And I think you're going to see this new wave of companies that are doing distributed data centers, right? So these data centers are mid-sized, right? They're mid-sized, and it's going to be even more important as you go into this agentic world”
Rodrigo Liang Jul 17, 2026 ▶ 11:27
Disclosure
Liang: SambaNova is deploying hardware in metropolitan cities for lower latency
“And so you're now seeing us coming in and saying, look, we're going to deploy the hardware where the users are in large metropolitan cities. Because that's where business is being run, and so latency is really important, so we're going to deploy that.”
Rodrigo Liang Jul 17, 2026 ▶ 13:15
Prediction Not checkable as stated
Liang: Frontier AI models are heading toward 10 trillion parameters
“The new models are heading towards 10 trillion. Even the open source models are already one to two trillion parameter models, and so now you're starting to see these models getting very big because people are looking for accuracy, right?”
Rodrigo Liang Jul 17, 2026 ▶ 16:12
Opinion
Liang: Groq and Cerebras can only run small models fast
“Or you look at services like Rock Cerebus that run fast, you can only run the small models, right?”
Rodrigo Liang Jul 17, 2026 ▶ 17:04
Disclosure
Liang: SambaNova runs the largest models in full precision without quantizing
“We take the biggest models in the world and we run them in the original precision. We don't, you know, we don't quantize. We don't, you know, quantizing is, you know, you chop half the, you know, weights off. So, so we don't chop the model down. We just run or…”
Rodrigo Liang Jul 17, 2026 ▶ 17:12
Prediction Not checkable as stated
Liang: Demand will converge on the fastest, most accurate large models
“And so, as the cost of delivering fast goes down, you're going to see most people switch over to the fastest. And this is why I feel like, you know, the premium inference, which is large models, which equals the most accurate. The most accurate models and fast…”
Rodrigo Liang Jul 17, 2026 ▶ 20:13
Disclosure
Liang: SambaNova builds 10-to-20 rack edge data centers inside shipping containers
“Because of somewhat of technology being as little as 10 kilowatts per rack, we can put them inside shipping containers. Right, so we build out these data centers, clusters of 10, 20 racks, inside the shipping containers that you see on these ships, right?”
Rodrigo Liang Jul 17, 2026 ▶ 22:04
Opinion
Liang: AI inference service providers lack sustainable profit margins today
“Today, inference services, they're not making enough margins. You're generating lots of revenue, but you're not generating enough margin, and in order for them to sustain, they gotta be more profitable”
Rodrigo Liang Jul 17, 2026 ▶ 27:18
Insight
Liang: AI compute economics are measured by token revenue generated per rack
“And this is the way we see most providers measuring. They purchase per rack, they operate per rack, and so they want to generate revenue per rack, and the revenue is generated per token, right? If I put a rack of hardware, I'm just seeing how many tokens I mig…”
Rodrigo Liang Jul 17, 2026 ▶ 30:04
Opinion
Liang: Nvidia GPUs are a commodity offering very limited cost advantage
“People forget, as much as NVIDIA costs, it's commodity. Right? Because what you offer is the same as what your neighbor offers and your differentiation is, I can save you a little bit of money because maybe I got a discount from NVIDIA, right? Or maybe I got a…”
Rodrigo Liang Jul 17, 2026 ▶ 33:47
Disclosure
Liang: SambaNova has raised $2.5 billion in total capital
“We're two and a half billion dollars raised in the history of the company.”
Rodrigo Liang Jul 17, 2026 ▶ 35:22
Prediction Not checkable as stated
Liang predicts data centers will support only three to four AI chips
“As much as these data centers and service providers are heterogeneous, right, they're using different chips, NVIDIA and other chips, it's not going to be a hundred different chips. It might be two or three. Maybe three or four, right? That's as heterogeneous a…”
Rodrigo Liang Jul 17, 2026 ▶ 36:08
Disclosure
Liang: SambaNova will not build proprietary cloud to compete with AWS
“Many of the chip companies have chosen to go build their own cloud. They compete with the AWSs of the world. We have chosen not to do that. What we decided that, you know, we want to do is focus our energy on creating technology that we can ship.”
Rodrigo Liang Jul 17, 2026 ▶ 38:09
Assertion Supported
Liang: Japan and South Korea are investing heavily in sovereign AI models
“Countries have started doing this work and you see this in Japan and Korea announced the same thing you see in other parts of the world where they're investing a significant, significant amount of money to actually train from scratch. Train their own national …”
Rodrigo Liang Jul 17, 2026 ▶ 42:29
Assertion Not checkable as stated
Liang: Infrastructure repatriation to on-premises is actively occurring across enterprises
“Look, repatriation of infrastructure into on-prem is definitely happening, right? You saw this big shift. Everybody's got a cloud called cloud, you know, everything's in the cloud, and 20 years later, you still have companies just starting the migration to the…”
Rodrigo Liang Jul 17, 2026 ▶ 44:48
Prediction Not checkable as stated
Liang: Relying on commodity AI models will compress enterprise margins within two years
“If you actually transfer all of those You know, services, that differentiation, to all using the same exact model. That's in the community. Where does the differentiation come from? Right? And so what most companies start to realize, if you just fast forward, …”
Rodrigo Liang Jul 17, 2026 ▶ 45:56
Insight
Liang: Demanding upfront AI ROI is like calculating the ROI of email
“It's kind of like early days of internet. If we say, what is it, what is the ROI for email? Well, show me, you know, I mean, there are companies that say, prove me ROI before you use it. You know, am I typing an email? Should I try to figure out what the ROI o…”
Rodrigo Liang Jul 17, 2026 ▶ 49:34
Prediction Not checkable as stated
Liang: AI energy, data center, and chip constraints will only worsen
“The hints that you're seeing today with energy constraints, data center constraint, chip availability constraints, cost constraints, all of those things are only getting exacerbated, right?”
Rodrigo Liang Jul 17, 2026 ▶ 54:03
Opinion
Liang: Disaggregated inference across SambaNova, Nvidia, and Xeon is most efficient
“This aggregated inference that we talked about with NVIDIA chips, with Summono we've already used, with Xeons, Is just the most efficient way of actually deploying inference at scale”
Rodrigo Liang Jul 17, 2026 ▶ 58:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 160 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.