Jul 16, 2026 · 1h 41m · latent-space

🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences

Andy Beam · 48m spoken Rafa Gomez-Bombarelli · 27m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space Science, Lila Sciences leaders Andy Beam and Rafa Gomez-Bombarelli discuss building a frontier scientific reasoning platform where automated physical laboratories act as empirical verifiers for reinforcement learning across biology and materials science.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.7 Guest teaching 2.9 Guest disagreement 1.2 The hosts pushing back 3.2
05100:0020:0040:001:00:001:20:001:40:003:08–5:37 · The hosts as informed peer 3/10 Rafa Gomez-Bombarelli's Journey in Computational Materials Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila.5:38–10:02 · The hosts as informed peer 7/10 Lila's Core Thesis: Nature as a Verifiable Reward Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads.10:02–14:20 · The hosts as informed peer 6/10 Lab Architecture and Modular Automation Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line.14:20–18:30 · The hosts as informed peer 6/10 Laboratory Safety and Non-Intuitive Discoveries Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails.18:30–24:41 · The hosts as informed peer 6/10 Scientific Rigor, Human Synergy, and Green Hydrogen RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts.24:41–28:20 · The hosts as informed peer 5/10 RL Pathologies, Chain of Thought, and Tool Calling Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls.28:21–33:13 · The hosts as informed peer 7/10 Lila's Model-First Strategy and Cross-Domain Generalization Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models.33:13–35:23 · The hosts as informed peer 6/10 Scientific Representation: Language vs Domain Architectures RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools.35:24–44:13 · The hosts as informed peer 6/10 Materials Science Capabilities and Rapid Formulations Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents.44:14–52:59 · The hosts as informed peer 5/10 In Vivo CAR-T Breakthrough and Virtual Startups Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform.52:59–59:43 · The hosts as informed peer 7/10 Translational Bottlenecks and Scientific Serendipity RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities.59:44–1:07:10 · The hosts as informed peer 4/10 Machine Creativity and Open-Ended Exploration Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware.1:07:10–1:14:05 · The hosts as informed peer 6/10 Lab Orchestration, Scheduling, and Fast Feedback Loops Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods.1:14:07–1:18:39 · The hosts as informed peer 5/10 Rapid Instrument Onboarding and Facility Expansion RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge.1:18:39–1:24:27 · The hosts as informed peer 7/10 The 10-Trillion Scientific Token Dataset Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models.1:24:27–1:31:41 · The hosts as informed peer 6/10 Flagship Pioneering Origins and Distinct Identity Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment.1:31:42–1:36:13 · The hosts as informed peer 6/10 Materials Science vs Biology: Challenges and Economics Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors.1:36:13–1:40:13 · The hosts as informed peer 5/10 Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute.3:08–5:37 · Guest teaching 1/10 Rafa Gomez-Bombarelli's Journey in Computational Materials Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila.5:38–10:02 · Guest teaching 3/10 Lila's Core Thesis: Nature as a Verifiable Reward Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads.10:02–14:20 · Guest teaching 2/10 Lab Architecture and Modular Automation Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line.14:20–18:30 · Guest teaching 4/10 Laboratory Safety and Non-Intuitive Discoveries Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails.18:30–24:41 · Guest teaching 3/10 Scientific Rigor, Human Synergy, and Green Hydrogen RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts.24:41–28:20 · Guest teaching 2/10 RL Pathologies, Chain of Thought, and Tool Calling Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls.28:21–33:13 · Guest teaching 4/10 Lila's Model-First Strategy and Cross-Domain Generalization Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models.33:13–35:23 · Guest teaching 5/10 Scientific Representation: Language vs Domain Architectures RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools.35:24–44:13 · Guest teaching 3/10 Materials Science Capabilities and Rapid Formulations Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents.44:14–52:59 · Guest teaching 2/10 In Vivo CAR-T Breakthrough and Virtual Startups Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform.52:59–59:43 · Guest teaching 3/10 Translational Bottlenecks and Scientific Serendipity RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities.59:44–1:07:10 · Guest teaching 2/10 Machine Creativity and Open-Ended Exploration Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware.1:07:10–1:14:05 · Guest teaching 3/10 Lab Orchestration, Scheduling, and Fast Feedback Loops Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods.1:14:07–1:18:39 · Guest teaching 2/10 Rapid Instrument Onboarding and Facility Expansion RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge.1:18:39–1:24:27 · Guest teaching 3/10 The 10-Trillion Scientific Token Dataset Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models.1:24:27–1:31:41 · Guest teaching 2/10 Flagship Pioneering Origins and Distinct Identity Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment.1:31:42–1:36:13 · Guest teaching 6/10 Materials Science vs Biology: Challenges and Economics Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors.1:36:13–1:40:13 · Guest teaching 3/10 Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute.3:08–5:37 · Guest disagreement 0/10 Rafa Gomez-Bombarelli's Journey in Computational Materials Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila.5:38–10:02 · Guest disagreement 1/10 Lila's Core Thesis: Nature as a Verifiable Reward Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads.10:02–14:20 · Guest disagreement 1/10 Lab Architecture and Modular Automation Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line.14:20–18:30 · Guest disagreement 2/10 Laboratory Safety and Non-Intuitive Discoveries Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails.18:30–24:41 · Guest disagreement 1/10 Scientific Rigor, Human Synergy, and Green Hydrogen RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts.24:41–28:20 · Guest disagreement 1/10 RL Pathologies, Chain of Thought, and Tool Calling Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls.28:21–33:13 · Guest disagreement 2/10 Lila's Model-First Strategy and Cross-Domain Generalization Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models.33:13–35:23 · Guest disagreement 4/10 Scientific Representation: Language vs Domain Architectures RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools.35:24–44:13 · Guest disagreement 1/10 Materials Science Capabilities and Rapid Formulations Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents.44:14–52:59 · Guest disagreement 0/10 In Vivo CAR-T Breakthrough and Virtual Startups Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform.52:59–59:43 · Guest disagreement 1/10 Translational Bottlenecks and Scientific Serendipity RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities.59:44–1:07:10 · Guest disagreement 0/10 Machine Creativity and Open-Ended Exploration Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware.1:07:10–1:14:05 · Guest disagreement 1/10 Lab Orchestration, Scheduling, and Fast Feedback Loops Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods.1:14:07–1:18:39 · Guest disagreement 0/10 Rapid Instrument Onboarding and Facility Expansion RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge.1:18:39–1:24:27 · Guest disagreement 1/10 The 10-Trillion Scientific Token Dataset Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models.1:24:27–1:31:41 · Guest disagreement 1/10 Flagship Pioneering Origins and Distinct Identity Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment.1:31:42–1:36:13 · Guest disagreement 4/10 Materials Science vs Biology: Challenges and Economics Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors.1:36:13–1:40:13 · Guest disagreement 1/10 Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute.3:08–5:37 · The hosts pushing back 0/10 Rafa Gomez-Bombarelli's Journey in Computational Materials Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila.5:38–10:02 · The hosts pushing back 4/10 Lila's Core Thesis: Nature as a Verifiable Reward Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads.10:02–14:20 · The hosts pushing back 3/10 Lab Architecture and Modular Automation Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line.14:20–18:30 · The hosts pushing back 6/10 Laboratory Safety and Non-Intuitive Discoveries Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails.18:30–24:41 · The hosts pushing back 3/10 Scientific Rigor, Human Synergy, and Green Hydrogen RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts.24:41–28:20 · The hosts pushing back 3/10 RL Pathologies, Chain of Thought, and Tool Calling Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls.28:21–33:13 · The hosts pushing back 5/10 Lila's Model-First Strategy and Cross-Domain Generalization Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models.33:13–35:23 · The hosts pushing back 3/10 Scientific Representation: Language vs Domain Architectures RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools.35:24–44:13 · The hosts pushing back 4/10 Materials Science Capabilities and Rapid Formulations Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents.44:14–52:59 · The hosts pushing back 1/10 In Vivo CAR-T Breakthrough and Virtual Startups Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform.52:59–59:43 · The hosts pushing back 5/10 Translational Bottlenecks and Scientific Serendipity RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities.59:44–1:07:10 · The hosts pushing back 1/10 Machine Creativity and Open-Ended Exploration Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware.1:07:10–1:14:05 · The hosts pushing back 4/10 Lab Orchestration, Scheduling, and Fast Feedback Loops Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods.1:14:07–1:18:39 · The hosts pushing back 2/10 Rapid Instrument Onboarding and Facility Expansion RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge.1:18:39–1:24:27 · The hosts pushing back 4/10 The 10-Trillion Scientific Token Dataset Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models.1:24:27–1:31:41 · The hosts pushing back 3/10 Flagship Pioneering Origins and Distinct Identity Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment.1:31:42–1:36:13 · The hosts pushing back 4/10 Materials Science vs Biology: Challenges and Economics Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors.1:36:13–1:40:13 · The hosts pushing back 2/10 Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 34:00 Rafa rejects language as the sole medium for chemistry

Rafa firmly pushes back on the host's assertion that the future of chemistry is language, arguing that chemists do not think in English and domain-specific geometric representations remain essential.

Hardest push from the hosts ▶ 15:20 Brandon challenges the relevance of AI safety in Lila's lab

Brandon directly challenges the premise that AI safety is a practical concern in Lila's current lab automation setting, arguing that malicious risk or catastrophic emergent outputs are negligible.

Biggest teaching moment ▶ 1:33:45 Rafa educates on materials economics versus biotech assets

Rafa dismantles the host's analogy between material product validation and clinical trials, demonstrating how approved drugs command massive pricing power while critical materials like superconductors remain anonymous and low-margin.

The host holds their own ▶ 29:45 Brandon cites Sri Kosuri on ML business model paradoxes

Brandon challenges Lila's core data-generation premise by citing Sri Kosuri's critique on whether models are even needed once requisite domain datasets have already been generated.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Rafa Gomez-Bombarelli's Journey in Computational Materials 3100 Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila.
Lila's Core Thesis: Nature as a Verifiable Reward 7314 Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads.
Lab Architecture and Modular Automation 6213 Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line.
Laboratory Safety and Non-Intuitive Discoveries 6426 Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails.
Scientific Rigor, Human Synergy, and Green Hydrogen 6313 RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts.
RL Pathologies, Chain of Thought, and Tool Calling 5213 Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls.
Lila's Model-First Strategy and Cross-Domain Generalization 7425 Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models.
Scientific Representation: Language vs Domain Architectures 6543 RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools.
Materials Science Capabilities and Rapid Formulations 6314 Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents.
In Vivo CAR-T Breakthrough and Virtual Startups 5201 Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform.
Translational Bottlenecks and Scientific Serendipity 7315 RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities.
Machine Creativity and Open-Ended Exploration 4201 Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware.
Lab Orchestration, Scheduling, and Fast Feedback Loops 6314 Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods.
Rapid Instrument Onboarding and Facility Expansion 5202 RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge.
The 10-Trillion Scientific Token Dataset 7314 Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models.
Flagship Pioneering Origins and Distinct Identity 6213 Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment.
Materials Science vs Biology: Challenges and Economics 6644 Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors.
Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization 5312 Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute.

Statements from this episode (37)

Opinion
Beam: Academia lacks the scaled compute needed for frontier AI
“Academia has a lot going for it. Access to scaled compute is not one of the things that it has going for it, or scaled resources.”
Andy Beam Jul 16, 2026 ▶ 2:33
Insight
Beam: Science acts as an infinite token generator for AI
“Science is as an infinite token generator to train models at scale.”
Andy Beam Jul 16, 2026 ▶ 2:46
Opinion
Gomez-Bombarelli: AI Must Solve Real Material Science Beyond Simulation
“But it was clear that we needed to bridge a gap and get this thing all the way out and make, do AI for actual material science and not just the computational version.”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 5:10
Opinion
Beam: AI research has exhausted human-generated internet data
“That data came from the internet, it was human generated, and we have used it all. You know, as Elia said at NeurIPS last year, We have but one internet. It's the fossil fuel. We fracked. We got every ounce of data that we could out of the internet, but it's g…”
Andy Beam Jul 16, 2026 ▶ 6:20
Opinion
Beam: Nature and scientific experiments are ultimate verifiers for RL
“But what at Lilo we believe is that actually science running the scientific method and using nature and experiments as verifier is like the ultimate version of that. And so what we're building, we'll talk about these things that we call AI science factories. T…”
Andy Beam Jul 16, 2026 ▶ 6:59
Disclosure
Beam: Lila Sciences prioritizes experimental flexibility over throughput
“So actually, the experimental platform that we're building prioritizes generalizability and flexibility over raw throughput. We want the model to be able to design a new experimental protocol, run the protocol, and receive the feedback, even if that's not an e…”
Andy Beam Jul 16, 2026 ▶ 9:34
Disclosure
Beam: Lila uses magnetically levitating planar motors to connect lab instruments
“We have a physical transport layer that connects almost every instrument that we have bought at Lila to each other. These are currently planar motor systems where there's an I-A-Six well plate, That magnetically levitates over a track. You have sort of millime…”
Andy Beam Jul 16, 2026 ▶ 10:40
Assertion Not checkable as stated
Beam: Lila's AI hits 80% zero-shot on gene editing, beating humans' 0%
“Certainly for expression protocols, for some gene editing work that we've done we have tested like the platform's ability to do that versus humans. Model gets like 80% of that zero shot. Humans get zero percent of that zero shot.”
Andy Beam Jul 16, 2026 ▶ 13:47
Insight
Beam: AI safety cannot wait because capability curves are sigmoid-shaped
“Safety is not something you can procrastinate on because capability curves tend to be sigmoid shaped and it can look like everything's fine and then all of a sudden there's something that you didn't Anticipate the model being able to do.”
Andy Beam Jul 16, 2026 ▶ 16:35
Assertion Not checkable as stated
Beam: Lila's best non-platinum electrocatalysts came from AI ideas experts called stupid
“Some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turns…”
Andy Beam Jul 16, 2026 ▶ 17:53
Insight
Gomez-Bombarelli: False Positives Are Terrible for Humans but Great for AI
“False positives are terrible for human scientists, right? Because you go to try something. It doesn't work for the model. It's fantastic. It reduces uncertainty a lot. For the operator, it's kind of a bummer, right? Because you thought you were going to get so…”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 20:34
Disclosure
Lila Sciences structures wet lab instruments as model tool calls in English
“At Lila, the fun thing is that the lab instruments are also tool calls or a series of tool calls for a workflow or but it's all human legible. It's all in English.”
Andy Beam Jul 16, 2026 ▶ 27:12
Insight
Beam: Chain of thought is an unreliable narrator of model computation
“It actually thinks in latent space, it emits tokens. So, like, the chain of thought is often an unreliable narrator for what the model, the computation of the model is actually doing.”
Andy Beam Jul 16, 2026 ▶ 27:51
Disclosure
Beam: Lila's core asset is its reasoning model, not clinical pipeline
“The model itself is the thing of value at Lila. So in that sense, we're much more of like a neolab trying to think of a new way to push forward capabilities of a core reasoning LLM based model.”
Andy Beam Jul 16, 2026 ▶ 28:48
Insight
Beam: Cross-domain training reduces domain data requirements in science models
“And so again, the core bet that we're making is that is true for science. That if the model is trained on an increasingly broad set of data, the amount of data that you need in a given domain, that data requirement is reduced. In some cases will be reduced to …”
Andy Beam Jul 16, 2026 ▶ 30:20
Assertion Not checkable as stated
Beam: Lila's 10-trillion-token general science model beats specialized AI
“So we have assembled this reasoning data set of 10 trillion scientific tokens reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences, and we have seen that this general model often beats the domain-specific mod…”
Andy Beam Jul 16, 2026 ▶ 32:39
Opinion
Gomez-Bombarelli: Pure language representation is not required for scientific superintelligence
“I don't think it's necessary. I don't think that's a necessary condition for a scientific super intelligence.”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 33:21
Assertion Not checkable as stated
Gomez-Bombarelli: Self-driving lab synthesizes custom quantum dots in 90 minutes
“So we have a cute demo where our self-driving lab, we ask our visitors to pick up wavelengths. What color do you want the quantum dot to be when they come into the office? And then we fire off the machine, The model reasons, even sometimes we've been throwing …”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 36:27
Assertion Not checkable as stated
Gomez-Bombarelli: Small-Molecule Drug Discovery Training Transferred to Metal-Organic Frameworks
“And it turns out that our models had been trained on small molecule drug discovery, and all of the chemistry that they had learned thinking about drug discovery carried over to start reasoning over this metal organic framework materials that we can use to take…”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 39:53
Assertion Not checkable as stated
Beam: Lila's in vivo CAR-T data outperformed Capstan in non-human primates
“So we have developed some monster UTRs, untranslated regions, which flank the protein coding region which dictate those expression properties. Something like Tenex, the references from Moderna and Pfizer. And over the course of six months, got to in vivo data …”
Andy Beam Jul 16, 2026 ▶ 48:48
Opinion
Beam: U.S. Biotech Lagging China Is Regulatory, Not an Innovation Deficit
“The reason why U.S. Biotech is losing to Chinese biotech is not because of an innovation problem. There's a regulatory framework too that has to go to enabling like fast clinical trials.”
Andy Beam Jul 16, 2026 ▶ 54:37
Disclosure
Gomez-Bombarelli: Lila AI Calls Process Simulators for Scale-Up Economics
“In the materials and chemistry, you know, our tools can call process engineering simulators. And go figure out what pipe diameters and what heat exchangers you should be using in order to scale up the process for the economics to be worth it.”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 56:28
What-if
Beam: CAR-T Therapy Likely Would Have Failed Without Serendipitous Intervention
“In so many counterfactual worlds, that doctor was not the one treating Emily Whitehead in that case, and CAR-T may have looked like, it may have been yet another gravestone in E-Room's law, you know, for you know, yet another failed drug. So I do think that li…”
Andy Beam Jul 16, 2026 ▶ 59:04
Disclosure
Beam: Lila builds custom firmware and drivers for lab instruments
“Because we have written our own custom drivers, our own custom firmware to get sort of low level granular control over a lot of these instruments and make them talk to each other.”
Andy Beam Jul 16, 2026 ▶ 1:04:30
Disclosure
Beam: Lila uses a vision-language model to automate Windows 95 machines
“We actually have a vision language model controlling a Windows 95 machine. Because that's the only way to automate it.”
Andy Beam Jul 16, 2026 ▶ 1:04:57
Opinion
Beam: The lab of the future should look like a data center
“But we think that like the lab of the future should not be made for people to easily walk into it. It should feel like a data center where you go and you see the rows of server racks. There's room for like a crash cart behind it to service the nodes. But it sh…”
Andy Beam Jul 16, 2026 ▶ 1:05:56
Disclosure
Beam: Future AI labs will be million-square-foot, lights-out data centers
“We do think about scaling it In the same way that you would think about scaling a data center, in that it's a multi-level building, occupies millions of square feet, and it's like a lights-out facility, as they say. It's like running 24 seven, generating data …”
Andy Beam Jul 16, 2026 ▶ 1:09:00
Assertion Not checkable as stated
Gomez-Bombarelli: Lila tests 96 MOFs in an hour using proxy measurements
“And that's something we built in the lab now where instead of measuring pressure, we're measuring another property we care about that is a readout for what actually pressure will tell us, but we can do 96 well plates for 96 metal organic frameworks in like an …”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 1:12:46
Prediction Not checkable as stated
Beam: Round-over-round experimentation yields more compound value than broad datasets
“The bet is that the sort of like as the model performance improves, the sample efficiency goes up, and therefore like the compound interest that you get from round over round experimentation will outweigh that, that you would get from a big noisy, but broad da…”
Andy Beam Jul 16, 2026 ▶ 1:13:48
Assertion Not checkable as stated
Gomez-Bombarelli: Modern lab benchtop instruments match decade-old government beamlines
“The instruments we have now are as powerful as a beam line would have been 10 years ago. We're taking measurements today that 10 years ago would have requested you to ask the federal government for a Time slot at two in the morning, somewhere out there you kno…”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 1:17:14
Disclosure
Beam: Lila avoids pre-training from scratch, builds on open-weight models
“We have not decided to take on pre-training as well, just because the black magic that you have to do is, is insane, and we've been gifted, you know, something like a billion dollars worth of compute in the form of open-weight models. So we start with an open-…”
Andy Beam Jul 16, 2026 ▶ 1:21:02
Disclosure
Beam: Lila will open source part of its 1,000 scientific RL environments
“So one of the things that we've developed along the way is a test suite of something like a thousand unique scientific RL environments where you can drop in a frontier model. You can drop in your own model. We drop in our models. So almost surely we're going t…”
Andy Beam Jul 16, 2026 ▶ 1:22:10
Insight
Beam: Verified scientific reasoning traces lift models despite parameter disadvantages
“So like we have just seen incredible lift from showing the model that even if we're at like a parameter disadvantage relative to the frontier models, just showing it an experimentally verified reasoning, reasoning trace, you see just immediate lift when we do …”
Andy Beam Jul 16, 2026 ▶ 1:24:12
Assertion Not checkable as stated
Beam: Lila Sciences likely holds a top 3 global biopharma GPU cluster
“I think that, you know, if we called ourselves a biopharma, we probably would have a top three GPU cluster in the world.”
Andy Beam Jul 16, 2026 ▶ 1:31:13
Opinion
Beam: Materials science is harder for AI than biology
“Well, I think materials are harder. So they have the benefit of like great simulators like that we don't have in bio. Like, in material science, you don't have the, like, mature, high-throughput automation that you have in biology. For me, materials as a subje…”
Andy Beam Jul 16, 2026 ▶ 1:32:26
Opinion
Gomez-Bombarelli: Meta's virtual materials datasets fail to solve real-world chemistry
“Meta has produced tens of millions, hundreds of millions of training data points, but they're all virtual simulations that just don't carry enough water for the thing we actually want to do.”
Rafa Gomez-Bombarelli Jul 16, 2026 ▶ 1:37:52
Assertion Supported
Beam: Reinforcement learning achieves only 5% to 6% GPU FLOP utilization
“And for reinforcement learning, it's always somewhere, like, around five to, like, six percent. So, said differently, that means that we're getting, like, five percent of the actual GPU computing power that we're paying for.”
Andy Beam Jul 16, 2026 ▶ 1:38:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.