Jul 16, 2026 · 1h 41m · latent-space
🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Latent Space Science, Lila Sciences leaders Andy Beam and Rafa Gomez-Bombarelli discuss building a frontier scientific reasoning platform where automated physical laboratories act as empirical verifiers for reinforcement learning across biology and materials science.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Rafa firmly pushes back on the host's assertion that the future of chemistry is language, arguing that chemists do not think in English and domain-specific geometric representations remain essential.
Hardest push from the hosts ▶ 15:20 Brandon challenges the relevance of AI safety in Lila's labBrandon directly challenges the premise that AI safety is a practical concern in Lila's current lab automation setting, arguing that malicious risk or catastrophic emergent outputs are negligible.
Biggest teaching moment ▶ 1:33:45 Rafa educates on materials economics versus biotech assetsRafa dismantles the host's analogy between material product validation and clinical trials, demonstrating how approved drugs command massive pricing power while critical materials like superconductors remain anonymous and low-margin.
The host holds their own ▶ 29:45 Brandon cites Sri Kosuri on ML business model paradoxesBrandon challenges Lila's core data-generation premise by citing Sri Kosuri's critique on whether models are even needed once requisite domain datasets have already been generated.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Rafa Gomez-Bombarelli's Journey in Computational Materials | 3 | 1 | 0 | 0 | Rafa introduces his background in computational materials science, early generative models for molecules, and the transition toward validating AI in wet labs. The hosts primarily listen as Rafa outlines his trajectory from Harvard and MIT to Lila. | |
| Lila's Core Thesis: Nature as a Verifiable Reward | 7 | 3 | 1 | 4 | Andy introduces the core thesis of reinforcement learning with nature as a verifiable reward. Brandon demonstrates solid industry domain knowledge by citing Escalante Bio's runtime essay and questioning token value versus commoditized NGS reads. | |
| Lab Architecture and Modular Automation | 6 | 2 | 1 | 3 | Andy describes the lab's modular architecture using a planar motor physical transport layer and the PCI bus analogy. RJ probes into how automation handles biological versus material systems and how humans fit below the API line. | |
| Laboratory Safety and Non-Intuitive Discoveries | 6 | 4 | 2 | 6 | Brandon pushes back against the premise that lab safety is an urgent problem for Lila's current scope. Rafa and Andy explain why chemical EHS, unexpected equipment stresses, and non-intuitive emergent behaviors require proactive guardrails. | |
| Scientific Rigor, Human Synergy, and Green Hydrogen | 6 | 3 | 1 | 3 | RJ references controversies surrounding automated synthesis measurements at Berkeley Lab to ask how Lila ensures rigor. Rafa and Andy explain how high-dimensional environmental logging and automated re-testing validate counterintuitive results like non-platinum catalysts. | |
| RL Pathologies, Chain of Thought, and Tool Calling | 5 | 2 | 1 | 3 | Andy shares humorous RL pathologies including models cursing during plate layout tasks and chain-of-thought repetition. RJ drills down to confirm whether RL is executing purely on reasoning tokens or triggering physical lab tool calls. | |
| Lila's Model-First Strategy and Cross-Domain Generalization | 7 | 4 | 2 | 5 | Brandon quotes Sri Kosuri on the circularity of ML data requirements, and RJ questions whether cross-domain transfer exists across fundamentally different physical scales. Andy argues that unified reasoning across 10 trillion multimodal scientific tokens beats narrow vertical models. | |
| Scientific Representation: Language vs Domain Architectures | 6 | 5 | 4 | 3 | RJ posits that the future of all chemistry is language. Rafa directly reframes this, noting that chemists do not think in English and citing Demis Hassabis regarding why geometric and domain-specific architectures remain essential tools. | |
| Materials Science Capabilities and Rapid Formulations | 6 | 3 | 1 | 4 | Rafa details physical science workflows including quantum dots, coatings, and liquid formulation. RJ raises the difficulty of scaling physical materials to market over 10-15 year horizons, prompting Rafa to discuss techno-economic evaluation agents. | |
| In Vivo CAR-T Breakthrough and Virtual Startups | 5 | 2 | 0 | 1 | Andy walks through Lila's in vivo CAR-T validation, contrasting their 6-month non-human primate milestone against Capstan's multi-year timeline. He explains the 'zero FTE virtual startup' commercial model enabled by the platform. | |
| Translational Bottlenecks and Scientific Serendipity | 7 | 3 | 1 | 5 | RJ and Brandon emphasize that clinical trial translation (5-8% success) is the primary bottleneck rather than early discovery. Andy and Rafa acknowledge the challenge and argue that high-throughput optimization shifts preclinical success probabilities. | |
| Machine Creativity and Open-Ended Exploration | 4 | 2 | 0 | 1 | Andy introduces Ken Stanley's open-endedness research at Lila and presents video footage of the automated lab, highlighting planar levitation and VLM-based automation of legacy hardware. | |
| Lab Orchestration, Scheduling, and Fast Feedback Loops | 6 | 3 | 1 | 4 | Brandon presses on experimental runtime tradeoffs between broad multiplexing and rapid serial iteration. Rafa explains how developing high-speed proxy measurements allowed 2500x acceleration in gas sorption testing over traditional BET methods. | |
| Rapid Instrument Onboarding and Facility Expansion | 5 | 2 | 0 | 2 | RJ asks if Lila risks saturating amenable experimental problems. Andy and Rafa explain that standardized interfaces will turn instrument onboarding into plug-and-play operations as new analytical devices emerge. | |
| The 10-Trillion Scientific Token Dataset | 7 | 3 | 1 | 4 | Brandon breaks down the arithmetic of Lila's 10-trillion token claim relative to genomes and foundation models. Andy clarifies that these are post-training reasoning traces with verified tool executions built atop open-weight frontier models. | |
| Flagship Pioneering Origins and Distinct Identity | 6 | 2 | 1 | 3 | Brandon asks how Lila differentiates from traditional single-asset Flagship Pioneering biotechs. Andy details their model-first identity, external capital structure, and heavy GPU infrastructure investment. | |
| Materials Science vs Biology: Challenges and Economics | 6 | 6 | 4 | 4 | Brandon asks whether materials or biology is harder. When Brandon argues that supply chains and validation are equivalent, Rafa counters sharply by detailing the stark economic divergence between named big pharma drugs and commoditized materials like superconductors. | |
| Magic Wand Bottlenecks: Sim-to-Real and Flop Utilization | 5 | 3 | 1 | 2 | Rafa identifies the sim-to-real gap in physics simulations as his ideal bottleneck to erase, while Andy targets low GPU mean flop utilization in RL pipelines. Andy explains how parallel expert models decouple physical rollout latency from GPU compute. |