Jul 21, 2026 · 1h 29m · latent-space
🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space AI for Science podcast, Xaira Therapeutics leaders Bo Wang and Ci Chu discuss X-Cell, a diffusion-based virtual cell foundation model trained on massive genome-wide Perturb-seq datasets to predict complex cellular responses to genetic perturbations. They explore how high-throughput causal biological data, novel machine learning architectures, and out-of-context generalization are transforming drug discovery and advancing the realization of dynamic virtual cells.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ci Chu immediately rejects RJ's premise that Xaira started primarily with stressed stem cells, explaining they began with robust cancer and immortalized cell lines before deliberately moving into multi-lineage stem cells.
Hardest push from the hosts ▶ 1:14:33 Challenging academic funding necessityRJ refuses the standard assumption that academia deserves funding over industry, directly demanding Bo justify why government money shouldn't just go to well-resourced private companies.
Biggest teaching moment ▶ 21:10 Why observational data cannot establish causal networksCi Chu provides a clear mathematical and biological breakdown demonstrating that correlated expression in descriptive atlases can fit infinite causal structures, proving observational models cannot solve counterfactual perturbation tasks.
The host holds their own ▶ 43:57 Brandon challenges diffusion versus set-based transformersBrandon demonstrates deep domain knowledge of transformer mechanics, directly challenging the field's bias toward autoregressive and diffusion models over natural permutation-invariant set operations.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions: Latent Space AI for Science and Xaira Leadership | 4 | 0 | 0 | 0 | RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily. | |
| Xaira's Core Mission and Three Foundational AI Platforms | 5 | 1 | 0 | 0 | Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation. | |
| Navigating Biological Bottlenecks and Causal Data Scarcity | 4 | 3 | 0 | 0 | RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data. | |
| Introducing X-Cell and Defining Genetic Perturbations | 3 | 2 | 0 | 0 | RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences. | |
| The Evolution of Virtual Cells from Differential Equations to scGPT | 6 | 2 | 0 | 0 | Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features. | |
| Why Causal Models Require High-Throughput Pooled Perturb-Seq | 5 | 4 | 1 | 0 | Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices. | |
| Industrializing High-Throughput Biology and Maximizing Dataset Diversity | 5 | 5 | 2 | 1 | RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity. | |
| Spatial Omics and Multimodal Horizons for Virtual Cell Modeling | 5 | 2 | 0 | 0 | RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments. | |
| X-Cell Architecture: Autoregressive Versus Diffusion Language Models | 7 | 2 | 1 | 1 | Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion. | |
| Conditioning with Biological Priors and Dissecting Ablation Hierarchies | 6 | 2 | 0 | 0 | RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata. | |
| Demonstrating Out-of-Context Generalization in Primary and Activated T Cells | 5 | 3 | 0 | 0 | RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors. | |
| Surpassing Linear Baselines and Capturing Context-Specific Biology | 6 | 3 | 1 | 0 | Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology. | |
| Biological Redundancy and Modeling Combinatorial Perturbations | 7 | 2 | 0 | 0 | Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training. | |
| Academic Freedom, Computational Disparities, and Industry Synergy | 5 | 3 | 1 | 3 | Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs. | |
| Commitment to Open Science, Cultivating Taste, and Collaborative Biology | 6 | 2 | 0 | 0 | Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards. | |
| Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing | 5 | 2 | 0 | 0 | RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing. |