Jul 21, 2026 · 1h 29m · latent-space

🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)

Bo Wang · 32m spoken Ci Chu · 32m spoken RJ Haneke · 10m spoken Brandon Anderson · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space AI for Science podcast, Xaira Therapeutics leaders Bo Wang and Ci Chu discuss X-Cell, a diffusion-based virtual cell foundation model trained on massive genome-wide Perturb-seq datasets to predict complex cellular responses to genetic perturbations. They explore how high-throughput causal biological data, novel machine learning architectures, and out-of-context generalization are transforming drug discovery and advancing the realization of dynamic virtual cells.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.3 Guest teaching 2.4 Guest disagreement 0.4 The hosts pushing back 0.3
05100:0020:0040:001:00:001:20:000:41–2:48 · The hosts as informed peer 4/10 Introductions: Latent Space AI for Science and Xaira Leadership RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily.2:48–5:41 · The hosts as informed peer 5/10 Xaira's Core Mission and Three Foundational AI Platforms Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation.5:42–11:02 · The hosts as informed peer 4/10 Navigating Biological Bottlenecks and Causal Data Scarcity RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data.11:03–13:24 · The hosts as informed peer 3/10 Introducing X-Cell and Defining Genetic Perturbations RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences.13:24–20:06 · The hosts as informed peer 6/10 The Evolution of Virtual Cells from Differential Equations to scGPT Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features.20:08–27:48 · The hosts as informed peer 5/10 Why Causal Models Require High-Throughput Pooled Perturb-Seq Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices.27:49–33:29 · The hosts as informed peer 5/10 Industrializing High-Throughput Biology and Maximizing Dataset Diversity RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity.33:29–40:23 · The hosts as informed peer 5/10 Spatial Omics and Multimodal Horizons for Virtual Cell Modeling RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments.40:25–46:52 · The hosts as informed peer 7/10 X-Cell Architecture: Autoregressive Versus Diffusion Language Models Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion.46:53–53:13 · The hosts as informed peer 6/10 Conditioning with Biological Priors and Dissecting Ablation Hierarchies RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata.53:20–1:01:59 · The hosts as informed peer 5/10 Demonstrating Out-of-Context Generalization in Primary and Activated T Cells RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors.1:02:00–1:07:07 · The hosts as informed peer 6/10 Surpassing Linear Baselines and Capturing Context-Specific Biology Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology.1:07:07–1:10:33 · The hosts as informed peer 7/10 Biological Redundancy and Modeling Combinatorial Perturbations Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training.1:10:33–1:18:38 · The hosts as informed peer 5/10 Academic Freedom, Computational Disparities, and Industry Synergy Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs.1:18:39–1:25:23 · The hosts as informed peer 6/10 Commitment to Open Science, Cultivating Taste, and Collaborative Biology Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards.1:25:24–1:28:49 · The hosts as informed peer 5/10 Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing.0:41–2:48 · Guest teaching 0/10 Introductions: Latent Space AI for Science and Xaira Leadership RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily.2:48–5:41 · Guest teaching 1/10 Xaira's Core Mission and Three Foundational AI Platforms Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation.5:42–11:02 · Guest teaching 3/10 Navigating Biological Bottlenecks and Causal Data Scarcity RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data.11:03–13:24 · Guest teaching 2/10 Introducing X-Cell and Defining Genetic Perturbations RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences.13:24–20:06 · Guest teaching 2/10 The Evolution of Virtual Cells from Differential Equations to scGPT Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features.20:08–27:48 · Guest teaching 4/10 Why Causal Models Require High-Throughput Pooled Perturb-Seq Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices.27:49–33:29 · Guest teaching 5/10 Industrializing High-Throughput Biology and Maximizing Dataset Diversity RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity.33:29–40:23 · Guest teaching 2/10 Spatial Omics and Multimodal Horizons for Virtual Cell Modeling RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments.40:25–46:52 · Guest teaching 2/10 X-Cell Architecture: Autoregressive Versus Diffusion Language Models Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion.46:53–53:13 · Guest teaching 2/10 Conditioning with Biological Priors and Dissecting Ablation Hierarchies RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata.53:20–1:01:59 · Guest teaching 3/10 Demonstrating Out-of-Context Generalization in Primary and Activated T Cells RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors.1:02:00–1:07:07 · Guest teaching 3/10 Surpassing Linear Baselines and Capturing Context-Specific Biology Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology.1:07:07–1:10:33 · Guest teaching 2/10 Biological Redundancy and Modeling Combinatorial Perturbations Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training.1:10:33–1:18:38 · Guest teaching 3/10 Academic Freedom, Computational Disparities, and Industry Synergy Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs.1:18:39–1:25:23 · Guest teaching 2/10 Commitment to Open Science, Cultivating Taste, and Collaborative Biology Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards.1:25:24–1:28:49 · Guest teaching 2/10 Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing.0:41–2:48 · Guest disagreement 0/10 Introductions: Latent Space AI for Science and Xaira Leadership RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily.2:48–5:41 · Guest disagreement 0/10 Xaira's Core Mission and Three Foundational AI Platforms Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation.5:42–11:02 · Guest disagreement 0/10 Navigating Biological Bottlenecks and Causal Data Scarcity RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data.11:03–13:24 · Guest disagreement 0/10 Introducing X-Cell and Defining Genetic Perturbations RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences.13:24–20:06 · Guest disagreement 0/10 The Evolution of Virtual Cells from Differential Equations to scGPT Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features.20:08–27:48 · Guest disagreement 1/10 Why Causal Models Require High-Throughput Pooled Perturb-Seq Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices.27:49–33:29 · Guest disagreement 2/10 Industrializing High-Throughput Biology and Maximizing Dataset Diversity RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity.33:29–40:23 · Guest disagreement 0/10 Spatial Omics and Multimodal Horizons for Virtual Cell Modeling RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments.40:25–46:52 · Guest disagreement 1/10 X-Cell Architecture: Autoregressive Versus Diffusion Language Models Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion.46:53–53:13 · Guest disagreement 0/10 Conditioning with Biological Priors and Dissecting Ablation Hierarchies RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata.53:20–1:01:59 · Guest disagreement 0/10 Demonstrating Out-of-Context Generalization in Primary and Activated T Cells RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors.1:02:00–1:07:07 · Guest disagreement 1/10 Surpassing Linear Baselines and Capturing Context-Specific Biology Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology.1:07:07–1:10:33 · Guest disagreement 0/10 Biological Redundancy and Modeling Combinatorial Perturbations Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training.1:10:33–1:18:38 · Guest disagreement 1/10 Academic Freedom, Computational Disparities, and Industry Synergy Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs.1:18:39–1:25:23 · Guest disagreement 0/10 Commitment to Open Science, Cultivating Taste, and Collaborative Biology Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards.1:25:24–1:28:49 · Guest disagreement 0/10 Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing.0:41–2:48 · The hosts pushing back 0/10 Introductions: Latent Space AI for Science and Xaira Leadership RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily.2:48–5:41 · The hosts pushing back 0/10 Xaira's Core Mission and Three Foundational AI Platforms Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation.5:42–11:02 · The hosts pushing back 0/10 Navigating Biological Bottlenecks and Causal Data Scarcity RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data.11:03–13:24 · The hosts pushing back 0/10 Introducing X-Cell and Defining Genetic Perturbations RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences.13:24–20:06 · The hosts pushing back 0/10 The Evolution of Virtual Cells from Differential Equations to scGPT Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features.20:08–27:48 · The hosts pushing back 0/10 Why Causal Models Require High-Throughput Pooled Perturb-Seq Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices.27:49–33:29 · The hosts pushing back 1/10 Industrializing High-Throughput Biology and Maximizing Dataset Diversity RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity.33:29–40:23 · The hosts pushing back 0/10 Spatial Omics and Multimodal Horizons for Virtual Cell Modeling RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments.40:25–46:52 · The hosts pushing back 1/10 X-Cell Architecture: Autoregressive Versus Diffusion Language Models Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion.46:53–53:13 · The hosts pushing back 0/10 Conditioning with Biological Priors and Dissecting Ablation Hierarchies RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata.53:20–1:01:59 · The hosts pushing back 0/10 Demonstrating Out-of-Context Generalization in Primary and Activated T Cells RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors.1:02:00–1:07:07 · The hosts pushing back 0/10 Surpassing Linear Baselines and Capturing Context-Specific Biology Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology.1:07:07–1:10:33 · The hosts pushing back 0/10 Biological Redundancy and Modeling Combinatorial Perturbations Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training.1:10:33–1:18:38 · The hosts pushing back 3/10 Academic Freedom, Computational Disparities, and Industry Synergy Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs.1:18:39–1:25:23 · The hosts pushing back 0/10 Commitment to Open Science, Cultivating Taste, and Collaborative Biology Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards.1:25:24–1:28:49 · The hosts pushing back 0/10 Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 30:52 Correcting flawed premise on stem cell models

Ci Chu immediately rejects RJ's premise that Xaira started primarily with stressed stem cells, explaining they began with robust cancer and immortalized cell lines before deliberately moving into multi-lineage stem cells.

Hardest push from the hosts ▶ 1:14:33 Challenging academic funding necessity

RJ refuses the standard assumption that academia deserves funding over industry, directly demanding Bo justify why government money shouldn't just go to well-resourced private companies.

Biggest teaching moment ▶ 21:10 Why observational data cannot establish causal networks

Ci Chu provides a clear mathematical and biological breakdown demonstrating that correlated expression in descriptive atlases can fit infinite causal structures, proving observational models cannot solve counterfactual perturbation tasks.

The host holds their own ▶ 43:57 Brandon challenges diffusion versus set-based transformers

Brandon demonstrates deep domain knowledge of transformer mechanics, directly challenging the field's bias toward autoregressive and diffusion models over natural permutation-invariant set operations.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions: Latent Space AI for Science and Xaira Leadership 4000 RJ Haneke opens by outlining the core premise of AI for science versus traditional enterprise software and sets up the guest introductions. The guests introduce their respective backgrounds at Xaira, University of Toronto, Insitro, and Verily.
Xaira's Core Mission and Three Foundational AI Platforms 5100 Brandon asks about Xaira's mission, thesis, and massive funding round. Ci and Bo explain the three platform pillars: protein design, virtual cell modeling, and patient representation.
Navigating Biological Bottlenecks and Causal Data Scarcity 4300 RJ asks where the biological bottlenecks lie and how clinical data connects back to early discovery. Ci explains that unlike protein structure prediction with decades of PDB data, virtual cell modeling and patient prediction severely lack high-quality causal data.
Introducing X-Cell and Defining Genetic Perturbations 3200 RJ prompts an explanation of X-Cell and perturbation for lay listeners. Ci and RJ walk through in silico gene knockdowns, gene regulatory pathways, and their phenotypic consequences.
The Evolution of Virtual Cells from Differential Equations to scGPT 6200 Brandon and Bo trace the history of virtual cells from differential equations to single-cell foundation models like scGPT. The hosts engage on batch effects and spurious features.
Why Causal Models Require High-Throughput Pooled Perturb-Seq 5410 Ci explains why observational datasets like CELLxGENE are underpowered for causality due to observational confounding. He details how pooled Perturb-seq with CRISPR guide barcodes enables scalable 2D perturbation matrices.
Industrializing High-Throughput Biology and Maximizing Dataset Diversity 5521 RJ asks about the validity of using stem cells as proxies for mature cell types. Ci clarifies that they began with immortalized lines before expanding into pan-differentiated iPSC multi-cell screens to capture biological diversity.
Spatial Omics and Multimodal Horizons for Virtual Cell Modeling 5200 RJ asks about moving beyond dissociated cells to spatial context. Ci and Bo describe spatial transcriptomics and multi-modal integration as critical frontiers for modeling tumor-immune microenvironments.
X-Cell Architecture: Autoregressive Versus Diffusion Language Models 7211 Brandon pushes back on using autoregressive or diffusion models over unordered gene sets, pointing out transformers natively operate on sets. Bo clarifies that generative full-transcriptome decoding requires iterative refinement via diffusion.
Conditioning with Biological Priors and Dissecting Ablation Hierarchies 6200 RJ and Brandon probe the impact of prior biological conditioning and ablation rankings. Bo details the five prior types used in X-Cell and ranks dataset scale/quality above architecture and prior metadata.
Demonstrating Out-of-Context Generalization in Primary and Activated T Cells 5300 RJ questions the ROI of large computational runs versus wet-lab spending. Ci presents validation data showing X-Cell generalizing from resting T cells to activated T cells and primary patient donors.
Surpassing Linear Baselines and Capturing Context-Specific Biology 6310 Brandon brings up the benchmark problem where foundation models struggle to beat linear baselines. Bo and Ci explain that sparse MAE metrics favor linear mean baselines, whereas Pearson delta on real shifts proves X-Cell's nonlinear capture of context-dependent biology.
Biological Redundancy and Modeling Combinatorial Perturbations 7200 Brandon cites developmental dosage compensation and regulatory redundancy to ask about multi-gene perturbations. Ci and Bo agree and explain how PPI networks and in silico combinatorial sampling expand single-knockout training.
Academic Freedom, Computational Disparities, and Industry Synergy 5313 Bo reflects on the resource gap between academia and industry. RJ pushes back directly, questioning why public money should subsidize academia instead of flowing directly to well-resourced industry labs.
Commitment to Open Science, Cultivating Taste, and Collaborative Biology 6200 Brandon asks how students develop scientific taste when code and synthesis can be outsourced to LLMs. Bo and Ci emphasize hands-on debugging, wet-lab validation, and open-source community standards.
Magic Wand Breakthroughs: High-Throughput Proteomics and Live Longitudinal Sequencing 5200 RJ asks for magic-wand technological breakthroughs. Ci wishes for high-throughput single-cell proteomics, while Bo wants live, non-destructive longitudinal single-cell sequencing.

Statements from this episode (34)

Insight
Haneke: Lab Experimentation Differentiates AI for Science From B2B SaaS
“One of the themes that has run through the podcast is how the lab and experimentation and the real world have probably the biggest impact and have the most relevance to whether something is AI for science or something like B to B SAS.”
RJ Haneke Jul 21, 2026 ▶ 0:53
Disclosure
Chu: Xaira is building three core AI platforms for drug discovery
“There are three main AI platforms that we're building here. The first one is Protein Design, work that spun out of our co-founder Dr. David Baker's group from UW. A lot of the current generation of protein designers are here in the company. So there, the think…”
Ci Chu Jul 21, 2026 ▶ 3:36
Assertion Not checkable as stated
Wang: X-Cell Is First to Predict Unseen Cell Line Perturbations
“One of the rewarding signals I receive after we develop Excel is that like it's a wow moment from biologists that this is the first time biologists actually find the model can predict exactly how these unseen cell lines kind of respond to different perturbatio…”
Bo Wang Jul 21, 2026 ▶ 7:54
Insight
Chu: 70 years of curated data drove rapid protein design AI progress
“And in Protein design space. I think that's where we have seen the most rapid progress so far. That's partially because we have a lot of data, high quality data over 70 years curated by the entire community.”
Ci Chu Jul 21, 2026 ▶ 9:10
Insight
Chu: Virtual cell models lag due to high-quality data limitations
“In the other domains, such as clinical model prediction, such as virtual cell, we are nowhere near the same kind of massive data that are high quality, and I think it's mainly a data limitation issue.”
Ci Chu Jul 21, 2026 ▶ 9:46
Disclosure
Wang: X-Cell predicts cell responses to genetic and chemical perturbations
“X-Cell is Xera's first virtual cell models. It is an AI model that can predict the response to genetic perturbations. Certainly we can extend it to other type of interventions such as drug perturbations chemical perturbations, et cetera.”
Bo Wang Jul 21, 2026 ▶ 11:11
Insight
Wang: Differential Equation Models Failed Because Biology Is Too Complex
“Largely speaking, that was a failed attempt in the sense that the biology is just way too complicated to write in a few predefined set of differential equations.”
Bo Wang Jul 21, 2026 ▶ 14:41
Prediction Not checkable as stated
Wang: Virtual cell AI models will eventually replace physical cellular experiments
“Eventually we kind of, we can replace all the cellular experiments by simply running simulations on computer without even running the actual Experiments.”
Bo Wang Jul 21, 2026 ▶ 17:35
Opinion
Wang: Virtual cells are a much broader concept than foundation models
“What's happening for this field is that we are lacking a concrete definition of virtual cells and people almost equate foundation model With virtual cell, but in my view, virtual cell is probably a much broader concept than just foundation models.”
Bo Wang Jul 21, 2026 ▶ 19:13
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Ci Chu Jul 21, 2026 ▶ 21:47
Insight
Chu: Observational Profiling Data Cannot Learn Biological Causality
“Fundamentally, we believe our provisional data are underpowered to learn causality, truly.”
Ci Chu Jul 21, 2026 ▶ 22:45
Opinion
Chu: 2D Perturb-Seq Datasets Will Power Biological Foundation Models
“And I think it's these type of rich two-D datasets that power the training of foundation models of biology.”
Ci Chu Jul 21, 2026 ▶ 27:43
Assertion Supported
Chu: Xaira's XLS Orion was the world's largest Perturb-seq release
“When we started data generation, so we put out the method that I talked about, as well as the first two datasets, which is the world's largest perturbsic data released at the time. Last June in the preprint, we call it dataset XLS Orion.”
Ci Chu Jul 21, 2026 ▶ 31:00
Assertion Supported
Chu: Xaira ran genome-scale perturbations across 10 differentiated iPSC cell types
“So effectively we differentiate iPSC into 10 different cell types in one single experiment without restriction. And we did a genome scale perturbation across them. So you can imagine instead of just generating 10,000 different biological experiments, We did 10…”
Ci Chu Jul 21, 2026 ▶ 32:34
Insight
Chu: Virtual cell models require biological diversity, not just raw cell counts
“We think that in the beginning phase of data collection, as Bo said, I think we're just in the early days of virtual cell building, context and diversity and richness of the data matters. It's not just the total number of cells or total number of sequencing re…”
Ci Chu Jul 21, 2026 ▶ 32:57
Prediction Open · timeframe Jul 2031
Wang: Virtual cell models will integrate multi-omics to predict cell states
“So eventually what I predict is that a virtual cell model will be able to integrate not only RNC, can integrate more functionally Related, for example, proteinomics or other regulatory side of omics such as ataxic to overall combine all your descriptive omics …”
Bo Wang Jul 21, 2026 ▶ 35:52
Prediction Open · timeframe Jul 2029
Wang: Next version of X-Cell will infer spatial cellular representations
“However, definitely our ongoing work and the next version of Excel will be able to infer the spatially aware representations for different cells.”
Bo Wang Jul 21, 2026 ▶ 40:13
Insight
Wang: Shuffling gene order does not alter expression biology
“For DNA sequences, the order of ATTG make total sense to us, right? But for expression data, they're literally just matrices. So it's really hard to assume, ah, an inherent order of genes. Even if we shuffle the order of genes, I think the biology doesn't chan…”
Bo Wang Jul 21, 2026 ▶ 42:03
Insight
Wang: Autoregressive Models Type Sequentially, Diffusion Models Iteratively Edit
“What's the difference between autoregressive training versus diffusion language models is that it kind of, you can think of autoregressive training as typing. There is, for example, I like coffee, you have to type I, and then like, and coffee. There's inherent…”
Bo Wang Jul 21, 2026 ▶ 42:53
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Bo Wang Jul 21, 2026 ▶ 52:00
Insight
Wang: Data Quality and Scale Matter More Than Architecture or Priors
“Overall, if we have to give an order, my order would be the quality among scale of the datasets, and then the architecture, and then the prior knowledge.”
Bo Wang Jul 21, 2026 ▶ 52:51
Assertion Supported
Chu: X-Cell Predicts Perturbation Effects In Unseen Activated T Cells
“Critically, Excel has not seen active cell T cells. And it's able to make accurate prediction, not only on the known biology, the TCR complex, predicting their effect accurately, that these are going to inactive the T cells, which It's exactly what we would ex…”
Ci Chu Jul 21, 2026 ▶ 57:35
Assertion Supported
Wang: Average single-cell profiles can beat technical replicates on MAE
“Because single cell data set are so sparse, the average profiles of all the cells, certainly you kind of, you can imagine is a great minimum kind of local optimum to minimize the MAEs. This is why sometimes the average profile of cells has lower MAEs even than…”
Bo Wang Jul 21, 2026 ▶ 1:02:57
Disclosure
Wang: X-Cell trains on causal perturbation data unlike static models
“What sets Excel different from these static expression models such as SGB or geneformers is that we actually, instead of training on gene expression datasets, we train on causal datasets. We train on massive amount of genome-wide perturbation datasets so that …”
Bo Wang Jul 21, 2026 ▶ 1:03:52
Prediction Not checkable as stated
Wang: Foundation models on causal data will beat linear baselines
“I believe that foundation model or other more complicated AI models that trend on the right data will outperform these linear models in harder tasks, particularly in generalization tasks.”
Bo Wang Jul 21, 2026 ▶ 1:04:25
Assertion Supported
Chu: Xaira is first to combine seven genome-wide Perturb-seq campaigns
“It is the first time that someone can put together not just one perturbseek, but seven genome-wide perturbseek campaigns together.”
Ci Chu Jul 21, 2026 ▶ 1:05:59
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Bo Wang Jul 21, 2026 ▶ 1:09:55
Opinion
Wang: Academia is the primary source of innovation in biotech
“I still deeply believe that academic is the main source of innovation for the whole field, and particularly when it comes to biotech.”
Bo Wang Jul 21, 2026 ▶ 1:14:10
Assertion Not checkable as stated
Wang: Academic researchers have unique access to restricted healthcare datasets
“By being a professor in academic, we also have access to lots of healthcare data sets, which are very hard for industry to access due to many illegal reasons or regulatory reasons.”
Bo Wang Jul 21, 2026 ▶ 1:16:27
Insight
Chu: Academia drives accidental discovery while industry excels at scaling data
“These innovations take so long and the discovery process can be so accidental, right, that It's perhaps not ideal for pure industry to take on, but once they show early promise, scaling them, and robustifying them, and generating data that's not only massive, …”
Ci Chu Jul 21, 2026 ▶ 1:17:57
Assertion Not checkable as stated
Wang: scGPT Is Widely Used by Pharma Single-Cell Teams
“This is also why SCGBT quickly become one of the most widely used single cell foundation model in pharma companies.”
Bo Wang Jul 21, 2026 ▶ 1:19:48
Insight
Wang: Agentic AI Inverts Research Time From Coding to Debugging
“We used to spend lots of time coding, a little bit of time just debugging, but now we let the agent do most of the coding, but we spend most of time debugging, which seems to be definitely interesting to me”
Bo Wang Jul 21, 2026 ▶ 1:22:55
Insight
Chu: High-throughput single-cell proteomics will power next-gen biological models
“RNA is amazing. It foreshadows which proteins are going to get made, but protein by and large are the functional units in a cell. Not only does their abundance matter, their post-translational modification matter, their localization in a cell matter. If you ca…”
Ci Chu Jul 21, 2026 ▶ 1:26:02
Insight
Wang: Destructive Cell Sequencing Limits Existing Models to Static Snapshots
“In order to sequence the cell, you have to kill the cell, right? So can we have a technology that can measure the cell states at different time points for the same set of cells? I think that will bring a very different dimension to the data set so that we can …”
Bo Wang Jul 21, 2026 ▶ 1:26:51
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.