Everything Ci Chu said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Chu: Observational Profiling Data Cannot Learn Biological Causality
“Fundamentally, we believe our provisional data are underpowered to learn causality, truly.”
Chu: X-Cell Predicts Perturbation Effects In Unseen Activated T Cells
“Critically, Excel has not seen active cell T cells. And it's able to make accurate prediction, not only on the known biology, the TCR complex, predicting their effect accurately, that these are going to inactive the T cells, which It's exactly what we would ex…”
Chu: Xaira is first to combine seven genome-wide Perturb-seq campaigns
“It is the first time that someone can put together not just one perturbseek, but seven genome-wide perturbseek campaigns together.”
Chu: Academia drives accidental discovery while industry excels at scaling data
“These innovations take so long and the discovery process can be so accidental, right, that It's perhaps not ideal for pure industry to take on, but once they show early promise, scaling them, and robustifying them, and generating data that's not only massive, …”
Chu: Virtual cell models lag due to high-quality data limitations
“In the other domains, such as clinical model prediction, such as virtual cell, we are nowhere near the same kind of massive data that are high quality, and I think it's mainly a data limitation issue.”
Chu: 2D Perturb-Seq Datasets Will Power Biological Foundation Models
“And I think it's these type of rich two-D datasets that power the training of foundation models of biology.”
Chu: Xaira's XLS Orion was the world's largest Perturb-seq release
“When we started data generation, so we put out the method that I talked about, as well as the first two datasets, which is the world's largest perturbsic data released at the time. Last June in the preprint, we call it dataset XLS Orion.”
Chu: Xaira ran genome-scale perturbations across 10 differentiated iPSC cell types
“So effectively we differentiate iPSC into 10 different cell types in one single experiment without restriction. And we did a genome scale perturbation across them. So you can imagine instead of just generating 10,000 different biological experiments, We did 10…”
Chu: Virtual cell models require biological diversity, not just raw cell counts
“We think that in the beginning phase of data collection, as Bo said, I think we're just in the early days of virtual cell building, context and diversity and richness of the data matters. It's not just the total number of cells or total number of sequencing re…”
Chu: High-throughput single-cell proteomics will power next-gen biological models
“RNA is amazing. It foreshadows which proteins are going to get made, but protein by and large are the functional units in a cell. Not only does their abundance matter, their post-translational modification matter, their localization in a cell matter. If you ca…”
Chu: Xaira is building three core AI platforms for drug discovery
“There are three main AI platforms that we're building here. The first one is Protein Design, work that spun out of our co-founder Dr. David Baker's group from UW. A lot of the current generation of protein designers are here in the company. So there, the think…”
Chu: 70 years of curated data drove rapid protein design AI progress
“And in Protein design space. I think that's where we have seen the most rapid progress so far. That's partially because we have a lot of data, high quality data over 70 years curated by the entire community.”