Feb 25, 2025 · 57m · no-priors

No Priors Ep. 103 | With Vevo Therapeutics and the Arc Institute

Nima Alidoust · 17m spoken Patrick Hsu · 10m spoken Dave Burke · 8m spoken Hani Goodarzi · 6m spoken Sarah Guo · 5m spoken Johnny Yu · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Leaders from Vevo Therapeutics and the Arc Institute introduce the Tahoe-100 dataset, discussing how massive single-cell perturbation data and foundation models will transform drug discovery through predictive virtual cell engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 10.9% of the talking time here. How this is scored →

The hosts as informed peer 5.4 Guest teaching 4.0 Guest disagreement 1.7 The hosts pushing back 2.0
05100:0015:0030:0045:001:34–4:18 · The hosts as informed peer 4/10 Significance of Tahoe-100 and Single-Cell Perturbation Datasets Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology.4:18–9:09 · The hosts as informed peer 5/10 The Virtual Cell Architecture and Transcriptomic Equalizers Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM.9:09–15:17 · The hosts as informed peer 5/10 Batch Effects, Causal Inference, and Biological Token Scaling Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens.15:17–19:09 · The hosts as informed peer 4/10 Vevo's Mosaic Platform and Hypothesis-Free Data Generation Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation.19:09–25:01 · The hosts as informed peer 7/10 Scaling Scientific Intuition and Sutton's Bitter Lesson Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization.25:01–32:42 · The hosts as informed peer 4/10 Open Sourcing the Atlas and Automated Scientific Agents Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents.32:42–41:37 · The hosts as informed peer 6/10 Evaluating Virtual Cells and Reimagining Drug Discovery Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate.41:37–44:15 · The hosts as informed peer 6/10 Choosing Abstraction Levels from Transcriptomes to Organoids Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals.44:15–52:21 · The hosts as informed peer 6/10 Platform Biotechs, Global Competition, and Lean Science Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models.52:21–57:20 · The hosts as informed peer 7/10 Overcoming AI Skepticism and Biology's GPT Progression Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2.1:34–4:18 · Guest teaching 3/10 Significance of Tahoe-100 and Single-Cell Perturbation Datasets Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology.4:18–9:09 · Guest teaching 4/10 The Virtual Cell Architecture and Transcriptomic Equalizers Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM.9:09–15:17 · Guest teaching 5/10 Batch Effects, Causal Inference, and Biological Token Scaling Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens.15:17–19:09 · Guest teaching 4/10 Vevo's Mosaic Platform and Hypothesis-Free Data Generation Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation.19:09–25:01 · Guest teaching 3/10 Scaling Scientific Intuition and Sutton's Bitter Lesson Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization.25:01–32:42 · Guest teaching 4/10 Open Sourcing the Atlas and Automated Scientific Agents Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents.32:42–41:37 · Guest teaching 5/10 Evaluating Virtual Cells and Reimagining Drug Discovery Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate.41:37–44:15 · Guest teaching 4/10 Choosing Abstraction Levels from Transcriptomes to Organoids Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals.44:15–52:21 · Guest teaching 4/10 Platform Biotechs, Global Competition, and Lean Science Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models.52:21–57:20 · Guest teaching 4/10 Overcoming AI Skepticism and Biology's GPT Progression Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2.1:34–4:18 · Guest disagreement 1/10 Significance of Tahoe-100 and Single-Cell Perturbation Datasets Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology.4:18–9:09 · Guest disagreement 1/10 The Virtual Cell Architecture and Transcriptomic Equalizers Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM.9:09–15:17 · Guest disagreement 2/10 Batch Effects, Causal Inference, and Biological Token Scaling Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens.15:17–19:09 · Guest disagreement 1/10 Vevo's Mosaic Platform and Hypothesis-Free Data Generation Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation.19:09–25:01 · Guest disagreement 2/10 Scaling Scientific Intuition and Sutton's Bitter Lesson Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization.25:01–32:42 · Guest disagreement 2/10 Open Sourcing the Atlas and Automated Scientific Agents Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents.32:42–41:37 · Guest disagreement 2/10 Evaluating Virtual Cells and Reimagining Drug Discovery Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate.41:37–44:15 · Guest disagreement 1/10 Choosing Abstraction Levels from Transcriptomes to Organoids Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals.44:15–52:21 · Guest disagreement 3/10 Platform Biotechs, Global Competition, and Lean Science Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models.52:21–57:20 · Guest disagreement 2/10 Overcoming AI Skepticism and Biology's GPT Progression Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2.1:34–4:18 · The hosts pushing back 1/10 Significance of Tahoe-100 and Single-Cell Perturbation Datasets Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology.4:18–9:09 · The hosts pushing back 2/10 The Virtual Cell Architecture and Transcriptomic Equalizers Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM.9:09–15:17 · The hosts pushing back 2/10 Batch Effects, Causal Inference, and Biological Token Scaling Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens.15:17–19:09 · The hosts pushing back 1/10 Vevo's Mosaic Platform and Hypothesis-Free Data Generation Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation.19:09–25:01 · The hosts pushing back 3/10 Scaling Scientific Intuition and Sutton's Bitter Lesson Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization.25:01–32:42 · The hosts pushing back 1/10 Open Sourcing the Atlas and Automated Scientific Agents Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents.32:42–41:37 · The hosts pushing back 3/10 Evaluating Virtual Cells and Reimagining Drug Discovery Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate.41:37–44:15 · The hosts pushing back 2/10 Choosing Abstraction Levels from Transcriptomes to Organoids Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals.44:15–52:21 · The hosts pushing back 2/10 Platform Biotechs, Global Competition, and Lean Science Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models.52:21–57:20 · The hosts pushing back 3/10 Overcoming AI Skepticism and Biology's GPT Progression Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 22.2% · guest 77.8%0:00 · the hosts 22.2% · guest 77.8%3:00 · the hosts 7.7% · guest 92.3%3:00 · the hosts 7.7% · guest 92.3%6:00 · the hosts 0.7% · guest 99.3%6:00 · the hosts 0.7% · guest 99.3%9:00 · the hosts 6% · guest 94%9:00 · the hosts 6% · guest 94%12:00 · the hosts 6.4% · guest 93.6%12:00 · the hosts 6.4% · guest 93.6%15:00 · the hosts 5.2% · guest 94.8%15:00 · the hosts 5.2% · guest 94.8%18:00 · the hosts 32% · guest 68%18:00 · the hosts 32% · guest 68%21:00 · the hosts 6% · guest 94%21:00 · the hosts 6% · guest 94%24:00 · the hosts 19.9% · guest 80.1%24:00 · the hosts 19.9% · guest 80.1%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 8.6% · guest 91.4%30:00 · the hosts 8.6% · guest 91.4%33:00 · the hosts 14.6% · guest 85.4%33:00 · the hosts 14.6% · guest 85.4%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 5.2% · guest 94.8%39:00 · the hosts 5.2% · guest 94.8%42:00 · the hosts 11.6% · guest 88.4%42:00 · the hosts 11.6% · guest 88.4%45:00 · the hosts 8.8% · guest 91.2%45:00 · the hosts 8.8% · guest 91.2%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 28.5% · guest 71.5%51:00 · the hosts 28.5% · guest 71.5%54:00 · the hosts 5.4% · guest 94.6%54:00 · the hosts 5.4% · guest 94.6%57:00 · the hosts 98.7% · guest 1.3%57:00 · the hosts 98.7% · guest 1.3%
Sharpest disagreement ▶ 49:23 Nima calls out biotech's slow multi-year planning culture

Nima Alidoust forcefully criticizes traditional biotech organizations for announcing multi-year timelines and operating with slow, bloated bureaucracies.

Hardest push from the hosts ▶ 52:21 Sarah confronts the decade-long lack of AI biotech treatments

Sarah presses the panel on why listeners should believe AI will deliver cures now when AI biotechs have pitched big claims for over a decade with few commercial therapies.

Biggest teaching moment ▶ 4:48 Dave Burke's CPU and graphic equalizer transcriptomic model

Dave Burke re-educates the conversation by laying out an engineering mental model of the cell, mapping DNA to ROM, RNA expression to dynamic equalizer RAM, and virtual cell AI to the central CPU.

The host holds their own ▶ 22:11 Sarah frames the shift using Sutton's Bitter Lesson

Sarah demonstrates deep domain expertise by connecting the panel's hypothesis-free data scaling strategy to Rich Sutton's foundational Bitter Lesson in computing.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Significance of Tahoe-100 and Single-Cell Perturbation Datasets 4311 Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology.
The Virtual Cell Architecture and Transcriptomic Equalizers 5412 Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM.
Batch Effects, Causal Inference, and Biological Token Scaling 5522 Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens.
Vevo's Mosaic Platform and Hypothesis-Free Data Generation 4411 Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation.
Scaling Scientific Intuition and Sutton's Bitter Lesson 7323 Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization.
Open Sourcing the Atlas and Automated Scientific Agents 4421 Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents.
Evaluating Virtual Cells and Reimagining Drug Discovery 6523 Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate.
Choosing Abstraction Levels from Transcriptomes to Organoids 6412 Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals.
Platform Biotechs, Global Competition, and Lean Science 6432 Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models.
Overcoming AI Skepticism and Biology's GPT Progression 7423 Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2.

Statements from this episode (23)

Assertion Supported
Johnny Yu: Tahoe-100 is the world's largest single-cell RNA sequencing dataset
“So Tahoe 100 is the world's biggest single cell RNA sequencing data set, and it enables basically a ton of machine learning applications, including things like the virtual cell, but it also enables a lot of drug discovery applications, and broadly in the conte…”
Johnny Yu Feb 25, 2025 ▶ 1:40
Assertion Not checkable as stated
Goodarzi: Cell state AI is data-limited, unlike compute-constrained DNA models
“When it comes to, for example, DNA language models, again, thanks to the field and decades of having sequenced a ton of genomes we are not as much data limited, but Compute and specifically context and how long we can actually consume DNA and what size of inpu…”
Hani Goodarzi Feb 25, 2025 ▶ 6:28
Assertion Supported
Alidoust: Prior public single-cell data totaled only 45M to 60M cells
“Before that, I think the number of human cells that we had, had been collated together it was in the order of 45, fifty million, if you are generous, sixty million single cell data points.”
Nima Alidoust Feb 25, 2025 ▶ 7:40
Assertion Not checkable as stated
Alidoust: Single-cell AI performance holds even after 99% data downsampling
“If you actually reduce the number of the sixty million, you down sample it by like, Even 99%. You know, you just use one percent of that data to train your models. Actually, the model's performance doesn't reduce that much. So it means that the information con…”
Nima Alidoust Feb 25, 2025 ▶ 8:10
Assertion Partly supported
Yu: Tahoe dataset doubles a decade of cumulative single-cell data
“And so this data set, it is basically doubling the size of all the data that's out there cumulatively over the past decade. It covers 50 different cancer models from different patients, so it's cells from 50 different patients. 1200 drug treatments, so it's a …”
Johnny Yu Feb 25, 2025 ▶ 9:51
Assertion Partly supported
Alidoust: Tahoe dataset expands public perturbational single-cell data fiftyfold
“I think when you put all of the perturbational data sets in the world together if you're generous, it's like one to two million single cell data points. And this is publicly available data. We don't know as much about, you know, what, what's inside different o…”
Nima Alidoust Feb 25, 2025 ▶ 12:14
Insight
Alidoust: 100 million single-cell data points equate to 200-300 billion tokens
“Think of it like a cell collection of for this data says 2000 to 5000 genes, and each gene and its expression is basically a token in what we're doing. So 200, like a hundred million single cell data points is akin to around 200 to three hundred billion tokens…”
Nima Alidoust Feb 25, 2025 ▶ 14:56
Assertion Not checkable as stated
Yu: Cancer perturbation pathways apply broadly to neuroscience and immunology
“This data set, even though it's heavily based around these kinds of chemical perturbations of cancer, they also, these pathways are so conserved and fundamental that they broadly apply to the neuroscience space or like to just immune cell development in genera…”
Johnny Yu Feb 25, 2025 ▶ 15:53
Insight
Alidoust: Hypothesis-driven biology slowed progress, but falling costs enable unbiased data
“One thing that has been a Has been slowing the progress in bio is the fact that we have always been super hypothesis driven, and I think it has, the reason is that a lot of these experiments are expensive, you know, that they take a lot of time, a lot of resou…”
Nima Alidoust Feb 25, 2025 ▶ 18:18
Prediction Not checkable as stated
Hsu: High-throughput token generation will dominate biology over mechanistic research
“The vast majority of, you know, mechanistic data that's been generated to date is really made to ask very specific, very well scoped questions and just, you know, way more tokens per experiment is, you know, just, it's just going to be the way to do it.”
Patrick Hsu Feb 25, 2025 ▶ 20:51
Insight
Hsu: Scientific research and creativity fundamentally involve hallucination
“People criticize these models for hallucinating, but if you think about it, the process of scientific research just involves hallucination, right? That's what creativity is.”
Patrick Hsu Feb 25, 2025 ▶ 22:01
Assertion Supported
Hsu: Biological scaling laws hold across proteins, DNA, and inference compute
“We've, we're seeing evidence of scaling laws in biology across proteins, right? That's been shown in the protein language models across DNA, which is what we've shown in our EVO series of models from ARC. We're also seeing inference time scaling laws which in,…”
Patrick Hsu Feb 25, 2025 ▶ 22:27
Disclosure
Burke: Arc Institute built SC Basecamp with 230M single cells
“We've created something called SC Basecamp and you can almost think of it like the Google crawler and index, so we've built this agent that goes onto the internet, And basically mines public single cell sequenced RNA data, and then curates it in a very sort of…”
Dave Burke Feb 25, 2025 ▶ 27:51
Prediction Not checkable as stated
Hsu: All dry lab workflows will be automated by AI agents
“I think it's very clear now that basically all dry lab workflows are going to get automated with agents or with co-pilots”
Patrick Hsu Feb 25, 2025 ▶ 28:52
Disclosure
Yu: Vevo generated Tahoe dataset with four people in three days
“Well, it ended up being actually four people from Vivo, and we did it, I think, over like three days in the end.”
Johnny Yu Feb 25, 2025 ▶ 31:20
Assertion Not checkable as stated
Burke: Current virtual cell models have only around 10% predictive accuracy
“The reality is today the best models are very poor at this. like the predictive ability of the DEGs as we call them is, is in the order of 10%.”
Dave Burke Feb 25, 2025 ▶ 33:19
Opinion
Burke: Virtual cell accuracy is limited by data quality, not architecture
“One of our conjectures is that one of the reasons the models aren't doing well is not just simply model structure. We have a lot of rich structures that we understand in the ML space and machine learning space that the issue is the Data quality.”
Dave Burke Feb 25, 2025 ▶ 33:38
Prediction Not checkable as stated
Yu: Virtual cell models will directly design disease-reversing drugs
“A big part of our future vision and roadmap is that we think there will be a moment where from a virtual cell model, a drug is spit out, and basically the drug will actually cause a healthy disease cell to become a healthy cell again.”
Johnny Yu Feb 25, 2025 ▶ 37:03
Prediction Not checkable as stated
Alidoust: Virtual cells will expand AI drug discovery to systems biology
“Virtual cells, in my opinion, are going to allow us to go beyond the language of structural biology and venture into the language of systems biology and understand how how the drug is interacting with the broader biological system.”
Nima Alidoust Feb 25, 2025 ▶ 41:18
Insight
Burke: Transcriptomic level is the right abstraction for cell modeling
“And so I think our belief around the room is the right level of abstraction is at the sort of transcriptomic level because you have these very complex gene pathways. And so whenever a cell is changing to its environment reacting, that it will be reflected and …”
Dave Burke Feb 25, 2025 ▶ 41:59
Insight
Hsu: The optimal biotech model is between fully virtual and vertically integrated
“Previously the virtual biotech was a concept that was very much in fashion, right? Folks found out just In reality, when you try to do this, even though it looks really good on paper, it's incredibly slow, right? So then folks tried the other way, which is let…”
Patrick Hsu Feb 25, 2025 ▶ 47:25
Assertion Supported
Burke: Evo 2 trained on 9.3 trillion nucleotides predicts BRCA1 pathogenicity
“Like, if you look at the EVO II model, we trained it on 9.3 trillion nucleotides, but we didn't tell it anything about DNA. We just like, here's a lot of DNA on the planet across, you know, every single piece of DNA we could get a hold of. And then what di…”
Dave Burke Feb 25, 2025 ▶ 54:27
Opinion
Alidoust: Protein AI is past GPT-3 while cell models are at GPT-1
“I think in the protein, protein models, we are past GPT-III. when it comes to single cell models and virtual cell models, yeah, I think GPT-I to two right now. I think we're closer to GPT-I than two.”
Nima Alidoust Feb 25, 2025 ▶ 55:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.