Feb 25, 2025 · 57m · no-priors
No Priors Ep. 103 | With Vevo Therapeutics and the Arc Institute
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Leaders from Vevo Therapeutics and the Arc Institute introduce the Tahoe-100 dataset, discussing how massive single-cell perturbation data and foundation models will transform drug discovery through predictive virtual cell engineering.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 10.9% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Nima Alidoust forcefully criticizes traditional biotech organizations for announcing multi-year timelines and operating with slow, bloated bureaucracies.
Hardest push from the hosts ▶ 52:21 Sarah confronts the decade-long lack of AI biotech treatmentsSarah presses the panel on why listeners should believe AI will deliver cures now when AI biotechs have pitched big claims for over a decade with few commercial therapies.
Biggest teaching moment ▶ 4:48 Dave Burke's CPU and graphic equalizer transcriptomic modelDave Burke re-educates the conversation by laying out an engineering mental model of the cell, mapping DNA to ROM, RNA expression to dynamic equalizer RAM, and virtual cell AI to the central CPU.
The host holds their own ▶ 22:11 Sarah frames the shift using Sutton's Bitter LessonSarah demonstrates deep domain expertise by connecting the panel's hypothesis-free data scaling strategy to Rich Sutton's foundational Bitter Lesson in computing.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Significance of Tahoe-100 and Single-Cell Perturbation Datasets | 4 | 3 | 1 | 1 | Sarah kicks off the interview with high-level framing about Tahoe-100 and why cellular models differ from protein prediction. The guests collaboratively explain how cellular perturbational datasets represent an ImageNet-style foundational leap for biology. | |
| The Virtual Cell Architecture and Transcriptomic Equalizers | 5 | 4 | 1 | 2 | Sarah prompts the panel on the distinction between protein models and virtual cell models. Dave Burke and Hani Goodarzi provide detailed technical analogies comparing DNA to ROM and RNA expression to an equalizer RAM. | |
| Batch Effects, Causal Inference, and Biological Token Scaling | 5 | 5 | 2 | 2 | Sarah asks for layman explanations of data quality and scaling laws for cell tokens. The panel breaks down batch effects in legacy academic data and explains why 100M single-cell data points equal hundreds of billions of training tokens. | |
| Vevo's Mosaic Platform and Hypothesis-Free Data Generation | 4 | 4 | 1 | 1 | Sarah asks how genetic perturbations are selected across the landscape. Johnny and Nima explain Vevo's mosaic tumor pooling platform and why massive scale enables hypothesis-free, unbiased biological data generation. | |
| Scaling Scientific Intuition and Sutton's Bitter Lesson | 7 | 3 | 2 | 3 | Sarah synthesizes Vevo's mosaic technology in layman terms and connects the guests' data scaling philosophy directly to Sutton's Bitter Lesson. The guests enthusiastically affirm the connection and discuss cross-domain AI generalization. | |
| Open Sourcing the Atlas and Automated Scientific Agents | 4 | 4 | 2 | 1 | Sarah inquires why a private biotech startup would open source its crown-jewel dataset. The panel highlights how open data mobilizes the scientific community while automating dry lab workflows with web-crawling AI agents. | |
| Evaluating Virtual Cells and Reimagining Drug Discovery | 6 | 5 | 2 | 3 | Sarah presses on how model quality is measured and questions the concrete therapeutic outcomes of virtual cell models. The guests honestly address low current DEG predictive baseline accuracy and how in silico models could overturn the 90% clinical failure rate. | |
| Choosing Abstraction Levels from Transcriptomes to Organoids | 6 | 4 | 1 | 2 | Sarah challenges the panel on whether single-cell resolution is sufficient versus multi-cell organoid systems. Dave and Nima defend the transcriptomic level as the optimal abstraction layer that still captures microenvironmental signals. | |
| Platform Biotechs, Global Competition, and Lean Science | 6 | 4 | 3 | 2 | Sarah brings up the commercial trade-offs of platform biotechs and the competitive pressure of Chinese biotechs. Patrick, Johnny, and Nima push back strongly against bloated, bureaucratic legacy pharma operating models. | |
| Overcoming AI Skepticism and Biology's GPT Progression | 7 | 4 | 2 | 3 | Sarah asks the key skeptic question: why should AI in bio work now after a decade of unmet promises? Dave illustrates progress with Evo II's unprompted discovery of BRCA1 pathogenicity, and the guests place current biology between GPT-1 and GPT-2. |