Oct 3, 2024 · 28m · a16z
AI at the Intersection of Bio | Vijay Pande, Surya Ganguli & Bowen Liu
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the 'Raising Health' podcast, Dr. Vijay Pande hosts AI experts Dr. Surya Ganguli and Dr. Bowen Liu to explore how deep learning, self-supervised foundation models, and generative AI are transforming drug discovery, target identification, and clinical development.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Surya explicitly frames his argument as controversial, claiming that deep learning models like AlphaFold 3 fail at out-of-distribution generalization and that simple physics-based docking outperforms them on rare ligands.
Hardest push from the host ▶ 9:20 Host challenges the practical value of protein structure predictionThe host interrupts the guest's overview to directly challenge the utility of structure prediction, demanding to know what specific value it adds to actual drug discovery.
Biggest teaching moment ▶ 14:54 Surya breaks down AlphaFold 3 failure modes with empirical dataSurya educates the host on model generalization limits by citing data from Inductive Bio, demonstrating that physics docking beat AlphaFold 3 by 8 percent on ligands outside the top 50 common data bank examples.
The host holds their own ▶ 16:28 Host uses classical physics to reframe extrapolation vs interpolationThe host draws on his background as a physicist to challenge Surya's definition of extrapolation, showing how Newton's jump from falling apples to orbiting planets was actually interpolation within the correct underlying latent space.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| 'Aha Moments' in AI and Biological Sciences | 4 | 4 | 1 | 1 | The host opens the episode asking about 'aha moments' and framing the shift from traditional computational chemistry to deep learning. He offers a helpful analogy comparing representations to Arabic vs. Roman numerals, while guests detail the evolution from physics-based models to self-supervised learning. | |
| Data Scale and Self-Supervised Learning Across Modalities | 4 | 5 | 1 | 2 | Surya provides a deep quantitative breakdown of dataset sizes across language, proteins, 3D structures, and chemical spaces. The host interjects briefly to add nuance about sequencing bias and deep learning architectures acting as complex physics models. | |
| Solving Labeled Data Scarcity in Drug Discovery | 6 | 5 | 1 | 6 | The host demonstrates strong knowledge of the drug discovery pipeline and explicitly pushes back on Bowen to clarify what protein structure prediction actually accomplishes for practical drug design. The guests respond by explaining binding mechanics, multi-objective optimization, and the economic burden of Eroom's law. | |
| Generative AI and Inverse Molecular Design | 6 | 5 | 2 | 6 | When Bowen notes that validating generated ideas in science is hard, the host explicitly pushes back by citing in silico AUC benchmarks. Bowen clarifies that wet lab synthesis remains the true bottleneck, leading the host to discuss target hit rates and experimentalist trust dynamics. | |
| Foundation Models vs. Specialized Models and Generalization Limits | 7 | 6 | 3 | 6 | Surya takes a contrarian stance that specialized physics models beat foundation models like AlphaFold 3 on out-of-distribution targets. The host counters Surya's framing of extrapolation by using Newton's laws of motion to demonstrate how finding the right latent space turns apparent extrapolation into interpolation. | |
| AI in Target Discovery and Disentangled Cellular Latent Spaces | 6 | 5 | 1 | 4 | The host probes the biological target discovery space, asking critical questions about cellular phenotypes and interpreting latent spaces versus physics formulas. Surya explains disentangled latent spaces using variational autoencoders and face generation analogies. | |
| Revolutionizing Clinical Trials and Countering Eroom's Law | 5 | 4 | 1 | 2 | The discussion turns to clinical trials, where the host calculates that shifting trial success from 20% to 30% represents a dramatic 50% increase in drug output. Surya details patient selection via EMR databases and strategies to counter Eroom's law. | |
| The 'Digital Human' Foundation Model and Future Outlook | 5 | 4 | 1 | 1 | The host synthesizes the discussion into a vision of an end-to-end 'digital human' foundation model over a 10-year timeline. The guests agree and outline the multi-modal biological hierarchy data currently being gathered to build it. |