Feb 12, 2026 · 1h 41m · latent-space

🔬Generating Molecules, Not Just Models

Gabriele Corso · 48m spoken Jeremy Wohlwend · 31m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

MIT PhD graduates and Boltz co-founders Gabriele Corso and Jeremy Wohlwend discuss the architectural evolution, open-source democratization, and wet-lab validation of generative biological foundation models. They explain how Boltz-1 and BoltzGen leverage full-atom diffusion and scalable infrastructure to empower researchers and accelerate de novo therapeutic drug discovery.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.1 Guest teaching 4.1 Guest disagreement 0.5 The hosts pushing back 0.5
05100:0020:0040:001:00:001:20:001:40:001:06–4:10 · The hosts as informed peer 4/10 The AlphaFold 2 Watershed Moment in Biology The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML.4:11–6:44 · The hosts as informed peer 2/10 Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity.6:45–12:21 · The hosts as informed peer 5/10 Benchmarking Generalization at the CASP Competitions When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation.12:22–15:40 · The hosts as informed peer 3/10 Biological Machinery: Conformations, Misfolding, and Disordered Proteins The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins.15:40–20:03 · The hosts as informed peer 5/10 The Complexity of Protein Folding and Energy Landscapes The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes.20:03–24:27 · The hosts as informed peer 4/10 Pairwise Representations and Dijkstra-Like Geometric Decoding Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions.24:27–27:24 · The hosts as informed peer 5/10 Architectural Shifts: Generative Diffusion Over Direct Regression Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior.27:25–30:10 · The hosts as informed peer 5/10 Pairwise Inductive Biases and Multi-Scale Representations The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases.30:11–34:39 · The hosts as informed peer 3/10 Visualizing Structural Motifs in the Boltz Lab Interface The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops.34:39–38:43 · The hosts as informed peer 7/10 Compute-Heavy Architectures and Iterative Recycling Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups.38:44–43:38 · The hosts as informed peer 5/10 Developing Open-Source Boltz-1 Under Severe Compute Constraints Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers.43:38–46:19 · The hosts as informed peer 3/10 Rigorous Benchmarking Against AlphaFold 3 Using PDB Data The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures.46:19–50:04 · The hosts as informed peer 5/10 Lessons from DiffDock, DockGen, and Open-Source Feedback The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks.50:04–54:40 · The hosts as informed peer 4/10 Founding Boltz as a Public Benefit Company The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories.54:41–58:23 · The hosts as informed peer 3/10 Nurturing a Self-Sustaining Scientific Community The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models.58:23–1:01:41 · The hosts as informed peer 3/10 Community Innovations, Inference-Time Search, and Ranking The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws.1:01:42–1:08:18 · The hosts as informed peer 5/10 BoltzGen: Unifying Protein Sequence and Structure Generation Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation.1:08:19–1:14:54 · The hosts as informed peer 4/10 Broad Experimental Validation Across Diverse Biological Modalities The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides.1:14:56–1:17:24 · The hosts as informed peer 4/10 The Wet-Lab Protein Expression and Binding Assay Workflow Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host.1:17:26–1:21:45 · The hosts as informed peer 3/10 Laboratory Hit Rates and Generalization on Zero-Homology Targets Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs.1:21:47–1:28:47 · The hosts as informed peer 4/10 Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs.1:28:47–1:33:02 · The hosts as informed peer 4/10 Frontier CRO Benchmarking and Open Therapeutic Candidates The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer.1:33:30–1:36:59 · The hosts as informed peer 4/10 Expanding Beyond Binding: Developability, ADME, and Cellular Context The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell.1:36:59–1:40:16 · The hosts as informed peer 5/10 Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners.1:06–4:10 · Guest teaching 2/10 The AlphaFold 2 Watershed Moment in Biology The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML.4:11–6:44 · Guest teaching 4/10 Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity.6:45–12:21 · Guest teaching 6/10 Benchmarking Generalization at the CASP Competitions When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation.12:22–15:40 · Guest teaching 5/10 Biological Machinery: Conformations, Misfolding, and Disordered Proteins The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins.15:40–20:03 · Guest teaching 5/10 The Complexity of Protein Folding and Energy Landscapes The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes.20:03–24:27 · Guest teaching 5/10 Pairwise Representations and Dijkstra-Like Geometric Decoding Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions.24:27–27:24 · Guest teaching 5/10 Architectural Shifts: Generative Diffusion Over Direct Regression Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior.27:25–30:10 · Guest teaching 4/10 Pairwise Inductive Biases and Multi-Scale Representations The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases.30:11–34:39 · Guest teaching 4/10 Visualizing Structural Motifs in the Boltz Lab Interface The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops.34:39–38:43 · Guest teaching 4/10 Compute-Heavy Architectures and Iterative Recycling Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups.38:44–43:38 · Guest teaching 4/10 Developing Open-Source Boltz-1 Under Severe Compute Constraints Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers.43:38–46:19 · Guest teaching 4/10 Rigorous Benchmarking Against AlphaFold 3 Using PDB Data The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures.46:19–50:04 · Guest teaching 3/10 Lessons from DiffDock, DockGen, and Open-Source Feedback The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks.50:04–54:40 · Guest teaching 4/10 Founding Boltz as a Public Benefit Company The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories.54:41–58:23 · Guest teaching 3/10 Nurturing a Self-Sustaining Scientific Community The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models.58:23–1:01:41 · Guest teaching 4/10 Community Innovations, Inference-Time Search, and Ranking The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws.1:01:42–1:08:18 · Guest teaching 4/10 BoltzGen: Unifying Protein Sequence and Structure Generation Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation.1:08:19–1:14:54 · Guest teaching 4/10 Broad Experimental Validation Across Diverse Biological Modalities The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides.1:14:56–1:17:24 · Guest teaching 3/10 The Wet-Lab Protein Expression and Binding Assay Workflow Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host.1:17:26–1:21:45 · Guest teaching 5/10 Laboratory Hit Rates and Generalization on Zero-Homology Targets Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs.1:21:47–1:28:47 · Guest teaching 4/10 Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs.1:28:47–1:33:02 · Guest teaching 4/10 Frontier CRO Benchmarking and Open Therapeutic Candidates The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer.1:33:30–1:36:59 · Guest teaching 4/10 Expanding Beyond Binding: Developability, ADME, and Cellular Context The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell.1:36:59–1:40:16 · Guest teaching 4/10 Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners.1:06–4:10 · Guest disagreement 1/10 The AlphaFold 2 Watershed Moment in Biology The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML.4:11–6:44 · Guest disagreement 0/10 Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity.6:45–12:21 · Guest disagreement 4/10 Benchmarking Generalization at the CASP Competitions When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation.12:22–15:40 · Guest disagreement 1/10 Biological Machinery: Conformations, Misfolding, and Disordered Proteins The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins.15:40–20:03 · Guest disagreement 1/10 The Complexity of Protein Folding and Energy Landscapes The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes.20:03–24:27 · Guest disagreement 0/10 Pairwise Representations and Dijkstra-Like Geometric Decoding Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions.24:27–27:24 · Guest disagreement 3/10 Architectural Shifts: Generative Diffusion Over Direct Regression Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior.27:25–30:10 · Guest disagreement 0/10 Pairwise Inductive Biases and Multi-Scale Representations The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases.30:11–34:39 · Guest disagreement 0/10 Visualizing Structural Motifs in the Boltz Lab Interface The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops.34:39–38:43 · Guest disagreement 0/10 Compute-Heavy Architectures and Iterative Recycling Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups.38:44–43:38 · Guest disagreement 0/10 Developing Open-Source Boltz-1 Under Severe Compute Constraints Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers.43:38–46:19 · Guest disagreement 0/10 Rigorous Benchmarking Against AlphaFold 3 Using PDB Data The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures.46:19–50:04 · Guest disagreement 0/10 Lessons from DiffDock, DockGen, and Open-Source Feedback The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks.50:04–54:40 · Guest disagreement 1/10 Founding Boltz as a Public Benefit Company The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories.54:41–58:23 · Guest disagreement 0/10 Nurturing a Self-Sustaining Scientific Community The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models.58:23–1:01:41 · Guest disagreement 0/10 Community Innovations, Inference-Time Search, and Ranking The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws.1:01:42–1:08:18 · Guest disagreement 0/10 BoltzGen: Unifying Protein Sequence and Structure Generation Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation.1:08:19–1:14:54 · Guest disagreement 0/10 Broad Experimental Validation Across Diverse Biological Modalities The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides.1:14:56–1:17:24 · Guest disagreement 0/10 The Wet-Lab Protein Expression and Binding Assay Workflow Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host.1:17:26–1:21:45 · Guest disagreement 0/10 Laboratory Hit Rates and Generalization on Zero-Homology Targets Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs.1:21:47–1:28:47 · Guest disagreement 0/10 Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs.1:28:47–1:33:02 · Guest disagreement 1/10 Frontier CRO Benchmarking and Open Therapeutic Candidates The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer.1:33:30–1:36:59 · Guest disagreement 0/10 Expanding Beyond Binding: Developability, ADME, and Cellular Context The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell.1:36:59–1:40:16 · Guest disagreement 1/10 Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners.1:06–4:10 · The hosts pushing back 1/10 The AlphaFold 2 Watershed Moment in Biology The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML.4:11–6:44 · The hosts pushing back 0/10 Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity.6:45–12:21 · The hosts pushing back 2/10 Benchmarking Generalization at the CASP Competitions When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation.12:22–15:40 · The hosts pushing back 0/10 Biological Machinery: Conformations, Misfolding, and Disordered Proteins The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins.15:40–20:03 · The hosts pushing back 1/10 The Complexity of Protein Folding and Energy Landscapes The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes.20:03–24:27 · The hosts pushing back 0/10 Pairwise Representations and Dijkstra-Like Geometric Decoding Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions.24:27–27:24 · The hosts pushing back 2/10 Architectural Shifts: Generative Diffusion Over Direct Regression Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior.27:25–30:10 · The hosts pushing back 0/10 Pairwise Inductive Biases and Multi-Scale Representations The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases.30:11–34:39 · The hosts pushing back 0/10 Visualizing Structural Motifs in the Boltz Lab Interface The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops.34:39–38:43 · The hosts pushing back 1/10 Compute-Heavy Architectures and Iterative Recycling Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups.38:44–43:38 · The hosts pushing back 1/10 Developing Open-Source Boltz-1 Under Severe Compute Constraints Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers.43:38–46:19 · The hosts pushing back 0/10 Rigorous Benchmarking Against AlphaFold 3 Using PDB Data The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures.46:19–50:04 · The hosts pushing back 0/10 Lessons from DiffDock, DockGen, and Open-Source Feedback The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks.50:04–54:40 · The hosts pushing back 1/10 Founding Boltz as a Public Benefit Company The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories.54:41–58:23 · The hosts pushing back 0/10 Nurturing a Self-Sustaining Scientific Community The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models.58:23–1:01:41 · The hosts pushing back 0/10 Community Innovations, Inference-Time Search, and Ranking The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws.1:01:42–1:08:18 · The hosts pushing back 0/10 BoltzGen: Unifying Protein Sequence and Structure Generation Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation.1:08:19–1:14:54 · The hosts pushing back 0/10 Broad Experimental Validation Across Diverse Biological Modalities The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides.1:14:56–1:17:24 · The hosts pushing back 0/10 The Wet-Lab Protein Expression and Binding Assay Workflow Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host.1:17:26–1:21:45 · The hosts pushing back 0/10 Laboratory Hit Rates and Generalization on Zero-Homology Targets Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs.1:21:47–1:28:47 · The hosts pushing back 0/10 Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs.1:28:47–1:33:02 · The hosts pushing back 1/10 Frontier CRO Benchmarking and Open Therapeutic Candidates The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer.1:33:30–1:36:59 · The hosts pushing back 0/10 Expanding Beyond Binding: Developability, ADME, and Cellular Context The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell.1:36:59–1:40:16 · The hosts pushing back 2/10 Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 8:57 Refusing the 'problem solved' framing

Jeremy firmly pushes back on the host's assertion that CASP-14 solved protein folding, pointing out the community's frustration with the term and explaining the limits of evolutionary co-variation.

Hardest push from the hosts ▶ 26:40 Challenging the bitter lesson simplification

The co-host challenges the architectural changes as an instance of Sutton's bitter lesson, prompting Gabriele to counter that specialized geometric architectures remain vastly superior to standard transformers in biology.

Biggest teaching moment ▶ 9:15 Co-evolutionary constraints vs folding mechanics

Jeremy educates the hosts on the difference between structure prediction via evolutionary sequence correlations and actual physical protein folding dynamics.

The host holds their own ▶ 35:58 Analyzing parameter efficiency and equivariant priors

The host demonstrates deep technical proficiency by citing AlphaFold 2's specific ~70M parameter scale, equivariant geometric layers, and how coevolutionary databases act as dynamic parameter priors.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The AlphaFold 2 Watershed Moment in Biology 4211 The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML.
Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids 2400 The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity.
Benchmarking Generalization at the CASP Competitions 5642 When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation.
Biological Machinery: Conformations, Misfolding, and Disordered Proteins 3510 The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins.
The Complexity of Protein Folding and Energy Landscapes 5511 The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes.
Pairwise Representations and Dijkstra-Like Geometric Decoding 4500 Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions.
Architectural Shifts: Generative Diffusion Over Direct Regression 5532 Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior.
Pairwise Inductive Biases and Multi-Scale Representations 5400 The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases.
Visualizing Structural Motifs in the Boltz Lab Interface 3400 The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops.
Compute-Heavy Architectures and Iterative Recycling 7401 Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups.
Developing Open-Source Boltz-1 Under Severe Compute Constraints 5401 Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers.
Rigorous Benchmarking Against AlphaFold 3 Using PDB Data 3400 The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures.
Lessons from DiffDock, DockGen, and Open-Source Feedback 5300 The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks.
Founding Boltz as a Public Benefit Company 4411 The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories.
Nurturing a Self-Sustaining Scientific Community 3300 The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models.
Community Innovations, Inference-Time Search, and Ranking 3400 The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws.
BoltzGen: Unifying Protein Sequence and Structure Generation 5400 Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation.
Broad Experimental Validation Across Diverse Biological Modalities 4400 The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides.
The Wet-Lab Protein Expression and Binding Assay Workflow 4300 Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host.
Laboratory Hit Rates and Generalization on Zero-Homology Targets 3500 Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs.
Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs 4400 Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs.
Frontier CRO Benchmarking and Open Therapeutic Candidates 4411 The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer.
Expanding Beyond Binding: Developability, ADME, and Cellular Context 4400 The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell.
Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows 5412 The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners.

Statements from this episode (24)

Assertion Not checkable as stated
Models Excel at Monomeric Proteins but Struggle on Other Biomolecular Modalities
“And, you know, we've seen in some of the latest GASP competitions, like, while we're becoming really, really good at proteins, especially monomeric proteins you know, other modalities still remain pretty difficult.”
Jeremy Wohlwend Feb 12, 2026 ▶ 8:10
Insight
Structure Prediction Models Degrade Without Co-Evolutionary Sequence Data
“What it implies also is that, you know, in absence of that co-evolutionary landscape, the models don't quite perform as well.”
Jeremy Wohlwend Feb 12, 2026 ▶ 10:24
Insight
AI Excels at Structure Prediction but Fails to Model Physical Folding
“Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state, and that I don't think we've made that much progress on, but the idea of, like, yeah, going straight to th…”
Jeremy Wohlwend Feb 12, 2026 ▶ 10:43
Insight
Current Models Fail to Predict Protein Conformational Dynamics and States
“Proteins are not static. They move, they take different shapes based on their energy states. And I think we are also not that good at understanding the different states that the protein can be in and at what frequency, what probability.”
Jeremy Wohlwend Feb 12, 2026 ▶ 11:28
Opinion
ML Folding Models Only Possess Localized Understanding of Physics
“I think one of the thing, at least I believe is that once you're in that sort of approximate, you know, area of the solution space, then the models have like some understanding, you know, of how to get you to like, you know, the low energy low energy state. An…”
Jeremy Wohlwend Feb 12, 2026 ▶ 19:32
Insight
AlphaFold Resolves Protein Structures Using Dijkstra-Like Pairwise Distance Decoding
“From this evolutionary information about potential contacts, Then it's almost as if the model is sort of running some kind of, you know, Diestro algorithm, where it's sort of decoding, okay, these have to be closed, okay, then if these are closed and this is c…”
Gabriele Corso Feb 12, 2026 ▶ 20:48
Insight
Diffusion Generates Distinct Structures While Regression Averages Them
“What's going to happen in a regression model is that, you know, I'm going to try to make an average of those different kind of answers that I had in mind. When you have a generative model, what you're going to do is, you know, sample all these different answer…”
Gabriele Corso Feb 12, 2026 ▶ 25:44
Opinion
Specialized Equivariant Architectures Vastly Outperform Simple Transformers in Molecular ML
“This field is one of the I would argue very few fields in applied machine learning where we still have kind of architecture. They are Very specialized. And, you know, there are many people that have tried to replace these architectures with, you know, simple t…”
Gabriele Corso Feb 12, 2026 ▶ 26:51
Insight
Pairwise Representation Mechanisms Have Survived Since AlphaFold 2
“And yeah, I think too, it's really survived the test of time. I mean, you know, this thing came out in 20, 21 and it's largely the same. I mean, there's been this change to the structure module that's been like largely simplified, but where a lot of the magic …”
Jeremy Wohlwend Feb 12, 2026 ▶ 28:56
Insight
Alternating Atomic and Token-Level Modeling Enabled AlphaFold 3 Small Molecules
“The other part, I think that's enough for three is sort of this moving away from modeling just at the amino acid level to actually sort of having the model sort of alternate between you know, sort of atomic resolution modeling and then more like it's got token…”
Jeremy Wohlwend Feb 12, 2026 ▶ 29:20
Assertion Supported
AlphaFold Models Have Fewer Parameters but Higher Compute Costs Than LLMs
“They, in terms of parameters, are actually not very big. They are definitely below a billion parameters. You know, if you're here these days in LLM space, you know, a model with less than a billion parameters, you'd think can't do anything. But on the other ha…”
Gabriele Corso Feb 12, 2026 ▶ 35:06
Insight
LLMs Need Massive Parameters for Memorization, Unlike Biology Models
“Part of the reason the LLMs are so large isn't just because of their reasoning capability, but it's also because of, like, the sheer quantity of information that they store. And I think here there's a little bit less of that, you know, and I think it's more ab…”
Jeremy Wohlwend Feb 12, 2026 ▶ 36:47
Assertion Supported
Boltz-1 Was the First Open-Source Model Matching AlphaFold 3 Accuracy
“We went ahead and built Pulse One, which was the first fully open source kind of model to approach the level of accuracy of half-fold-free.”
Gabriele Corso Feb 12, 2026 ▶ 40:10
Assertion Not checkable as stated
Compute Constraints Forced Boltz-1 to Train Once With In-Flight Bug Fixes
“And actually we only trained the big model once. That's how much compute we had. We could only train it once. And so like, while the model was training, we were like finding bugs left and right. A lot of them that I wrote. Yeah. And like, I would, I remember l…”
Jeremy Wohlwend Feb 12, 2026 ▶ 42:10
Assertion Supported
AlphaFold 3 Still Holds an Edge in Antibody-Antigen Prediction
“I would still say that, you know, even to this day, there are, you know, some specific instances where AlphaFold III works better. I think one common example is antibody antigen prediction, where, you know, AlphaFold III still seems to have an edge in, in many…”
Gabriele Corso Feb 12, 2026 ▶ 44:09
Insight
Open-Sourcing Models on GitHub Is Insufficient for Biotech Industry Adoption
“One of the reasons why we realized that Bolts needed to be a company, it couldn't just be an academic project, is that putting a model on GitHub is definitely not enough To get, you know, chemists and biologists, you know, across, you know both academia, biote…”
Gabriele Corso Feb 12, 2026 ▶ 50:41
Assertion Supported
Generative Molecular Models Exhibit Dramatic Inference-Time Compute Scaling
“These days we're seeing kind of pretty dramatic inference time scaling of these models where, you know, the more you run them, the better the results are.”
Gabriele Corso Feb 12, 2026 ▶ 51:59
Insight
Structure Generation Reduces to a Ranking Problem When Sampling Is Scaled
“If you can sample a ton and you assume that, like, you know, if you sample enough, you're likely to have, like, you know, the good structure, then it really just becomes a ranking problem.”
Jeremy Wohlwend Feb 12, 2026 ▶ 1:01:12
Insight
Model Structural Confidence Is a Poor Predictor of Binding Affinity
“Unfortunately, confidence is not a very good predictor of affinity.”
Gabriele Corso Feb 12, 2026 ▶ 1:05:24
Insight
Demand for Smaller Protein Modalities Is Driven by Manufacturing Ease
“There's a general pattern. I think in a, in trying to design things that are smaller, you know, like it's easier to manufacture. At the same time, like that comes with like potentially other challenges, like maybe a little bit less selectivity than like, if yo…”
Jeremy Wohlwend Feb 12, 2026 ▶ 1:14:19
Assertion Not checkable as stated
Most Protein Design Validation Uses Targets With Training Data Overlap
“One of the things that, you know, we found, ah, with the field was that a lot of the validation, especially outside of the validation that was done on specific problems, was done on targets that have a lot of, you know, known interactions in, in the training d…”
Gabriele Corso Feb 12, 2026 ▶ 1:19:49
Assertion Supported
Boltz Achieved Nanomolar Binders on Two-Thirds of Zero-Homology Targets
“And the very cool thing that we saw was that on two-thirds of those targets, we were able to, from these 15 designs get nanomolar binders.”
Gabriele Corso Feb 12, 2026 ▶ 1:21:10
Disclosure
Boltz Will Cede Discovered Molecules Rather Than Develop Therapeutic Drugs
“When we say we have no interest in making Dress, we're serious. Like, you know I mean, when it was with the academic labs, basically the, you know, it was, they keep it, they do whatever they want with it. And with the CRO so far, yeah, we've been very, yeah, …”
Jeremy Wohlwend Feb 12, 2026 ▶ 1:31:47
Insight
AI-Designed Proteins and Molecules Are Not Ready-to-Use Drugs
“When we say that we design new proteins, or we say that we design new molecules, you know you know, go and bind these particular targets. We should be very clear. You know, these are not drugs. You know, these are not things that are ready to be put into a hum…”
Gabriele Corso Feb 12, 2026 ▶ 1:32:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.