Feb 12, 2026 · 1h 41m · latent-space
🔬Generating Molecules, Not Just Models
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
MIT PhD graduates and Boltz co-founders Gabriele Corso and Jeremy Wohlwend discuss the architectural evolution, open-source democratization, and wet-lab validation of generative biological foundation models. They explain how Boltz-1 and BoltzGen leverage full-atom diffusion and scalable infrastructure to empower researchers and accelerate de novo therapeutic drug discovery.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jeremy firmly pushes back on the host's assertion that CASP-14 solved protein folding, pointing out the community's frustration with the term and explaining the limits of evolutionary co-variation.
Hardest push from the hosts ▶ 26:40 Challenging the bitter lesson simplificationThe co-host challenges the architectural changes as an instance of Sutton's bitter lesson, prompting Gabriele to counter that specialized geometric architectures remain vastly superior to standard transformers in biology.
Biggest teaching moment ▶ 9:15 Co-evolutionary constraints vs folding mechanicsJeremy educates the hosts on the difference between structure prediction via evolutionary sequence correlations and actual physical protein folding dynamics.
The host holds their own ▶ 35:58 Analyzing parameter efficiency and equivariant priorsThe host demonstrates deep technical proficiency by citing AlphaFold 2's specific ~70M parameter scale, equivariant geometric layers, and how coevolutionary databases act as dynamic parameter priors.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| The AlphaFold 2 Watershed Moment in Biology | 4 | 2 | 1 | 1 | The host sets up the context of AlphaFold 2 and asks what made that breakthrough moment so significant. Gabriele walks through structural biology fundamentals and his transition from theoretical ML. | |
| Biomolecular Foundations: Proteins, Small Molecules, and Nucleic Acids | 2 | 4 | 0 | 0 | The co-host asks for basic definitions of proteins, small molecules, and nucleic acids. Jeremy provides a detailed educational breakdown of amino acid sequences and molecular diversity. | |
| Benchmarking Generalization at the CASP Competitions | 5 | 6 | 4 | 2 | When the host asserts CASP-14 comprehensively solved the problem, Jeremy immediately rejects the premise, explaining why the community dislikes calling it solved and elaborating on evolutionary co-variation. | |
| Biological Machinery: Conformations, Misfolding, and Disordered Proteins | 3 | 5 | 1 | 0 | The hosts inquire about intermediate states and folding dynamics. Gabriele corrects the co-host on the ribosome's specific role before describing fold-switching and intrinsically disordered proteins. | |
| The Complexity of Protein Folding and Energy Landscapes | 5 | 5 | 1 | 1 | The co-host references Andrew White and custom ASICs, asking why co-evolution signals physical proximity. Jeremy and Gabriele explain combinatorial NP complexity and how models navigate energy landscapes. | |
| Pairwise Representations and Dijkstra-Like Geometric Decoding | 4 | 5 | 0 | 0 | Gabriele explains the pairwise context representation and compares the decoding mechanism to Dijkstra's algorithm traversing contact matrices before expanding into multimolecular interactions. | |
| Architectural Shifts: Generative Diffusion Over Direct Regression | 5 | 5 | 3 | 2 | Gabriele explains the shift from regression to diffusion-based generative modeling for structural uncertainty. When the co-host suggests this reflects the 'bitter lesson', Gabriele clarifies specialized architectures remain vastly superior. | |
| Pairwise Inductive Biases and Multi-Scale Representations | 5 | 4 | 0 | 0 | The host brings up triangle multiplicative update layers. Jeremy explains second-order pairwise representations, multi-scale residue-to-atom modeling, and strong inductive biases. | |
| Visualizing Structural Motifs in the Boltz Lab Interface | 3 | 4 | 0 | 0 | The co-host asks Jeremy to walk through the 3D visualization on Boltz Lab. Jeremy breaks down secondary structures including alpha helices, beta sheets, and flexible binding loops. | |
| Compute-Heavy Architectures and Iterative Recycling | 7 | 4 | 0 | 1 | Gabriele highlights how biological models have low parameter counts but massive FLOPs. The host demonstrates deep technical domain knowledge, citing AlphaFold 2's ~70M parameters, equivariant layers, and coevolution lookups. | |
| Developing Open-Source Boltz-1 Under Severe Compute Constraints | 5 | 4 | 0 | 1 | Gabriele and Jeremy discuss DeepMind's decision not to open-source AlphaFold 3 and describe training Boltz-1 on a single unrepeatable run with mid-training bug fixes on DOE supercomputers. | |
| Rigorous Benchmarking Against AlphaFold 3 Using PDB Data | 3 | 4 | 0 | 0 | The co-host asks how models are benchmarked against AlphaFold 3. Gabriele details using temporal cutoffs on the Protein Data Bank (PDB) to test zero-shot generalization on novel structures. | |
| Lessons from DiffDock, DockGen, and Open-Source Feedback | 5 | 3 | 0 | 0 | The host cites Gabriele's past papers (DiffDock, DiffDock-L, DockGen). Gabriele details how open-source failure modes and collaboration with Harvard led to rigorous out-of-distribution docking benchmarks. | |
| Founding Boltz as a Public Benefit Company | 4 | 4 | 1 | 1 | The co-host asks how open-source models align with commercial viability. Gabriele explains the rationale for a Public Benefit Company, emphasizing the engineering and compute required beyond raw GitHub repositories. | |
| Nurturing a Self-Sustaining Scientific Community | 3 | 3 | 0 | 0 | The host inquires about community growth. Jeremy describes how the Boltz Slack community became self-sustaining, and Gabriele outlines their commitment to a continuous suite of open-weight models. | |
| Community Innovations, Inference-Time Search, and Ranking | 3 | 4 | 0 | 0 | The guests recount unexpected community contributions, highlighting Tim O'Donnell's residue scanning trick that paved the way for inference-time search, ranking, and scaling laws. | |
| BoltzGen: Unifying Protein Sequence and Structure Generation | 5 | 4 | 0 | 0 | Gabriele explains BoltzGen's unified sequence and structure generation, using atomic coordinate placement to deduce amino acid identity. The host confirms the underlying formulation. | |
| Broad Experimental Validation Across Diverse Biological Modalities | 4 | 4 | 0 | 0 | The co-host highlights the extensive wet-lab validation in the BoltzGen paper. Gabriele and Jeremy detail coordinating with ten external academic and industry labs across diverse targets, nanobodies, and peptides. | |
| The Wet-Lab Protein Expression and Binding Assay Workflow | 4 | 3 | 0 | 0 | Jeremy walks through the wet-lab expression pipeline, from DNA synthesis to yeast expression, tag-based purification, and binding affinity assays, humorously summarized by the co-host. | |
| Laboratory Hit Rates and Generalization on Zero-Homology Targets | 3 | 5 | 0 | 0 | Gabriele details wet-lab success rates on zero-homology targets, achieving nanomolar affinity binders on two-thirds of previously unseen proteins across 15 candidate designs. | |
| Launching Boltz Lab: Design Agents, Scalable Infrastructure, and APIs | 4 | 4 | 0 | 0 | Jeremy introduces Boltz Lab's architecture: specialized protein/small-molecule design agents, massive parallel GPU search infrastructure, and collaborative multi-chemist consensus UIs. | |
| Frontier CRO Benchmarking and Open Therapeutic Candidates | 4 | 4 | 1 | 1 | The host asks what Boltz does when they discover potent binders. Jeremy and Gabriele clarify they open-source candidate hits because they are toolmakers rather than a pharmaceutical drug developer. | |
| Expanding Beyond Binding: Developability, ADME, and Cellular Context | 4 | 4 | 0 | 0 | The co-host asks about expanding into downstream drug properties. Gabriele discusses optimizing developability, ADME profiles, and iterative in-vivo feedback loops without building a full virtual cell. | |
| Overcoming Medicinal Chemist Skepticism with Domain-Driven Workflows | 5 | 4 | 1 | 2 | The host notes medicinal chemists are notoriously skeptical of ML. Gabriele describes integrating an in-house medicinal chemist to co-design workflows, and Jeremy emphasizes that only wet-lab hits ultimately convince practitioners. |