May 27, 2026 · 1h 10m · latent-space

🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub

Alex Rives · 46m spoken RJ Haneke · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Alex Rives, Head of Science at the Chan Zuckerberg Biohub, joins the Latent Space podcast to discuss ESM-Cambrian (ESMC), explaining how scaling transformer foundation models on massive metagenomic datasets turns evolutionary biology into a programmable world model capable of de novo antibody design, structural atlas generation, and virtual cell simulation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.3 Guest teaching 4.2 Guest disagreement 1.1 The hosts pushing back 2.6
05100:0015:0030:0045:001:00:001:27–5:58 · The hosts as informed peer 5/10 Evolutionary Constraints and Foundations of Protein Language Models Brandon challenges the premise that natural language transformer insights necessarily translate to protein sequences, noting sampling at infinite temperature produces valid proteins. Alex explains how evolutionary co-variation constraints provide the fundamental basis for masked language models.5:59–10:21 · The hosts as informed peer 4/10 Introducing ESM-Cambrian and the 1.1-Billion-Structure Atlas Alex introduces ESMC and the 1.1 billion predicted structure atlas. RJ asks clarifying questions regarding how clustering 6.8 billion non-redundant sequences at 70% sequence identity ensures complete structural coverage.10:22–16:02 · The hosts as informed peer 5/10 Mechanistic Interpretability and Linguistic Parallels in Protein Representations RJ summarizes mechanistic interpretability concepts and asks why disjoint sequences share internal feature activations. Alex educates the hosts using Zellig Harris's 1954 distributional hypothesis and the concept of latent compression.16:02–25:48 · The hosts as informed peer 6/10 Metagenomics and Overcoming Data Plateaus in Scaling Laws Brandon and RJ dig into the transition from UniRef to noisy metagenomic sequencing from extreme environments. Alex explains that metagenomics broke the plateau of diminishing returns observed in ESM-2.25:48–31:56 · The hosts as informed peer 7/10 De Novo Antibody and Binder Design Through World Models Brandon presses Alex on whether ESM-3's incorporation of structural priors was a detour from scaling, and displays deep domain knowledge regarding antibody diversity versus standard MSA evolutionary constraints.31:56–36:04 · The hosts as informed peer 5/10 Accelerating Scientific Discovery and Multimer Prediction with ESM RJ highlights the significance of matching AlphaFold 3 capabilities without relying on multiple sequence alignments, while Alex discusses identifying uncharacterized gene editing systems in the atlas.36:04–42:14 · The hosts as informed peer 5/10 Mapping the Interactome and Building Active Experimental Feedback Loops RJ synthesizes Alex's points about cryo-electron tomography into a closed-loop active learning framework. Alex details how digital reasoning oracles will reduce vast hypothesis spaces to targeted empirical tests.42:15–47:19 · The hosts as informed peer 4/10 Biohub's Vision: The Virtual Biology Initiative and Open Philanthropy Brandon links the conversation back to the Chan Zuckerberg Biohub vision from previous episodes. Alex articulates the mission of building open-source foundational measurement tools across cellular scales.47:19–1:04:07 · The hosts as informed peer 7/10 Scaling Cellular Biology: Perturbation, Spatial Mapping, and Cross-Modality Brandon provides detailed context on the $13B historical cost of the PDB and challenges the utility of static structures versus dynamic interactions. Alex explains the information-theoretic paradigm for cellular biology and breaks down the $500M Virtual Biology Initiative.1:04:07–1:10:04 · The hosts as informed peer 5/10 Data Bottlenecks, Fine Sequence Variations, and Open-Source Release Alex rejects Brandon's characterization of large sequence databases as redundant, explaining that fine sequence variations and point mutations are essential for learning protein function rather than just coarse structural topology.1:27–5:58 · Guest teaching 4/10 Evolutionary Constraints and Foundations of Protein Language Models Brandon challenges the premise that natural language transformer insights necessarily translate to protein sequences, noting sampling at infinite temperature produces valid proteins. Alex explains how evolutionary co-variation constraints provide the fundamental basis for masked language models.5:59–10:21 · Guest teaching 3/10 Introducing ESM-Cambrian and the 1.1-Billion-Structure Atlas Alex introduces ESMC and the 1.1 billion predicted structure atlas. RJ asks clarifying questions regarding how clustering 6.8 billion non-redundant sequences at 70% sequence identity ensures complete structural coverage.10:22–16:02 · Guest teaching 5/10 Mechanistic Interpretability and Linguistic Parallels in Protein Representations RJ summarizes mechanistic interpretability concepts and asks why disjoint sequences share internal feature activations. Alex educates the hosts using Zellig Harris's 1954 distributional hypothesis and the concept of latent compression.16:02–25:48 · Guest teaching 5/10 Metagenomics and Overcoming Data Plateaus in Scaling Laws Brandon and RJ dig into the transition from UniRef to noisy metagenomic sequencing from extreme environments. Alex explains that metagenomics broke the plateau of diminishing returns observed in ESM-2.25:48–31:56 · Guest teaching 4/10 De Novo Antibody and Binder Design Through World Models Brandon presses Alex on whether ESM-3's incorporation of structural priors was a detour from scaling, and displays deep domain knowledge regarding antibody diversity versus standard MSA evolutionary constraints.31:56–36:04 · Guest teaching 3/10 Accelerating Scientific Discovery and Multimer Prediction with ESM RJ highlights the significance of matching AlphaFold 3 capabilities without relying on multiple sequence alignments, while Alex discusses identifying uncharacterized gene editing systems in the atlas.36:04–42:14 · Guest teaching 4/10 Mapping the Interactome and Building Active Experimental Feedback Loops RJ synthesizes Alex's points about cryo-electron tomography into a closed-loop active learning framework. Alex details how digital reasoning oracles will reduce vast hypothesis spaces to targeted empirical tests.42:15–47:19 · Guest teaching 3/10 Biohub's Vision: The Virtual Biology Initiative and Open Philanthropy Brandon links the conversation back to the Chan Zuckerberg Biohub vision from previous episodes. Alex articulates the mission of building open-source foundational measurement tools across cellular scales.47:19–1:04:07 · Guest teaching 5/10 Scaling Cellular Biology: Perturbation, Spatial Mapping, and Cross-Modality Brandon provides detailed context on the $13B historical cost of the PDB and challenges the utility of static structures versus dynamic interactions. Alex explains the information-theoretic paradigm for cellular biology and breaks down the $500M Virtual Biology Initiative.1:04:07–1:10:04 · Guest teaching 6/10 Data Bottlenecks, Fine Sequence Variations, and Open-Source Release Alex rejects Brandon's characterization of large sequence databases as redundant, explaining that fine sequence variations and point mutations are essential for learning protein function rather than just coarse structural topology.1:27–5:58 · Guest disagreement 1/10 Evolutionary Constraints and Foundations of Protein Language Models Brandon challenges the premise that natural language transformer insights necessarily translate to protein sequences, noting sampling at infinite temperature produces valid proteins. Alex explains how evolutionary co-variation constraints provide the fundamental basis for masked language models.5:59–10:21 · Guest disagreement 1/10 Introducing ESM-Cambrian and the 1.1-Billion-Structure Atlas Alex introduces ESMC and the 1.1 billion predicted structure atlas. RJ asks clarifying questions regarding how clustering 6.8 billion non-redundant sequences at 70% sequence identity ensures complete structural coverage.10:22–16:02 · Guest disagreement 1/10 Mechanistic Interpretability and Linguistic Parallels in Protein Representations RJ summarizes mechanistic interpretability concepts and asks why disjoint sequences share internal feature activations. Alex educates the hosts using Zellig Harris's 1954 distributional hypothesis and the concept of latent compression.16:02–25:48 · Guest disagreement 1/10 Metagenomics and Overcoming Data Plateaus in Scaling Laws Brandon and RJ dig into the transition from UniRef to noisy metagenomic sequencing from extreme environments. Alex explains that metagenomics broke the plateau of diminishing returns observed in ESM-2.25:48–31:56 · Guest disagreement 2/10 De Novo Antibody and Binder Design Through World Models Brandon presses Alex on whether ESM-3's incorporation of structural priors was a detour from scaling, and displays deep domain knowledge regarding antibody diversity versus standard MSA evolutionary constraints.31:56–36:04 · Guest disagreement 1/10 Accelerating Scientific Discovery and Multimer Prediction with ESM RJ highlights the significance of matching AlphaFold 3 capabilities without relying on multiple sequence alignments, while Alex discusses identifying uncharacterized gene editing systems in the atlas.36:04–42:14 · Guest disagreement 0/10 Mapping the Interactome and Building Active Experimental Feedback Loops RJ synthesizes Alex's points about cryo-electron tomography into a closed-loop active learning framework. Alex details how digital reasoning oracles will reduce vast hypothesis spaces to targeted empirical tests.42:15–47:19 · Guest disagreement 0/10 Biohub's Vision: The Virtual Biology Initiative and Open Philanthropy Brandon links the conversation back to the Chan Zuckerberg Biohub vision from previous episodes. Alex articulates the mission of building open-source foundational measurement tools across cellular scales.47:19–1:04:07 · Guest disagreement 1/10 Scaling Cellular Biology: Perturbation, Spatial Mapping, and Cross-Modality Brandon provides detailed context on the $13B historical cost of the PDB and challenges the utility of static structures versus dynamic interactions. Alex explains the information-theoretic paradigm for cellular biology and breaks down the $500M Virtual Biology Initiative.1:04:07–1:10:04 · Guest disagreement 3/10 Data Bottlenecks, Fine Sequence Variations, and Open-Source Release Alex rejects Brandon's characterization of large sequence databases as redundant, explaining that fine sequence variations and point mutations are essential for learning protein function rather than just coarse structural topology.1:27–5:58 · The hosts pushing back 3/10 Evolutionary Constraints and Foundations of Protein Language Models Brandon challenges the premise that natural language transformer insights necessarily translate to protein sequences, noting sampling at infinite temperature produces valid proteins. Alex explains how evolutionary co-variation constraints provide the fundamental basis for masked language models.5:59–10:21 · The hosts pushing back 2/10 Introducing ESM-Cambrian and the 1.1-Billion-Structure Atlas Alex introduces ESMC and the 1.1 billion predicted structure atlas. RJ asks clarifying questions regarding how clustering 6.8 billion non-redundant sequences at 70% sequence identity ensures complete structural coverage.10:22–16:02 · The hosts pushing back 1/10 Mechanistic Interpretability and Linguistic Parallels in Protein Representations RJ summarizes mechanistic interpretability concepts and asks why disjoint sequences share internal feature activations. Alex educates the hosts using Zellig Harris's 1954 distributional hypothesis and the concept of latent compression.16:02–25:48 · The hosts pushing back 3/10 Metagenomics and Overcoming Data Plateaus in Scaling Laws Brandon and RJ dig into the transition from UniRef to noisy metagenomic sequencing from extreme environments. Alex explains that metagenomics broke the plateau of diminishing returns observed in ESM-2.25:48–31:56 · The hosts pushing back 5/10 De Novo Antibody and Binder Design Through World Models Brandon presses Alex on whether ESM-3's incorporation of structural priors was a detour from scaling, and displays deep domain knowledge regarding antibody diversity versus standard MSA evolutionary constraints.31:56–36:04 · The hosts pushing back 2/10 Accelerating Scientific Discovery and Multimer Prediction with ESM RJ highlights the significance of matching AlphaFold 3 capabilities without relying on multiple sequence alignments, while Alex discusses identifying uncharacterized gene editing systems in the atlas.36:04–42:14 · The hosts pushing back 1/10 Mapping the Interactome and Building Active Experimental Feedback Loops RJ synthesizes Alex's points about cryo-electron tomography into a closed-loop active learning framework. Alex details how digital reasoning oracles will reduce vast hypothesis spaces to targeted empirical tests.42:15–47:19 · The hosts pushing back 1/10 Biohub's Vision: The Virtual Biology Initiative and Open Philanthropy Brandon links the conversation back to the Chan Zuckerberg Biohub vision from previous episodes. Alex articulates the mission of building open-source foundational measurement tools across cellular scales.47:19–1:04:07 · The hosts pushing back 4/10 Scaling Cellular Biology: Perturbation, Spatial Mapping, and Cross-Modality Brandon provides detailed context on the $13B historical cost of the PDB and challenges the utility of static structures versus dynamic interactions. Alex explains the information-theoretic paradigm for cellular biology and breaks down the $500M Virtual Biology Initiative.1:04:07–1:10:04 · The hosts pushing back 4/10 Data Bottlenecks, Fine Sequence Variations, and Open-Source Release Alex rejects Brandon's characterization of large sequence databases as redundant, explaining that fine sequence variations and point mutations are essential for learning protein function rather than just coarse structural topology.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:06:27 Refuting sequence redundancy assumption

Alex firmly counters Brandon's premise that vast sequence databases are redundant, emphasizing that nuanced single-point mutations dictate protein function.

Hardest push from the hosts ▶ 25:48 Pushing on ESM-3 inductive biases

Brandon directly challenges Alex on whether ESM-3 was a detour into hand-crafted priors before returning to pure unconstrained scaling with ESMC.

Biggest teaching moment ▶ 1:06:27 Fine sequence variations vs structural diversity

Alex educates the hosts on how coarse sequence diversity informs structural topology while near-identical fine variations govern biochemical function.

The host holds their own ▶ 30:00 Analyzing antibody design constraints

Brandon demonstrates advanced biological expertise by explaining why antibodies defy standard MSA evolutionary assumptions due to pressure for diversity.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Evolutionary Constraints and Foundations of Protein Language Models 5413 Brandon challenges the premise that natural language transformer insights necessarily translate to protein sequences, noting sampling at infinite temperature produces valid proteins. Alex explains how evolutionary co-variation constraints provide the fundamental basis for masked language models.
Introducing ESM-Cambrian and the 1.1-Billion-Structure Atlas 4312 Alex introduces ESMC and the 1.1 billion predicted structure atlas. RJ asks clarifying questions regarding how clustering 6.8 billion non-redundant sequences at 70% sequence identity ensures complete structural coverage.
Mechanistic Interpretability and Linguistic Parallels in Protein Representations 5511 RJ summarizes mechanistic interpretability concepts and asks why disjoint sequences share internal feature activations. Alex educates the hosts using Zellig Harris's 1954 distributional hypothesis and the concept of latent compression.
Metagenomics and Overcoming Data Plateaus in Scaling Laws 6513 Brandon and RJ dig into the transition from UniRef to noisy metagenomic sequencing from extreme environments. Alex explains that metagenomics broke the plateau of diminishing returns observed in ESM-2.
De Novo Antibody and Binder Design Through World Models 7425 Brandon presses Alex on whether ESM-3's incorporation of structural priors was a detour from scaling, and displays deep domain knowledge regarding antibody diversity versus standard MSA evolutionary constraints.
Accelerating Scientific Discovery and Multimer Prediction with ESM 5312 RJ highlights the significance of matching AlphaFold 3 capabilities without relying on multiple sequence alignments, while Alex discusses identifying uncharacterized gene editing systems in the atlas.
Mapping the Interactome and Building Active Experimental Feedback Loops 5401 RJ synthesizes Alex's points about cryo-electron tomography into a closed-loop active learning framework. Alex details how digital reasoning oracles will reduce vast hypothesis spaces to targeted empirical tests.
Biohub's Vision: The Virtual Biology Initiative and Open Philanthropy 4301 Brandon links the conversation back to the Chan Zuckerberg Biohub vision from previous episodes. Alex articulates the mission of building open-source foundational measurement tools across cellular scales.
Scaling Cellular Biology: Perturbation, Spatial Mapping, and Cross-Modality 7514 Brandon provides detailed context on the $13B historical cost of the PDB and challenges the utility of static structures versus dynamic interactions. Alex explains the information-theoretic paradigm for cellular biology and breaks down the $500M Virtual Biology Initiative.
Data Bottlenecks, Fine Sequence Variations, and Open-Source Release 5634 Alex rejects Brandon's characterization of large sequence databases as redundant, explaining that fine sequence variations and point mutations are essential for learning protein function rather than just coarse structural topology.

Statements from this episode (22)

Opinion
Rives states he believes in scaling laws for biology
“Well, I'll take that. I believe in scaling laws, so”
Alex Rives May 27, 2026 ▶ 1:24
Assertion Supported
Rives: Meta FAIR trained the first protein transformer language model
“And so my team, when we were at MetaFair, trained really the first transformer language model for protein biology.”
Alex Rives May 27, 2026 ▶ 1:34
Assertion Not checkable as stated
Scaling protein models by orders of magnitude unlocks emergent biological capabilities
“So our team has really explored that idea over a number of different years, and we've really kind of, I think, seen the scaling curve and really seen as we have increased models by an order of magnitude kind of in each generation that, you know, there's this e…”
Alex Rives May 27, 2026 ▶ 1:54
Assertion Supported
Biohub predicted 3D structures for 1.1 billion proteins from global sequence databases
“So we put together kind of all the world's largest protein sequence databases. And so that kind of amounts to 6.8 billion non-redundant proteins, and then we've resolved predicted structures for 1.1 billion of those, and we've also computed features across all…”
Alex Rives May 27, 2026 ▶ 7:41
Assertion Supported
Rives: The ESMC family includes 300M, 600M, and 6B parameter models
“So there's actually three models in that family. There's a three hundred million parameter model, a six hundred million. Parameter model and a six billion parameter model.”
Alex Rives May 27, 2026 ▶ 10:31
Assertion Supported
Sparse autoencoders found a single learned feature for nucleophilic elbows in ESMC
“You know, what we found basically is that the model has a kind of a single feature for this nucleophilic elbow and is activating across these like very evolutionarily diverse families, you know, really completely different structural topologies, proteins that …”
Alex Rives May 27, 2026 ▶ 12:59
Insight
Rives: Amino acid sequence patterns allow protein language models to learn biology
“The contexts in which an amino acid can occur are really determined by, you know, the structure, the function of the protein, its biological roles, you know, these I mean, very complex phenomenon both the intrinsic biology of the protein and its relation to al…”
Alex Rives May 27, 2026 ▶ 15:24
Assertion Supported
Rives: ESMC Added Billions of Metagenomic Sequences Beyond UniRef
“ESM-II is trained on Uniref. And for ESMC, we added metagenomics. So we added billions more sequences to the training data.”
Alex Rives May 27, 2026 ▶ 19:31
Assertion Supported
Incorporating metagenomic training data eliminated diminishing returns in ESMC protein foundation models
“And then, you know, what we saw basically is, is, is there are no longer diminishing returns to scale. So that's really saying that ESM two was kind of data limited rather than compute limited for ESMC.”
Alex Rives May 27, 2026 ▶ 24:23
Assertion Supported
Rives: ESMC Has Been Used to Successfully Design scFv Antibodies
“We've been able to use this to actually now go and design many protein binders, but I think sort of most excitingly, we've been able to use this to actually design antibodies, SCFVs, and we're seeing really, I think, exciting success rates and a small number o…”
Alex Rives May 27, 2026 ▶ 28:26
Assertion Supported
ESMC model search generates novel antibodies achieving therapeutic-grade binding affinity levels
“What we're able to see is that, you know, you can search ESMC and you can actually find antibodies that are reaching the level of affinity that are, I should say, are really at the level of affinity that is needed for therapeutic function and activity.”
Alex Rives May 27, 2026 ▶ 29:46
Assertion Supported
Feng Zhang's laboratory used the ESM atlas to discover novel gene editors
“Actually, the first version of the ESM atlas was used by Fang Zhang's group to find A new gene editing system.”
Alex Rives May 27, 2026 ▶ 34:39
Assertion Supported
Rives: ESMC is state of the art among open models for multimer prediction
“Yeah, I mean, I think we're state of the art for open models.”
Alex Rives May 27, 2026 ▶ 36:02
Assertion Supported
Rives: ESMFold 2 yields atomic resolution predictions in seconds without MSAs
“So the other thing about ESM fold two is a really fast model because it doesn't require the multiple sequence alignment. So You know, you can do inference kind of, you know, directly from the sequence it takes seconds, you know, you can get an atomic resolutio…”
Alex Rives May 27, 2026 ▶ 36:26
Disclosure
Biohub is building high-contrast cryo-electron tomography for atomic cell imaging
“At Biohub, I mean, the other thing that we're thinking about is can we actually experimentally resolve this? And so one of the things that we are building is cryo-electron tomography, and we're really building systems that can greatly increase the contrast Whe…”
Alex Rives May 27, 2026 ▶ 36:45
Prediction Not checkable as stated
AI cell simulations will enable parallel reasoning across millions of biological hypotheses
“We're gonna have increasingly capable and accurate digital representations of molecules, genomes, cells, ultimately physiology. That's where you want to get. We're gonna have to go up that, that complexity scale, the levels of biological complexity that requir…”
Alex Rives May 27, 2026 ▶ 39:49
Insight
Rives: Curing disease requires personalized computational models, not conventional pills
“What is, you know, what is the cure to disease look like, right? It's not a pill, right? It's not a medicine in the conventional sense. You know, it's going to have to be a system that is capable of modeling and understanding, you know, the underlying physiolo…”
Alex Rives May 27, 2026 ▶ 44:42
Opinion
Rives: Current Virtual Cell Models Cannot Predict Novel Interventions
“I think with, you know, kind of the current generation of models that are being called virtual cells, they are good representations of the underlying data, but, you know, they have a very limited ability to predict what will happen when you make a novel interv…”
Alex Rives May 27, 2026 ▶ 49:16
Disclosure
CZ Biohub commits $500M to scale biological data creation and technology development
“We announced a few weeks ago, the virtual biology initiative we basically said, you know, we're gonna invest four hundred million internally in data creation and development of technology to scale data generation to be able to increase the number of modalities…”
Alex Rives May 27, 2026 ▶ 57:03
Prediction Not checkable as stated
Rives: Existing Tech Can Scale Biological Data 10x–100x Before Hitting Limits
“I think with current technology, you know, we can definitely kind of get, get data, 10 X to a hundred X where it is today with like relatively reasonable investments, you know, but then to get another 10 X or more in there, that's going to require, require a l…”
Alex Rives May 27, 2026 ▶ 1:03:19
Insight
Rives: Sequence diversity teaches models structure while small variations teach function
“Having a vast diversity of sequences across a wide range of protein families is, you know, really critical for the emergence of this kind of structure Prediction capability, because I think kind of large diversity is what trains the model to understand, to dev…”
Alex Rives May 27, 2026 ▶ 1:06:48
Disclosure
CZ Biohub is open-sourcing the ESMC protein foundation model under MIT license
“At the time that this podcast comes out, we will have announced ESMC and this world model for protein biology. It's gonna be open source. It's gonna be MIT licensed, and we want people to use it.”
Alex Rives May 27, 2026 ▶ 1:09:10
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.