Rives: Amino acid sequence patterns allow protein language models to learn biology
Alex Rives · 🔬 The Bitter Lesson is Coming for Proteins - Alex Rives, BioHub · May 27, 2026 · at 15:24
Alex Rives, Head of Science at Chan Zuckerberg Biohub, adapts Zellig Harris's distributional structure theory of linguistics to explain why protein language models learn biological meaning.
“The contexts in which an amino acid can occur are really determined by, you know, the structure, the function of the protein, its biological roles, you know, these I mean, very complex phenomenon both the intrinsic biology of the protein and its relation to all of the other proteins and the function and evolution. And so, but those are what determine the context sets. And so you would imagine that then those statistical patterns in the use of amino acids, they directly reflect those underlying hidden variables. And so the model is going to learn something about those hidden variables.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →