Protein Language Models Capture Structural and Functional Epistasis in a Zero-Shot Setting
Nambiar, A.; Littlefield, S. B.; Cuellar, C.; Khorana, R.; Maslov, S.
Show abstract
Protein language models (PLMs) learn from large collections of natural sequences and achieve striking success across prediction tasks, yet it remains unclear what biological principles underlie their representations. We use epistasis, the dependence of a mutations effect on its sequence context, as a lens to probe what PLMs capture about proteins. Comparing PLM-derived scores with deep mutational scanning data, we find that epistasis emerges naturally from pretrained models, without supervision on experimental fitness. Raw model scores align with residue-residue contacts, indicating that PLMs internalize structural proximity. Applying a nonlinear transformation to bring model outputs onto the experimental scale, however, shifts the signal toward functional couplings between distant sites. These findings show that PLMs capture both structural and functional dependencies from sequence data alone, and that epistasis provides a powerful window into the biological principles embedded in their representations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 96%
- Pervasive, conserved secondary structure in highly charged protein regions 95%
Similar papers in this journal
Similar papers in this journal
- The landscape of antibody binding affinity in SARS-CoV-2 Omicron BA.1 evolution 95%
- Epistasis facilitates functional evolution in an ancient transcription factor 95%
- Deep mutational scanning and machine learning reveal structural and molecular rules governing allosteric hotspots in homologous proteins 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.