Back

Protein Language Models Capture Structural and Functional Epistasis in a Zero-Shot Setting

Nambiar, A.; Littlefield, S. B.; Cuellar, C.; Khorana, R.; Maslov, S.

2025-09-17 bioinformatics
10.1101/2025.09.14.676130 bioRxiv
Show abstract

Protein language models (PLMs) learn from large collections of natural sequences and achieve striking success across prediction tasks, yet it remains unclear what biological principles underlie their representations. We use epistasis, the dependence of a mutations effect on its sequence context, as a lens to probe what PLMs capture about proteins. Comparing PLM-derived scores with deep mutational scanning data, we find that epistasis emerges naturally from pretrained models, without supervision on experimental fitness. Raw model scores align with residue-residue contacts, indicating that PLMs internalize structural proximity. Applying a nonlinear transformation to bring model outputs onto the experimental scale, however, shifts the signal toward functional couplings between distant sites. These findings show that PLMs capture both structural and functional dependencies from sequence data alone, and that epistasis provides a powerful window into the biological principles embedded in their representations.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.