Beyond additivity: zero-shot methods cannot predict impact of epistasis on protein properties and function
Kolchina, A.; Dubanevics, I.; Kondrashov, F. A.; Kalinina, O. V.
Show abstract
Accurate prediction of properties and function of mutated proteins is crucial for both research and industrial applications. Experimental assessment of mutations relies on biochemical techniques, which, while accurate, are costly and labour-intensive. As an alternative, computational methods have emerged as a scalable and cost-effective solution. A key challenge for predicting functional consequences of mutations is epistasis, a phenomenon where the effect of one mutation is influenced by others. We evaluated the ability of 95 zero-shot models to predict the impact of epistasis on proteins using datasets from ProteinGym. Our results demonstrate that while the current models perform well for single mutations and non-epistatic combinations of mutations, they fail to predict the effect of strongly epistasic combinations of mutations. This exposes deficiencies of the state-of-the-art models and the need for focusing on capturing complex mutational interactions, which is essential for advancing both evolutionary studies and protein design.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CodonTransformer: a multispecies codon optimizer using context-aware neural networks 95%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 95%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 95%
Similar papers in this journal
- Engineering indel and substitution variants of diverse and ancient enzymes using Graphical Representation of Ancestral Sequence Predictions (GRASP) 95%
- SUITOR: selecting the number of mutational signatures through cross-validation 94%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.