Deep learning models for unbiased sequence-based PPI prediction plateau at an accuracy of 0.65
Reim, T.; Hartebrodt, A.; Blumenthal, D. B.; Bernett, J.; List, M.
Show abstract
As most proteins interact with other proteins to perform their respective functions, methods to computationally predict these interactions have been developed. However, flawed evaluation schemes and data leakage in test sets have obscured the fact that sequence-based protein-protein interaction (PPI) prediction is still an open problem. Recently, methods achieving better-than-random performance on leakage-free PPI data have been proposed. Here, we show that the use of ESM-2 protein embeddings explains this performance gain irrespective of model architecture. We compared the performance of models with varying complexity, per-protein, and per-token embeddings, as well as the influence of self- or cross-attention, where all models plateaued at an accuracy of 0.65. Moreover, we show that the tested sequence-based models cannot implicitly learn a contact map as an intermediate layer. These results imply that other input types, such as structure, might be necessary for producing reliable PPI predictions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 98%
- ProteinBERT: A universal deep-learning model of protein sequence and function 98%
- Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.