A flaw in using pre-trained pLLMs in protein-protein interaction inference models
Szymborski, J.; Emad, A.
Show abstract
With the growing pervasiveness of pre-trained protein large language models (pLLMs), pLLM-based methods are increasingly being put forward for the protein-protein interaction (PPI) inference task. Here, we identify and confirm that existing pre-trained pLLMs are a source of data leakage for the downstream PPI task. We characterize the extent of the data leakage problem by training and comparing small and efficient pLLMs on a dataset that controls for data leakage ("strict") with one that does not ("non-strict"). While data leakage from pre-trained pLLMs cause measurable inflation of testing scores, we find that this does not necessarily extend to other, non-paired biological tasks such as protein keyword annotation. Further, we find no connection between the context-lengths of pLLMs and the performance of pLLM-based PPI inference methods on proteins with sequence lengths that surpass it. Furthermore, we show that pLLM-based and non-pLLM-based models fail to generalize in tasks such as prediction of the human-SARS-CoV-2 PPIs or the effect of point mutations on binding-affinities. This study demonstrates the importance of extending existing protocols for the evaluation of pLLM-based models applied to paired biological datasets and identifies areas of weakness of current pLLM models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 96%
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 96%
Similar papers in this journal
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 97%
- FAPM: Functional Annotation of Proteins using Multi-Modal Models Beyond Structural Modeling 97%
- UDSMProt: Universal Deep Sequence Models for Protein Classification 96%
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 95%
- Predicting compound-protein interaction using hierarchical graph convolutional networks 95%
- ProtAttn-QuadNet: An attention-based deep learning framework for protein-protein interaction prediction using ProtBERT embeddings 95%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 97%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Prop3D: A Flexible, Python-based Platform for Machine Learning with Protein Structural Properties and Biophysical Data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.