PhosphoLingo: protein language models for phosphorylation site prediction
Zuallaert, J.; Ramasamy, P.; Bouwmeester, R.; Callewaert, N.; Degroeve, S.
Show abstract
Protein post-translational modifications (PTMs) play an important role in numerous biological processes by significantly affecting protein structure and dynamics. Effective computational methods that provide a sequence-based prediction of PTM sites are desirable to guide functional experiments. Whereas these methods typically train neural networks on one-hot encoded amino acid sequences, protein language models carry higher-level pattern information that may improve sequence based prediction performance and hence constitute the current edge of the field. In this study, we first evaluate the training of convolutional neural networks on top of various protein language models for sequence based PTM prediction. Our results show substantial prediction accuracy improvements for various PTMs with current procedures of dataset compilation and model performance evaluation. We then used model interpretation methods to study what these advanced models actually base their learning on. Importantly for the entire field of PTM site predictors trained on proteomics-derived data, our model interpretation and transferability experiments reveal that the current approach to compile training datasets based on proteomics data leads to an artefactual protease-specific training bias that is exploited by the prediction models. This results in an overly optimistic estimation of prediction accuracy, an important caveat in the application of advanced machine learning approaches to PTM prediction based on proteomics data. We suggest a partial solution to reduce this data bias by implementing negative sample filtering, only allowing candidate PTM sites in matched peptides that are present in the experimental metadata. Availability and implementationThe prediction tool, with training and evaluation code, trained models, datasets, and predictions for various PTMs are available at https://github.com/jasperzuallaert/PhosphoLingo. Contactsven.degroeve@vib-ugent.be and nico.callewaert@vib-ugent.be Supplementary informationSupplementary materials are available at bioRxiv.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- In vitro Kinase-to-Phosphosite database (iKiP-DB) predicts kinase activity in phosphoproteomic datasets 97%
- To fly, or not to fly, that is the question: A deep learning model for peptide detectability prediction in mass spectrometry 96%
- Scop3P: a comprehensive resource of human phosphosites within their full context 96%
Similar papers in this journal
- AlphaPeptDeep: A modular deep learning framework to predict peptide properties for proteomics 96%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 96%
- Retention Time Prediction Using Neural Networks Increases Identifications in Crosslinking Mass Spectrometry 96%
Similar papers in this journal
- Sitetack: A Deep Learning Model that Improves PTM Predictionby Using Known PTMs 96%
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 95%
- Pepsickle rapidly and accurately predicts proteasomal cleavage sites for improved neoantigen identification 94%
Similar papers in this journal
- Deep Learning Prediction of Glycopeptide Tandem Mass Spectra Powers Glycoproteomics 94%
- Deep Domain Adversarial Neural Network for the Deconvolution of Cell Type Mixtures in Tissue Proteome Profiling 94%
- Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.