Prot-SCL: State Of The Art Prediction Of Protein Subcellular Localization From Primary Sequence Using Contrastive Learning
Giannakoulias, S.; Ferrie, J. J.; Apicello, A.; Mitchell, C.
Show abstract
Protein subcellular localization is a critically important parameter to consider when designing expression constructs and production strategies for industry scale protein production. In this study, we present Prot-SCL an innovative self-supervised machine learning approach to predict protein subcellular localization exclusively from primary sequence. The models herein were learned from a dataset of subcellular localizations derived by exhaustively analyzing the Uniprot database. The set of localization data was rigorously curated for machine learning by employing group sampling following clustering of the protein sequences. The novel component of this approach lies in the development of a triplet neural network architecture capable of generating meaningful embeddings for classification of protein subcellular localization. We observed a robust predictive power for our classical gradient boosted machine learning models trained on these triplet embeddings in both cross validation and in generalization to the testing set. Importantly, we have made this extensive dataset of protein subcellular localizations publicly accessible, facilitating future, need-based, localization studies. Finally, we provide the relevant codebase to encourage a wider adoption and expansion of this methodology. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=123 SRC="FIGDIR/small/555932v1_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@1d9b205org.highwire.dtl.DTLVardef@1368c23org.highwire.dtl.DTLVardef@2a88edorg.highwire.dtl.DTLVardef@8380b7_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting the subcellular location of prokaryotic proteins with DeepLocPro 94%
- CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models 93%
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 93%
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 93%
- Positional SHAP (PoSHAP) for Interpretation of Machine Learning Models Trained from Biological Sequences 93%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 92%
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 94%
- DeepSS2GO: protein function prediction from secondary structure 93%
- A machine learning-based approach to identify reliable gold standards for protein complex composition prediction 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.