Back

Prot-SCL: State Of The Art Prediction Of Protein Subcellular Localization From Primary Sequence Using Contrastive Learning

Giannakoulias, S.; Ferrie, J. J.; Apicello, A.; Mitchell, C.

2023-09-05 bioinformatics
10.1101/2023.09.01.555932 bioRxiv
Show abstract

Protein subcellular localization is a critically important parameter to consider when designing expression constructs and production strategies for industry scale protein production. In this study, we present Prot-SCL an innovative self-supervised machine learning approach to predict protein subcellular localization exclusively from primary sequence. The models herein were learned from a dataset of subcellular localizations derived by exhaustively analyzing the Uniprot database. The set of localization data was rigorously curated for machine learning by employing group sampling following clustering of the protein sequences. The novel component of this approach lies in the development of a triplet neural network architecture capable of generating meaningful embeddings for classification of protein subcellular localization. We observed a robust predictive power for our classical gradient boosted machine learning models trained on these triplet embeddings in both cross validation and in generalization to the testing set. Importantly, we have made this extensive dataset of protein subcellular localizations publicly accessible, facilitating future, need-based, localization studies. Finally, we provide the relevant codebase to encourage a wider adoption and expansion of this methodology. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=123 SRC="FIGDIR/small/555932v1_ufig1.gif" ALT="Figure 1"> View larger version (52K): org.highwire.dtl.DTLVardef@1d9b205org.highwire.dtl.DTLVardef@1368c23org.highwire.dtl.DTLVardef@2a88edorg.highwire.dtl.DTLVardef@8380b7_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.