Back

A hybrid approach for predicting multi-label subcellular localization of mRNA at genome scale

Choudhury, S.; Bajiya, N.; Patiyal, S.; Raghava, G. P. S.

2023-01-19 bioinformatics
10.1101/2023.01.17.524365 bioRxiv
Show abstract

In the past, number of methods have been developed for predicting single label subcellular localization of mRNA in a cell. Only limited methods had been built to predict multi-label subcellular localization of mRNA. Most of the existing methods are slow and cannot be implemented at transcriptome scale. In this study, a fast and reliable method had been developed for predicting multi-label subcellular localization of mRNA that can be implemented at genome scale. Firstly, deep learning method based on convolutional neural network method have been developed using one-hot encoding and attained an average AUROC - 0.584 (0.543 - 0.605). Secondly, machine learning based methods have been developed using mRNA sequence composition, our XGBoost classifier achieved an average AUROC - 0.709 (0.668 - 0.732). In addition to alignment free methods, we also developed alignment-based methods using similarity and motif search techniques. Finally, a hybrid technique has been developed that combine XGBoost models and motif-based searching and achieved an average AUROC 0.742 (0.708 - 0.816). Our method - MRSLpred, developed in this study is complementary to the existing method. One of the major advantages of our method over existing methods is its speed, it can scan all mRNA of a transcriptome in few hours. A publicly accessible webserver and a standalone tool has been developed to facilitate researchers (Webserver: https://webs.iiitd.edu.in/raghava/mrslpred/). Key PointsO_LIPrediction of Subcellular localization of mRNA C_LIO_LIClassification of mRNA based on Motif and BLAST search C_LIO_LICombination of alignment based and alignment free techniques C_LIO_LIA fast method for subcellular localization of mRNA C_LIO_LIA web server and standalone software C_LI

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.