Back

Machine Learning and Deep Learning challenges for building 2'O site prediction

Mostavi, M.; Huang, Y.

2020-05-12 bioinformatics
10.1101/2020.05.10.087189 bioRxiv
Show abstract

2'-O-methylation (2'O) is one of the abundant post-transcriptional RNA modifications which can be found in all types of RNA. Detection and functional analysis of 2'O methylation have become challenging problems for biologists ever since its discovery. This paper addresses computational challenges for building Machine Learning and Deep Learning models for predicting 2'O sites. In particular, the impact of sequence length containing 2'O site, embedding method and the type of predictive model are each investigated separately. 30 different predictive models are built and each showed the impact of the mentioned parameters. The area under the precision-recall and receiving operating characteristics curves are utilized to test imbalanced case scenarios in the real world. By comparing the performance of these models, it is shown that embedding methods are crucial for Machine Learning models. However, they do not improve the performance of Deep Learning models. Furthermore, the best predictive model was further investigated to extract significant nucleotides surrounding 2'O sites. Interestingly, based on the significant score matrix achieved by all 2'O samples, it is depicted that model pays the highest attention at the location that the dominant 2'O motifs exist. Dataset and all of the codes are available at https://github.com/MMostavi/2_O_Me_sitePred

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.