Deep learning models for identification of splice junctions across species
Dutta, A.; Singh, K. K.; Anand, A.
Show abstract
Deep learning models like convolutional neural networks (CNN) and recurrent neural networks (RNN) have been frequently used to identify splice sites from genome sequences. Most of the deep learning applications identify splice sites from a single species. Furthermore, the models generally identify and interpret only the canonical splice sites. However, a model capable of identifying both canonical and non-canonical splice sites from multiple species with comparable accuracy is more generalizable and robust. We choose some state-of-the-art CNN and RNN models and compare their performances in identifying novel canonical and non-canonical splice sites in homo sapiens, mus musculus, and drosophila melanogaster. The RNN-based model named SpliceViNCI outperforms its counterparts in identifying splice sites from multiple species as well as on unseen species. SpliceViNCI maintains its performance when trained with imbalanced data making it more robust. We observe that all the models perform better when trained with more than one species. SpliceViNCI outperforms the counterparts when trained with such an augmented dataset. We further extract and compare the features learned by SpliceViNCI when trained with single and multiple species. We validate the extracted features with knowledge from the literature.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rapid Reconstruction of Time-varying Gene Regulatory Networks with Limited Main Memory 94%
- Floating search methodology for combining classification models for site recognition in DNA sequences 94%
- GenoM7GNet: An Efficient N7-methylguanosine Site Prediction Approach Based on a Nucleotide Language Model 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.