Sentences, Words, Attention: A "Transforming" Aphorism of miRNA Discovery
Gupta, S.; Saini, V.; Kumar, R.; Shankar, R.
Show abstract
Discovering pre-miRNAs is the core of miRNA discovery. Using traditional sequence/structural features many tools have been published to discover miRNAs. However, in practical applications like genomic annotations, their actual performance has been far away from acceptable. This becomes more grave in plants where unlike animals pre-miRNAs are much more complex and difficult to identify. This is reflected by the huge gap between the available software for miRNA discovery and species specific miRNAs information for animals and plants. Here, we present miWords, an attention based genomic language processing transformer and context scoring deep-learning approach, with an optional sRNA-seq guided CNN module to accurately identify pre-miRNA regions in plant genomes. During a comprehensive bench-marking the transformer part of miWords alone significantly outperformed the compared published tools with consistent performance while breaching accuracy of 98% across a large number of experimentally validated data. Performance of miWords was also evaluated across Arabidopsis genome where also miWords, even without using its sRNA-seq reads module, outperformed those software which essentially require sRNA-seq reads to identify miRNAs. miWords was run across the Tea genome, reporting 803 pre-miRNA regions, all validated by sRNA-seq reads from multiple samples, and 10 randomly selected cases re-validated by qRT-PCR.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comprehensive machine-learning-based analysis of microRNA-target interactions reveals variable transferability of interaction rules across species 96%
- miTAR: a hybrid deep learning-based approach for predicting miRNA targets 96%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 94%
Similar papers in this journal
Similar papers in this journal
- DBpred: A deep learning method for the prediction of DNA interacting residues in protein sequences 95%
- Blood-based transcriptomic signature panel identification for cancer diagnosis: Benchmarking of feature extraction methods 95%
- PTFSpot: Deep co-learning on transcription factors and their binding regions attains impeccable universality in plants 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.