HextractoR: an R package for automatic extraction of hairpins from genome-wide data
Yones, C. A.; Macchiaroli, N.; Kamenetzky, L.; Stegmayer, G.; Milone, D.
Show abstract
Extracting stem-loop sequences (hairpins) from genome-wide data is very important nowadays for some data mining tasks in bioinformatics. The genome preprocessing is very important because it has a strong influence on the later steps and the final results. For example, for novel miRNA prediction, all well-known hairpins must be properly located. Although there are some scripts that can be adapted and put together to achieve this task, they are outdated, none of them guarantees finding correspondence to well-known structures in the genome under analysis, and they do not take advantage of the latest advances in secondary structure prediction. We present here an R package for automatic extraction of hairpins from genome-wide data (HextractorR). HextractoR makes an exhaustive and smart analysis of the genome in order to obtain a very good set of short sequences for further processing. Moreover, genomes can be processed in parallel and with low memory requirements. Results obtained showed that HextractoR has effectively outperformed other methods. HextractoR it is freely available at CRAN and Sourceforge.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comprehensive machine-learning-based analysis of microRNA-target interactions reveals variable transferability of interaction rules across species 95%
- miTAR: a hybrid deep learning-based approach for predicting miRNA targets 94%
- VARUS: Sampling Complementary RNA Reads from the Sequence Read Archive 94%
Similar papers in this journal
Similar papers in this journal
- Evaluating DCA-based method performances for RNA contact prediction by a well-curated dataset 94%
- Unraveling Unbreakable Hairpins: Characterizing RNA secondary structures that are persistent after dinucleotide shuffling 94%
- bpRNA-align: Improved RNA Secondary Structure Global Alignment for Comparing and Clustering RNA Structures 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.