Impact of sequencing technologies on long non-coding RNA computational identification
Chiquitto, A. G.; Silva, L. O. L.; Oliveira, L. S.; Domingues, D. S.; Paschoal, A. R.
Show abstract
The correct annotation of non-coding RNAs, especially long non-coding RNAs (lncRNAs), is still an important critial challenge in genome analyses. One crucial issue in lncRNA transcript annotation is the transcriptome resource that supports lncRNA loci. Long-read technologies now bring the potential to improve the quality of transcriptome annotation. Consequently, long non-coding RNAs (lncRNA) are probably the most benefited class of transcripts that would have improved annotation using this novel technology. However, there is a gap regarding benchmarking studies that highlighted if the direct use of lncRNA predictors in long-reads makes more precise identification of these transcripts. Considering that these lncRNA tools were not trained with these reads, we want to address: how is the performance of these tools? Are they also able to efficiently identify lncRNAs? We could provide evidence of where and how to make potential better approaches for the lncRNA annotation by understanding these issues. Keywords: Non-coding RNAs, high-throughput sequencing technologies, coding, methods, benchmarking, tools, NGS, transcripts
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Comprehensive benchmark of differential transcript usage analysis for static and dynamic conditions 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 95%
Similar papers in this journal
Similar papers in this journal
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 93%
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 93%
- Comprehensive genome-wide identification of angiosperm upstream ORFs with peptide sequences conserved in various taxonomic ranges using a novel pipeline, ESUCA 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.