Back

Assessing the impacts of various factors related to identification, conservation, biogenesis, and function on circular RNA reliability

Chuang, T.-J.; Chiang, T.-W.; Chen, C.-Y.

2022-10-28 bioinformatics
10.1101/2022.10.28.514164 bioRxiv
Show abstract

Circular RNAs (circRNAs) are non-polyadenylated RNAs with a continuous loop structure characterized by a non-co-linear back-splice junction (BSJ). While dozens of computational tools have been developed and identified millions of circRNA candidates in diverse species, it remains a major challenge for determining circRNA reliability due to various types of false positives. Here, we systematically assess the impacts of numerous factors related to identification, conservation, biogenesis, and function on circRNA reliability by comparisons of circRNA expression from mock (total RNAs) and the corresponding co-linear/polyadenylated RNA-depleted datasets based on three different RNA treatment approaches. Eight important indicators of circRNA reliability are determined. The relative contribution to variability explained analyses further reveal that the relative importance of these factors in affecting circRNA reliability is conservation level of circRNA > full-length circular sequences > supporting BSJ read count > both BSJ donor and acceptor splice sites at the same co-linear transcript isoforms > both BSJ donor and acceptor splice sites at the annotated exon boundaries > BSJs detected by multiple tools > supporting functional features > both BSJ donor and acceptor splice sites undergoing alternative splicing. By extracting RT-independent circRNAs, circRNAs passing multiple experimental validations, and database-specific circRNAs, we showed the additive effects of these important factors in determining circRNA reliability. This study thus provides a useful guideline and an important resource for selecting high-confidence circRNAs for further investigations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.