Use of short-read RNA-Seq data to identify transcripts that can translate novel ORFs
Erady, C.; Puntambekar, S.; Prabakaran, S.
Show abstract
Identification of as of yet unannotated or undefined novel open reading frames (nORFs) and exploration of their functions in multiple organisms has revealed that vast regions of the genome have remained unexplored or hidden. Present within both protein-coding and noncoding regions, these nORFs signify the presence of a much more diverse proteome than previously expected. Given the need to study nORFs further, proper identification strategies must be in place, especially because they cannot be identified using conventional gene signatures. Although Ribo-Seq and proteogenomics are frequently used to identify and investigate nORFs, in this study, we propose a workflow for identifying nORF containing transcripts using our precompiled database of nORFs with translational evidence, using sample transcript information. Further, we discuss the potential uses of this identification, the caveats involved in such a transcript identification and finally present a few representative results from our analysis of naive mouse B and T cells, human post-mortem brain and cichlid fish transcriptome. Our proposed workflow can identify noncoding transcripts that can potentially translate intronic, intergenic and several other classes of nORFs. One-line summaryA systematic workflow to identify nORF containing transcripts using sample transcript information.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 94%
- Covering all your bases: incorporating intron signal from RNA-seq data 93%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 93%
Similar papers in this journal
- Transcriptome profiling of mouse samples using nanopore sequencing of cDNA and RNA molecules 94%
- Characterization of the nuclear and cytosolic transcriptomes in human brain tissue reveals new insights into the subcellular distribution of RNA transcripts 94%
- Transcriptome-wide high-throughput mapping of protein-RNA occupancy profiles using POP-seq 94%
Similar papers in this journal
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 95%
- FastCAR: Fast Correction for Ambient RNA to facilitate differential gene expression analysis in single-cell RNA-sequencing datasets 93%
- Revealing the Prevalence of Suboptimal Cells and Organs in Reference Cell Atlases: An Imperative for Enhanced Quality Control 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.