Assessment of mapping strategies for determining the 5'-end of mRNAs and long-noncoding RNAs with short read sequences
Noguchi, S.; Kawaji, H.; Kasukawa, T.
Show abstract
BackgroundGenome mapping is an essential step in data processing for transcriptome analysis, and many previous studies have evaluated various methods and strategies for mapping RNA-seq data. Cap Analysis of Gene Expression (CAGE) is a sequencing-based protocol particularly designed to capture the 5{square}-ends of transcripts for quantitatively measuring the expression levels of transcription start sites genome-wide. Because CAGE analysis can also predict the activities of promoters and enhancers, this protocol has been an essential tool in studies of transcriptional regulation. Typically, the same mapping software is used to align both RNA-seq data and CAGE reads to a reference genome, but which mapping software and options are most appropriate for mapping the 5{square}-end sequence reads obtained through CAGE has not previously been evaluated systematically. ResultsHere we assessed various strategies for aligning CAGE reads, particularly [~]50-bp sequences, with the human genome by using the HISAT2, LAST, and STAR programs both with and without a reference transcriptome. One of the major inconsistencies among the tested strategies involves alignments to pseudogenes and parent genes: some of the strategies prioritized alignments with pseudogenes even when the read could be aligned with coding genes with fewer mismatches. Another inconsistency concerned the detection of exon-exon junctions. These preferences depended on the program applied and whether a reference transcriptome was included. Overall, the choice of strategy yielded different mapping results for approximately 2% of all promoters. ConclusionsAlthough the various alignment strategies produced very similar results overall, we noted several important and measurable differences. In particular, using the reference transcriptome in STAR yielded alignments with the fewest mismatches. In addition, the inconsistencies among the strategies were especially noticeable regarding alignments to pseudogenes and novel splice junctions. Our results indicate that the choice of alignment strategy is important because it might affect the biological interpretation of the data.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
- Covering all your bases: incorporating intron signal from RNA-seq data 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 94%
Similar papers in this journal
- Common tissue-specific expressions and regulatory mechanisms of c-KIT isoforms with and without GNNK and GNSK sequences across five mammals 94%
- Finding differentially expressed sRNA-Seq regions with srnadiff 94%
- Crinet: A computational tool to infer genome-wide competing endogenous RNA (ceRNA) interactions 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.