Coral accurately bridges paired-end RNA-seq reads alignment
Shi, Q.; Shao, M.
Show abstract
MotivationThe established high-throughput RNA-seq technologies usually produce paired-end reads. A challenging problem is therefore to computationally infer the alignment of entire fragments given the alignment of the two mate ends. Solving this problem essentially provide longer RNA-seq reads, and hence benefits downstream RNA-seq analysis. ResultsWe introduce Coral, a new tool that can accurately bridge paired-end RNA-seq reads. The core of Coral is a novel optimization formulation that can capture the most reliable bridging path while also filter out false paths. An efficient dynamic programming algorithm is designed to calculate the top N optimum. Coral implements a consensus approach to select the best solution among the N candidates by taking into account the distribution of fragment length. Coral is modular, can be easily incorporated into existing RNA-seq analysis pipeline. We show that Coral can improve transcript assembly by a large margin: on average over 2377 RNA-seq samples from GTEx, the improvement (measured with adjusted precision) is 7.5% and 11.2% when Coral is incorporated with StringTie and Scallop, respectively. AvailabilityCoral is open-source, freely available at GitHub (https://github.com/Shao-Group/coral) and Bioconda. Scripts, datasets and documentations that can reproduce all experimental results in this paper are available at https://github.com/Shao-Group/coraltest.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Toward a dynamic threshold for quality-score distortion in reference-based alignment 95%
- De novo clustering of long-read transcriptome data using a greedy, quality-value based algorithm 94%
- An Efficient, Scalable and Exact Representation of High-Dimensional Color Information Enabled via de Bruijn Graph Search 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.