KuPID: Kmer-based Upstream Preprocessing of Long Reads forIsoform Discovery
Borowiak, M.; Yu, Y. W.
Show abstract
Eukaryotic genes can encode multiple protein isoforms based on alternative splicing of their transcribed regions. Most modern novel isoform discovery methods function by identifying and assembling exon splice junctions from an RNAseq sample. However, splice junctions can only be accurately annotated with time-intensive dynamic programming alignment. This manuscript introduces KuPID, a method for preprocessing long RNAseq reads with the goal of better identifying novel isoform transcripts. KuPID utilizes kmer sketching as a pre-filter to quickly pseudo-align reads to known reference isoforms. Full alignment need only then be applied to reads that are most relevant to isoform discovery. Not only does KuPID speed up the discovery pipeline, it also increases downstream accuracy by filtering out extraneous reads. KuPID preprocessing simultaneously increases the f1 accuracy of isoform discovery pipelines by up to 16.7 points while decreasing the runtime by a factor of 2-3x. An optional mode permits a KuPID sample to be paired with both isoform discovery and transcript quantification. Code availability: https://github.com/mboro2000/KuPID.git
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying and quantifying isoforms from accurate full-length transcriptome sequencing reads with Mandalorion 97%
- Enhancing transcriptome expression quantification through accurate assignment of long RNA sequencing reads with TranSigner 96%
- Measuring, visualizing and diagnosing reference bias with biastools 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.