QueryFuse Is A Sensitive Algorithm For Detection Of Gene-Specific Fusions
Tan, Y.
Show abstract
Recurrent chromosomal translocations, known as fusions, play important roles in carcinogenesis. They can serve as valuable diagnostic and therapeutic targets. RNA-seq is an ideal platform for detecting transcribed fusions, and computational methods have been developed to identify fusion transcripts from RNA-seq data. However, some transciptome realignment procedures for these methods are unnecessary, making this task computationally expensive and time consuming. Therefore, we have developed QueryFuse, a novel hypothesis-based algorithm that identifies gene-specific fusion from pre-aligned RNA-seq data. It is designed to help biologists quickly find and/or computationally validate fusions of interest, together with visualization and detailed properties of supporting reads. By aligning reads to Query genes at the pre-processing step with a more sensitive, memory intensive local aligner, QueryFuse can reduce alignment time and improve detection sensitivity. QueryFuse performed better or at comparable levels with two popular tools (deFuse and TopHatFusion) on both simulated and well-annotated cell-line datasets. Finally, using QueryFuse, we identified a novel fusion event with a potential therapeutic implication in clinical samples. Taken together, our results showed that QueryFuse is efficient and reliable for detecting gene-specific fusion events.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Boosting variant-calling performance with multi-platform sequencing data using Clair3-MP 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 96%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 96%
Similar papers in this journal
- Accurate Identification of Extrachromosomal Circular DNA from Long-read Sequences. 96%
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 96%
- Assessing deep learning algorithms in cis-regulatory motif finding based on genomic sequencing data 95%
Similar papers in this journal
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.