Precise and ultrafast tandem repeat variant detection in massively parallel sequencing reads
Wang, X.; Huang, M.; Budowle, B.; Ge, J.
Show abstract
Calling tandem repeat (TR) variants from DNA sequences is of both theoretical and practical significance. A large number of software tools have been developed for detecting TRs. However, little study has been done to detect TR alleles from long-read sequences, and the effectiveness of detecting TR alleles from whole genome sequence (WGS) data still needs to be improved. Herein, a novel algorithm is described to retrieve TR regions from sequence alignment, and a software program, TRcaller, has been developed to call TR alleles from both short- and long-read sequences, both whole genome and targeted sequences generated from multiple sequencing platforms. The results showed that TRcaller could provide substantially higher accuracy in detecting TR alleles with magnitudes faster than the mainstream software tools. TRcaller is able to facilitate scalable, accurate, and ultrafast TR allele calling from large-scale sequence datasets in various applications, such as DNA forensics, medical research, disease diagnosis, evolution, and breeding programs. AvailabilityTRcaller is available at www.trcaller.com.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- An accurate assignment test for extremely low-coverage whole-genome sequence data 93%
- SambaR: an R package for fast, easy and reproducible population-genetic analyses of biallelic SNP datasets 93%
- Simulation with RADinitio Improves RADseq Experimental Design and Sheds Light on Sources of Missing Data 93%
Similar papers in this journal
- ZSeeker: An optimized algorithm for Z-DNA detection in genomic sequences 92%
- Kernel Local Fisher Discriminant Analysis of Principal Components (KLFDAPC) significantly improves the accuracy of predicting geographic origin of individuals 92%
- SatXplor - A comprehensive pipeline for satellite DNA analyses in complex genome assemblies 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.