Back

Efficient Identification of Short Tandem Repeats via Context-Aware Motif Discovery and Ultra-Fast Sequence Alignment

Liao, X.; Wen, L.; Jing, M.; Li, X.; Chen, B.; Zhang, B.; Gao, X.; Shang, X.

2025-11-28 bioinformatics
10.1101/2025.11.25.690584 bioRxiv
Show abstract

Tandem repeats (TRs) are highly polymorphic genomic elements, associated with diverse molecular traits and implicated in numerous human diseases. However, large-scale analysis of TRs has been limited by computational challenges, including motif recognition, detection in complex regions, and excessive computational cost. Here we present FastSTR, a computationally efficient tool for precise detection and characterization of TRs. FastSTR integrates a context-aware N-gram motif model with a segmented global alignment algorithm to enable accurate motif identification and boundary definition, even for repeat units up to 8 bp. Across 13 species, FastSTR achieved >90% recall and 99% precision, running several times faster than existing methods white outperforming them in both sensitivity and accuracy. Applied to the human genome, FastSTR uncovered previously unannotated HSATII elements, resolved population-specific TR demonstrate, and identified recurrent STR alterations in lung cancer. These results demonstrate FastSTR as a versatile framework for TR annotation and discovery, advancing studies of genome evolution, genetic diversity, and disease.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.