Defining a tandem repeat catalog and variation clusters for genome-wide analyses and population databases
Weisburd, B.; Dolzhenko, E.; Bennett, M. F.; Danzi, M. C.; English, A.; Hiatt, L.; Tanudisastro, H.; Kurtas, N. E.; Jam, H. Z.; Brand, H.; Sedlazeck, F. J.; Gymrek, M.; Dashnow, H.; Eberle, M. A.; Rehm, H. L.
Show abstract
Tandem repeat (TR) catalogs are important components of repeat genotyping studies as they define the genomic coordinates and expected motifs of all TR loci being analyzed. In recent years, genome-wide studies have used catalogs ranging in size from fewer than 200,000 to over 7 million loci. Where these catalogs overlapped, they often disagreed on locus boundaries, hindering the comparison and reuse of results across studies. Now, with multiple groups developing public databases of TR variation in large population cohorts, there is a risk that, without sufficient consensus in the choice of locus definitions, the use of divergent repeat catalogs will lead to confusion, fragmentation, and incompatibility across resources. In this paper, we compare existing TR catalogs and discuss desirable features of a comprehensive genome-wide catalog. We then present a new, richly annotated catalog designed for large-scale analyses and population databases. This new catalog, which we call the TRExplorer catalog v1.0, contains 4.86 million TR loci and, unlike most catalogs, is designed to be useful for both short-read and long-read analyses. It consists of 4,803,366 STRs and 59,675 VNTRs, of which 780,607 STRs and 21,888 VNTRs are both polymorphic and entirely absent from widely-used catalogs previously developed for short-read analyses. Additionally, our catalog stratifies TRs into two groups: 1) isolated TRs suitable for repeat copy number analysis using short-read or long-read data and 2) so-called variation clusters that contain TRs within wider polymorphic regions that are best studied through sequence-level analysis. To define variation clusters, we present a novel algorithm that leverages long-read HiFi sequencing data to group repeats with surrounding polymorphisms. We show that the human genome contains at least 25,000 complex variation clusters, most of which span over 120 bp and contain five or more TRs. Resolving the sequence of entire variation clusters instead of individually genotyping constituent TRs leads to a more accurate analysis of these regions and enables us to profile variation that would have been missed otherwise. We also share the trexplorer.broadinstitute.org portal which allows anyone to search, visualize, and download the catalog along with variation clusters and annotations.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Estimating gene conversion rates from population data using multi-individual identity by descent 95%
- HiFi long-read genomes for difficult-to-detect clinically relevant variants 95%
- A phenome-wide association study identifies effects of copy number variation of VNTRs and multicopy genes on multiple human traits 94%
Similar papers in this journal
Similar papers in this journal
- Benchmarking challenging small variants with linked and long reads 96%
- Targeted long-read sequencing facilitates phased diploid assembly and genotyping of the human T cell receptor alpha, delta and beta loci 95%
- Polymorphic short tandem repeats make widespread contributions to blood and serum traits 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.