Detailed tandem repeat allele profiling in 1,027 long-read genomes reveals genome-wide patterns of pathogenicity
Danzi, M. C.; Xu, I. R. L.; Fazal, S.; Dolzhenko, E.; Pellerin, D.; Weisburd, B.; Reuter, C.; Sampson, J.; Folland, C.; Wheeler, M. T.; O'Donnell-Luria, A.; Wuchty, S.; Ravenscroft, G.; Eberle, M. A.; All of Us Research Program Long Read Working Group, ; Zuchner, S.
Show abstract
Tandem repeats are a highly polymorphic class of genomic variation that play causal roles in rare diseases but are notoriously difficult to sequence using short-read techniques1,2. Most previous studies profiling tandem repeats genome-wide have reduced the description of each locus to the singular value of the length of the entire repetitive locus3,4. Here we introduce a comprehensive database of 3.6 billion tandem repeat allele sequences from over one thousand individuals using HiFi long-read sequencing. We show that the previously identified pathogenic loci are among the most variable tandem repeat loci in the genome, when incorporating nucleotide resolution sequence content to measure the longest pure motif segment. More broadly, we introduce a novel measure, tandem repeat constraint, that assists in distinguishing potentially pathogenic from benign loci. Our approach of measuring variation as the length of the longest pure segment successfully prioritizes pathogenic repeats within their previously published linkage regions. We also present evidence for two novel pathogenic repeat expansion candidates. In summary, this analysis significantly clarifies the potential for short tandem repeat pathogenicity at over 1.7 million tandem repeat loci and will aid the identification of disease-causing repeat expansions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Characterising tandem repeat complexities across long-read sequencing platforms with TREAT and otter 96%
- A general framework for identifying rare variant combinations in complex disorders 95%
- Nanopore sequencing of 1000 Genomes Project samples to build a comprehensive catalog of human genetic variation 95%
Similar papers in this journal
- Advancing long-read nanopore genome assembly and accurate variant calling for rare disease detection 96%
- Impact of genome build on RNA-seq interpretation and diagnostics 96%
- A survey of rare epigenetic variation in 23,116 human genomes identifies disease-relevant epivariations and novel CGG expansions 96%
Similar papers in this journal
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 95%
- Multi-modal investigation of the schizophrenia-associated 3q29 genomic interval reveals global genetic diversity with unique haplotypes and segments that increase the risk for non-allelic homologous recombination 95%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.