Dissecting the relationship between haplotypes around ATXN2 CAG repeats and the number of CAA interruptions by long-read sequencing
Lee, B. H.; Chan, J.; McMillan, C.; NYGC ALS Consortium, ; Song, Y.; Amado, D. A.; Wang, K.
Show abstract
CAG repeat expansions in ATXN2 are implicated as risk factors for neurological diseases, including amyotrophic lateral sclerosis (ALS) when 27-33 CAG (intermediate) repeats are present. However, how haplotypes around the repeats and CAA interruptions within the repeats are associated with diseases remains poorly understood. Here, we used long-read sequencing on the Oxford Nanopore technologies (ONT) platform to simultaneously infer haplotypes around ATXN2, the number of CAG repeats, and the number of CAA interruptions. We found that haplotypes around ATXN2 and the number of interruptions show ethnicity-specific and ALS-specific distribution. Three CAA interruptions are present at low prevalence ([~]1%) in control populations in multiple ancestry groups, but high prevalence ([~]55%) in ALS individuals with intermediate repeats. Furthermore, we examined 159 individuals with ALS ([~]90% European ancestry) with intermediate ATXN2 repeats and found a unique haplotype in ALS individuals with three CAA interruptions, which can be tagged by an SNV, rs148019457. We further sequenced 41 individuals (EUR = 39) with neurological diseases with intermediate repeats by ONT, and validated that the rs148019457-G allele is only present in haplotypes with three CAA interruptions. Our study shows that 3 CAA interruptions are rare in healthy controls but are common in individuals with intermediate ATXN2 CAG repeats and neurological disorders, and that rs148019457 tags a specific haplotype with 3 CAA interruptions in individuals of European ancestry. These results have implications for the development of precision genomic medicine for neurological disorders, and the tag SNV may help identify those with interruptions from existing microarray genotyping data.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-Tissue Neocortical Transcriptome-Wide Associations Study Implicates 8 Genes Across 6 Genomic Loci in Alzheimer's Disease 93%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 92%
- Genotype–phenotype correlations and novel molecular insights into the DHX30 -associated neurodevelopmental disorders 92%
Similar papers in this journal
- Tissue-specific and repeat length-dependent somatic instability of the X-linked dystonia parkinsonism-associated CCCTCT repeat 95%
- Transcriptional profiling of Multiple System Atrophy cerebellar tissue highlights differences between the parkinsonian and cerebellar sub-types of the disease 94%
- Single-nucleus RNA-seq identifies Huntington disease astrocyte states 93%
Similar papers in this journal
- Shared regulatory pathways reveal novel genetic correlations between grip strength and neuromuscular disorders 93%
- Transcriptomic changes resulting from STK32B overexpression identifies pathways potentially relevant to essential tremor 91%
- Transcriptome analysis reveals higher levels of mobile element-associated abnormal gene transcripts in temporal lobe epilepsy patients 91%
Similar papers in this journal
- Huntington's disease age at motor onset is modified by the tandem hexamer repeat in TCERG1 94%
- IBD analysis of Australian amyotrophic lateral sclerosis SOD1-mutation carriers identifies five founder events and links sporadic cases to existing ALS families 93%
- Synaptosome microRNAs regulate synapse functions in Alzheimer's disease 91%
Similar papers in this journal
- Repeat length increases disease penetrance and severity in C9orf72 ALS/FTD BAC transgenic mice 94%
- Familial ALS/FTD-associated RNA-Binding deficient TDP-43 mutants cause neuronal and synaptic transcript dysregulation in vitro 92%
- The Dynamic Nature of Genetic Risk for Schizophrenia Within Genes Regulated by FOXP1 During Neurodevelopment 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.