HORoSCOPE: Decoding human centromere architecture from short reads using k-mer signatures
Hain, C.; Rausch, T.; Human Genome Structural Variation Consortium, ; Human Pangenome Reference Consortium, ; Korbel, J. O.
Show abstract
By directing kinetochore formation and chromosome segregation, centromeres safeguard genome integrity. Yet, the roles of centromeres in human disease are understudied, as their repetitive architecture comprising -satellite higher-order repeats (HORs) renders them largely inaccessible to short-read sequencing approaches. Here we develop HORoSCOPE (Higher-Order Repeat organization and Size of Centromeres using Oligonucleotide Profiles for Estimation), a computational framework for k-mer-based inference of centromere structure and length from short-read data. Based on a reference atlas of 11,836 human centromeres extracted from completely assembled (telomere-to-telomere) haplotypes, we systematically interrogate chromosome-specific centromere architectures, deriving architecture-specific k-mer signatures as well as centromeric length-informative k-mers from common to rare centromere architectures. Leveraging these diagnostic k-mers, HORoSCOPE achieves a precision of 99.3% and a recall of 99.5% in classifying chromosome-specific centromere architectures from short-read de Bruijn graphs. We perform a population-scale analysis of global centromeric architectures in 4,029 human samples with ancestry from 80 human populations sequenced with short reads, uncovering continental haplotype structure for centromeric regions and highlighting African-enriched rare centromeric architectures. Furthermore, by analyzing 1,359 cancer genomes, we link graded HOR truncation events to arm-level copy-number alterations, and uncover a general dependency of chromosomal rearrangement locations on the position of the centromere dip region (CDR) defining the kinetochore attachment site. HORoSCOPE enables large-scale centromere genomics directly from short-read sample cohorts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gaps and complex structurally variant loci in phased genome assemblies 96%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 95%
- Genome-wide variability in recombination activity is associated with meiotic chromatin organization 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.