Biobank-scale genotyping of Robertsonian translocations reveals hidden structural variation on the human acrocentric chromosomes
Rhie, A.; Kim, J.; Rodriguez-Algarra, F.; Solar, S.; Koren, S.; Antipov, D.; Wilczewski, C. M.; Maxwell, G. L.; Gerton, J.; Paschall, J.; Potapova, T.; Wolfsberg, T. G.; Singh, S.; del Castillo del Rio, S. O.; Human Pangenome Reference Consortium, ; Turner, C.; Rakyan, V. K.; Phillippy, A. M.
Show abstract
Balanced Robertsonian translocation (ROB) is the most common chromosomal rearrangement, with an estimated occurrence of 1:800 in newborns. Carriers are at increased risk of cancer and often diagnosed after facing recurrent miscarriages, infertility, or aneuploid offspring. Genotyping carriers with DNA sequencing has been challenging due to incomplete sequences of the fusion site in the human reference. Only recently, the acrocentric short arms were successfully characterized, including the most common ROB fusion site. A ROB results in loss of two rDNA arrays and its adjacent distal sequences, including the highly conserved distal junction (DJ). Here, we present a novel method to type ROB carriers directly from short reads. Applying to cohorts of healthy newborns (n=4,172) and the UK Biobank (n=490,416), we find candidate ROBs at a consistent frequency (0.11-0.12%). We report uncharacterized structural variations of one DJ loss (2.8-3.4%) or gain (8.4-9.3%) found in near telomere-to-telomere genomes of the HPRC.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Genotyping sequence-resolved copy number variationusing pangenomes reveals paralog-specific global diversityand expression divergence of duplicated genes 97%
- Inferring compound heterozygosity from large-scale exome sequencing data 96%
- Targeted profiling of human extrachromosomal DNA by CRISPR-CATCH 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.